NockAI

What we measured so far

50 short coding prompts, the same for both models, run with Ollama on a Kaggle T4 GPU at temperature 0. v1 was measured on an RTX 3050 laptop. The base model is Qwen2.5 Coder 1.5B Instruct, the model Nock Coder was tuned from.

This page will grow as more tests finish. We publish the numbers we get, including the ones that are not flattering.

Answer lengthWords per answer, lower is better
Nock v251 words
Nock v148 words
Base model185 words
Explanation around the codeWords outside code blocks, lower is better
Nock v214 words
Nock v117 words
Base model134 words
Starts with codeShare of answers, higher is better
Nock v250 of 50
Nock v127 of 50
Base model1 of 50
No code at allShare of answers, lower is better
Nock v20 of 50
Nock v112 of 50
Base model0 of 50
SpeedTokens per second on a Kaggle T4 GPU
Nock v2110 per second
Base model133 per second
HumanEvalPython problems whose answer passes the tests, higher is better
Nock v2104 of 164
Base model112 of 164
SolidityContract tasks whose answer compiles with solc, higher is better
Nock v211 of 15
Base model8 of 15

Each model used its own default system prompt. Compiling is a low bar: it does not mean the contract is correct or safe. We publish every answer.

Raw outputs

How we tested

Models
Nock Coder from Nock-AI/nock-coder-1.5b-GGUF, and qwen2.5-coder:1.5b from the Ollama library. Both are 4 bit quantized files of about 1 GB.
Settings
Temperature 0, fixed seed, up to 900 tokens, one answer per prompt. Each model used its own default system prompt.
Prompts
50 short requests across Python, JavaScript, TypeScript, SQL, shell, Git, Go, Rust, Java, C, regex and web.
HumanEval
164 Python problems. Each answer is run against the tests that come with the problem; it passes or it fails.
Solidity
15 contract tasks. Each answer is compiled with solc 0.8, unedited. Compiling says nothing about whether a contract is correct or safe.
Where it ran
Both models on a Kaggle T4 GPU, with one exception: on Kaggle, Ollama stopped the base model's first Solidity answer because it kept repeating itself, and that stage ended there. We reran the base model's 15 Solidity tasks on an RTX 3050 laptop with the same settings; the aborted answer counts as a failure.
Counting
Words are counted by splitting the answer on spaces. Explanation words are the words outside code blocks. An answer starts with code when its first characters are a code fence.
Limits
One run per model. Minor differences mean little: on HumanEval, 8 problems separate the two models. The shorter answers and the code first habit are large enough to be real.