What we measured so far
50 short coding prompts, the same for both models, run with Ollama on a Kaggle T4 GPU at temperature 0. v1 was measured on an RTX 3050 laptop. The base model is Qwen2.5 Coder 1.5B Instruct, the model Nock Coder was tuned from.
This page will grow as more tests finish. We publish the numbers we get, including the ones that are not flattering.
- Answer lengthWords per answer, lower is better
- Explanation around the codeWords outside code blocks, lower is better
- Starts with codeShare of answers, higher is better
- No code at allShare of answers, lower is better
- SpeedTokens per second on a Kaggle T4 GPU
- HumanEvalPython problems whose answer passes the tests, higher is better
- SolidityContract tasks whose answer compiles with solc, higher is better
Each model used its own default system prompt. Compiling is a low bar: it does not mean the contract is correct or safe. We publish every answer.
Raw outputsHow we tested
- Models
- Nock Coder from Nock-AI/nock-coder-1.5b-GGUF, and qwen2.5-coder:1.5b from the Ollama library. Both are 4 bit quantized files of about 1 GB.
- Settings
- Temperature 0, fixed seed, up to 900 tokens, one answer per prompt. Each model used its own default system prompt.
- Prompts
- 50 short requests across Python, JavaScript, TypeScript, SQL, shell, Git, Go, Rust, Java, C, regex and web.
- HumanEval
- 164 Python problems. Each answer is run against the tests that come with the problem; it passes or it fails.
- Solidity
- 15 contract tasks. Each answer is compiled with solc 0.8, unedited. Compiling says nothing about whether a contract is correct or safe.
- Where it ran
- Both models on a Kaggle T4 GPU, with one exception: on Kaggle, Ollama stopped the base model's first Solidity answer because it kept repeating itself, and that stage ended there. We reran the base model's 15 Solidity tasks on an RTX 3050 laptop with the same settings; the aborted answer counts as a failure.
- Counting
- Words are counted by splitting the answer on spaces. Explanation words are the words outside code blocks. An answer starts with code when its first characters are a code fence.
- Limits
- One run per model. Minor differences mean little: on HumanEval, 8 problems separate the two models. The shorter answers and the code first habit are large enough to be real.