Result A
Qwen3 0.6B
llama.cpp 9a286ac · CPU · Matched server request
- Generation
- 6.725 tok/s
- Prompt
- 10.48 tok/s
- Peak
- 81 °C
Three measured requests completed with a clean kernel window.
Open Qwen3 0.6B resultsResult A is the starting point. Result B is the runtime, backend, thread count, or MTP setting you want to check against it. The presets below only pair rows that belong together.
Each option is a prepared pair from one test group. Result A is the baseline. Result B is the setting being checked against it.
Four exact Q4_K_M GGUFs received the same raw prompt, context, thread count, batch settings, sampler, and 96-token output through Ollama 0.32.1 and native llama.cpp commit 9a286ac.
Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.
Result A
llama.cpp 9a286ac · CPU · Matched server request
Three measured requests completed with a clean kernel window.
Open Qwen3 0.6B resultsResult B
Ollama 0.32.1 · CPU · Matched server request
Ollama reported 16.11% higher internal generation throughput than the paired llama.cpp row.
Open Qwen3 0.6B resultsResult B generated 16.1% faster than Result A in this test group. The cards below keep the artifact, runtime, backend, mode, temperature, and outcome next to that difference.
4 prepared comparisons. Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.
Open full test2 prepared comparisons. Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.
Open full test1 prepared comparisons. The four-thread row used a disclosed 90 °C abort after the first attempt reached the normal 85 °C guard.
Open full test12 prepared comparisons. Each on/off pair belongs to its own requested draft depth. Do not compare the baseline from one pair as if it were the same request block as another.
Open full test