Matched Ollama and llama.cpp CPU Test
Four exact Q4_K_M GGUFs received the same raw prompt, context, thread count, batch settings, sampler, and 96-token output through Ollama 0.32.1 and native llama.cpp commit 9a286ac.
Local AI Benchmarks, Inference Results, and Hardware Testing
The results outgrew the articles. At some point, another wall of bash output stopped being useful. The Hub puts the individual runs, hardware, model files, speeds, thermals, failures, and source material in one place. Now you can search by device or model and pull up the numbers, test settings, compatibility results, and evidence behind each run.
Comparison boundary: compare rows inside the same test group. Matched server requests, pp128/tg96, RKLLM NPU tests, and outside Pi references answer different questions.
| Model and artifact | Runtime and mode | Generation | Prompt | Peak | State | Full test |
|---|---|---|---|---|---|---|
| Qwen3 0.6B Q4_K_M | llama.cpp 9a286ac CPU · Matched server request | 6.725 tok/s | 10.48 tok/s | 81 °C | Published result | Matched CPU runtimes |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S. Three measured requests completed with a clean kernel window. Recorded setup
Artifact SHA-256 | ||||||
| Qwen3 0.6B Q4_K_M | Ollama 0.32.1 CPU · Matched server request | 7.809 tok/s | 13.24 tok/s | 81 °C | Published result | Matched CPU runtimes |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S. Ollama reported 16.11% higher internal generation throughput than the paired llama.cpp row. Recorded setup
Artifact SHA-256 | ||||||
| Qwen3 1.7B Q4_K_M | llama.cpp 9a286ac CPU · Matched server request | 2.852 tok/s | 3.849 tok/s | 83 °C | Published result | Matched CPU runtimes |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S. Three measured requests completed with a clean kernel window. Recorded setup
Artifact SHA-256 | ||||||
| Qwen3 1.7B Q4_K_M | Ollama 0.32.1 CPU · Matched server request | 3.321 tok/s | 5.075 tok/s | 83 °C | Published result | Matched CPU runtimes |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S. Ollama reported 16.46% higher internal generation throughput than the paired llama.cpp row. Recorded setup
Artifact SHA-256 | ||||||
| Qwen3 4B Instruct Q4_K_M | llama.cpp 9a286ac CPU · Matched server request | 1.484 tok/s | 1.899 tok/s | 81 °C | Published result | Matched CPU runtimes |
Open result detailsWhy this model is hereThe Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan. It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S. Three measured requests completed with a clean kernel window. Recorded setup
Artifact SHA-256 | ||||||
| Qwen3 4B Instruct Q4_K_M | Ollama 0.32.1 CPU · Matched server request | 2.002 tok/s | 2.789 tok/s | 81 °C | Published result | Matched CPU runtimes |
Open result detailsWhy this model is hereThe Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan. It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S. Ollama reported 34.89% higher internal generation throughput than the paired llama.cpp row. Recorded setup
Artifact SHA-256 | ||||||
| Phi-4 Mini Q4_K_M | llama.cpp 9a286ac CPU · Matched server request | 1.57 tok/s | 2.115 tok/s | 82 °C | Published result | Matched CPU runtimes |
Open result detailsWhy this model is herePhi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang. It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S. Three measured requests completed with a clean kernel window. Recorded setup
Artifact SHA-256 | ||||||
| Phi-4 Mini Q4_K_M | Ollama 0.32.1 CPU · Matched server request | 2.072 tok/s | 3.131 tok/s | 82 °C | Published result | Matched CPU runtimes |
Open result detailsWhy this model is herePhi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang. It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S. Ollama reported 32.00% higher internal generation throughput than the paired llama.cpp row. Recorded setup
Artifact SHA-256 | ||||||
| Qwen3 0.6B Q4_K_M | llama.cpp 9a286ac CPU · True CPU | 7.397 tok/s | 10.67 tok/s | 81 °C | Published result | CPU and Vulkan modes |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S. Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window. Recorded setup
| ||||||
| Qwen3 0.6B Q4_K_M | llama.cpp 9a286ac CPU + GPU · Mixed host operations | 7.425 tok/s | 34.59 tok/s | 80 °C | Published result | CPU and Vulkan modes |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S. Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved. Recorded setup
| ||||||
| Qwen3 0.6B Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | 8.593 tok/s | 37.35 tok/s | 57 °C | Published result | CPU and Vulkan modes |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. Full Vulkan improved generation 16.2% over true CPU and cut the recorded package peak by 24 °C. Recorded setup
| ||||||
| Qwen3 1.7B Q4_K_M | llama.cpp 9a286ac CPU · True CPU | 2.986 tok/s | 3.893 tok/s | 84 °C | Published result | CPU and Vulkan modes |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S. Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window. Recorded setup
| ||||||
| Qwen3 1.7B Q4_K_M | llama.cpp 9a286ac CPU + GPU · Mixed host operations | 2.997 tok/s | 12.37 tok/s | 83 °C | Published result | CPU and Vulkan modes |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S. Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved. Recorded setup
| ||||||
| Qwen3 1.7B Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | 3.462 tok/s | 12.88 tok/s | 58 °C | Published result | CPU and Vulkan modes |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. Full Vulkan improved generation 15.9% over true CPU and cut the recorded package peak by 26 °C. Recorded setup
| ||||||
| Qwen3 4B Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | n/a | n/a | n/a | Failed run | CPU and Vulkan modes |
Open result detailsWhy this model is hereThe Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan. It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. The run reached an i915 reset timeout. No speed score is published for it. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Phi-4 Mini Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | n/a | n/a | n/a | Failed run | CPU and Vulkan modes |
Open result detailsWhy this model is herePhi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang. It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. The run window contained an i915 GPU hang even though llama-bench returned zero. The kernel event overrides the process exit code. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Qwen3 8B Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | n/a | n/a | n/a | Failed run | CPU and Vulkan modes |
Open result detailsWhy this model is hereQwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang. It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. The run window contained an i915 GPU hang even though llama-bench returned zero. No clean performance row is claimed. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Gemma 3 4B Text-only derivative | llama.cpp 9a286ac Vulkan · Full Vulkan | n/a | n/a | n/a | Failed run | CPU and Vulkan modes |
Open result detailsWhy this model is hereA separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss. It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix. What ranText-only derivative ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. Fence and preemption timeouts led to an i915 reset and vk::DeviceLostError. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Ling-mini-2.0 IQ4_XS, pinned revision 8be84a0 | llama.cpp 9a286ac CPU · 3 threads, 85 °C guard | 3.091 tok/s | 4.375 tok/s | 74 °C | Published result | Ling-mini thread test |
Open result detailsWhy this model is hereThe exact pinned IQ4_XS file completed true-CPU tests at three and four threads. The thread test shows a real speed and temperature tradeoff on the N5095 using one verified 8.8 GB artifact rather than a mutable model tag. What ranIQ4_XS, pinned revision 8be84a0 ran through llama.cpp 9a286ac using CPU in 3 threads, 85 °C guard mode on Youyeetoo X1S. The clean three-thread row stayed well below the normal thermal guard. Recorded setup
Artifact SHA-256 | ||||||
| Ling-mini-2.0 IQ4_XS, pinned revision 8be84a0 | llama.cpp 9a286ac CPU · 4 threads, 90 °C guard | 4.083 tok/s | 5.817 tok/s | 85 °C | Published result | Ling-mini thread test |
Open result detailsWhy this model is hereThe exact pinned IQ4_XS file completed true-CPU tests at three and four threads. The thread test shows a real speed and temperature tradeoff on the N5095 using one verified 8.8 GB artifact rather than a mutable model tag. What ranIQ4_XS, pinned revision 8be84a0 ran through llama.cpp 9a286ac using CPU in 4 threads, 90 °C guard mode on Youyeetoo X1S. The selected four-thread row improved generation 32.1% and completed under the disclosed 90 °C ceiling. Recorded setup
Artifact SHA-256 | ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Depth 1, MTP off | 11.34 tok/s | n/a | 74 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 1. Recorded setup
| ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 1, MTP on | 6.043 tok/s | n/a | 75 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S. MTP was active and ran 46.73% slower than its paired baseline. Recorded setup
| ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Depth 2, MTP off | 11.49 tok/s | n/a | 77 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 2. Recorded setup
| ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 2, MTP on | 4.019 tok/s | n/a | 79 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S. MTP was active and ran 65.03% slower than its paired baseline. Recorded setup
| ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Depth 3, MTP off | 11.21 tok/s | n/a | 76 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 3. Recorded setup
| ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 3, MTP on | 3.133 tok/s | n/a | 77 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S. MTP was active and ran 72.06% slower than its paired baseline. Recorded setup
| ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Depth 4, MTP off | 11.16 tok/s | n/a | 78 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 4. Recorded setup
| ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 4, MTP on | 2.624 tok/s | n/a | 79 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S. MTP was active and ran 76.48% slower than its paired baseline. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Depth 1, MTP off | 4.891 tok/s | n/a | 78 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 1. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 1, MTP on | 3.778 tok/s | n/a | 77 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S. MTP was active and ran 22.76% slower than its paired baseline. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Depth 2, MTP off | 5.154 tok/s | n/a | 79 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 2. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 2, MTP on | 2.853 tok/s | n/a | 79 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S. MTP was active and ran 44.65% slower than its paired baseline. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Depth 3, MTP off | 4.853 tok/s | n/a | 80 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 3. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 3, MTP on | 2.599 tok/s | n/a | 80 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S. MTP was active and ran 46.45% slower than its paired baseline. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Depth 4, MTP off | 5.121 tok/s | n/a | 80 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 4. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 4, MTP on | 1.522 tok/s | n/a | 79 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S. MTP was active and ran 70.28% slower than its paired baseline. Recorded setup
| ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Depth 1, MTP off | 3.269 tok/s | n/a | 80 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 1. Recorded setup
| ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU + draft MTP · Depth 1, MTP on | 2.266 tok/s | n/a | 84 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S. MTP was active and ran 30.68% slower than its paired baseline. Recorded setup
| ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Depth 2, MTP off | 3.359 tok/s | n/a | 80 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 2. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard. Recorded setup
| ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU + draft MTP · Depth 2, MTP on | 1.96 tok/s | n/a | 84 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S. MTP was active and ran 41.63% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard. Recorded setup
| ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Depth 3, MTP off | 3.352 tok/s | n/a | 82 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 3. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard. Recorded setup
| ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU + draft MTP · Depth 3, MTP on | 1.474 tok/s | n/a | 86 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S. MTP was active and ran 56.03% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard. Recorded setup
| ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Depth 4, MTP off | 3.397 tok/s | n/a | 81 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 4. Recorded setup
| ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU + draft MTP · Depth 4, MTP on | 1.333 tok/s | n/a | 84 °C | Published result | MTP depth sweep |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S. MTP was active and ran 60.75% slower than its paired baseline. Recorded setup
| ||||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Two-run warm average | 10.44 tok/s | n/a | 70 °C | Published result | Additional models |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S. The fastest normal text-generation result in the X1S work. Throughput does not establish answer quality. Recorded setup
| ||||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Two-run warm average | 4.89 tok/s | n/a | 75 °C | Published result | Additional models |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S. A clean middle-tier CPU result from the normal non-MTP Ollama path. Recorded setup
| ||||||
| Qwen3.5 9B Q4_K_M | Ollama 0.32.1 CPU · Five-prompt average | 1.11 tok/s | n/a | 82 °C | Published result | Additional models |
Open result detailsWhy this model is hereThe official Q4_K_M Ollama artifact fit in 16 GB and averaged 1.110 tok/s across five prompts. This was a fit-and-completion test for a much larger model on the X1S, not a claim that 1.110 tok/s is comfortable interactive speed. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Five-prompt average mode on Youyeetoo X1S. The 9B artifact fit in 16 GB and completed five prompts, but generation requires patience. Recorded setup
Artifact SHA-256 | ||||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Two-run warm average | 3.38 tok/s | n/a | 77 °C | Published result | Additional models |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S. The normal non-MTP Ollama result. The separate MTP sweep is recorded in its own test group. Recorded setup
| ||||||
Gemma 4 E4B Q4_K_M | Ollama 0.32.1 CPU · Two-run warm average | 1.78 tok/s | n/a | 78 °C | Published result | Additional models |
Open result detailsWhy this model is hereGemma 4 E4B completed the original warm Ollama test at 1.78 tok/s. The row is useful as a recorded X1S operating point, but it does not have enough separate material for an indexable model page yet. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S. A clean normal Ollama row retained as an operating point, not a quality ranking. Recorded setup
| ||||||
Granite 4 Tiny-H Original Ollama tag | Ollama 0.32.1 CPU · Single original observation | 5.43 tok/s | n/a | 79 °C | Partial result | Additional models |
Open result detailsWhy this model is hereThe X1S recorded one 5.43 tok/s Ollama observation. One observation is retained for completeness but is not promoted into a deeper standalone page. What ranOriginal Ollama tag ran through Ollama 0.32.1 using CPU in Single original observation mode on Youyeetoo X1S. One retained observation, not a repeated average. | ||||||
LFM2.5 Mutable latest tag | Ollama 0.32.1 CPU · Single recorded tag state | 4.95 tok/s | n/a | 80 °C | Partial result | Additional models |
Open result detailsWhy this model is hereThe X1S recorded one 4.95 tok/s result from a mutable latest tag. The tag was not content-pinned, so the row stays visible as a partial result without becoming its own search page. What ranMutable latest tag ran through Ollama 0.32.1 using CPU in Single recorded tag state mode on Youyeetoo X1S. The mutable latest tag was not content-pinned, so this remains a partial result. | ||||||
Gemma 3n E2B gemma3n:e2b | Ollama 0.32.1 CPU · Single original observation | 3.24 tok/s | n/a | 80 °C | Partial result | Additional models |
Open result detailsWhy this model is hereThe X1S recorded one 3.24 tok/s Ollama result. The row is kept separate from Gemma 4 E2B and does not yet have enough unique material for its own page. What rangemma3n:e2b ran through Ollama 0.32.1 using CPU in Single original observation mode on Youyeetoo X1S. One retained observation. This is Gemma 3n E2B, not Gemma 4 E2B. | ||||||
| BitCPM-CANN 1B Official TQ2_0 GGUF | llama.cpp 9a286ac CPU · True CPU, five repetitions | 8.584 tok/s | 13.01 tok/s | 76 °C | Published result | Runtime compatibility |
Open result detailsWhy this model is hereThe exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1. This is a runtime-support result, not a speed contest. The same verified file produced inference in one runtime and a tensor-size overflow in the other. What ranOfficial TQ2_0 GGUF ran through llama.cpp 9a286ac using CPU in True CPU, five repetitions mode on Youyeetoo X1S. Pinned llama.cpp accepted the verified file and completed the controlled true-CPU workload. Recorded setup
Artifact SHA-256 | ||||||
| BitCPM-CANN 1B Official TQ2_0 GGUF | Ollama 0.32.1 CPU · Model load | n/a | n/a | n/a | Failed run | Runtime compatibility |
Open result detailsWhy this model is hereThe exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1. This is a runtime-support result, not a speed contest. The same verified file produced inference in one runtime and a tensor-size overflow in the other. What ranOfficial TQ2_0 GGUF ran through Ollama 0.32.1 using CPU in Model load mode on Youyeetoo X1S. Ollama rejected the same verified file with a tensor-size overflow before inference. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. Artifact SHA-256 | ||||||
| Gemma 3 4B Separate text-only derivative | llama.cpp 9a286ac CPU · True CPU, five repetitions | 1.624 tok/s | 2.172 tok/s | 81 °C | Published result | Runtime compatibility |
Open result detailsWhy this model is hereA separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss. It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix. What ranSeparate text-only derivative ran through llama.cpp 9a286ac using CPU in True CPU, five repetitions mode on Youyeetoo X1S. The derived copy completed true CPU. The original multimodal source artifact was not overwritten. Recorded setup
Artifact SHA-256 | ||||||
| Qwen3 0.6B Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 6.788 tok/s | n/a | 74 °C | Published result | Original CPU matrix |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Fastest row, least complete response on the single prompt. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Qwen3 1.7B Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 3.129 tok/s | n/a | 77 °C | Published result | Original CPU matrix |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Best interactive starting tier tested. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Qwen3 4B Instruct Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 1.824 tok/s | n/a | 77 °C | Published result | Original CPU matrix |
Open result detailsWhy this model is hereThe Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan. It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Patient local batch use. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Phi-4 Mini Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 1.991 tok/s | n/a | 77 °C | Published result | Original CPU matrix |
Open result detailsWhy this model is herePhi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang. It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Patient local batch use. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Gemma 3 4B Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 1.995 tok/s | n/a | 77 °C | Published result | Original CPU matrix |
Open result detailsWhy this model is hereA separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss. It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Patient local batch use. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Qwen3 8B Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 0.924 tok/s | n/a | 80 °C | Published result | Original CPU matrix |
Open result detailsWhy this model is hereQwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang. It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Fit comfortably with sub-1 tok/s generation. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| DeepSeek-R1 Distill Qwen 1.5B W8A8 RKLLM | RKLLM 1.2.1 NPU · 50 state-capitals prompt | 11.5 tok/s | n/a | n/a | Published result | Nova NPU models |
Open result detailsWhy this model is hereThe Nova NPU run started at 11.5 tok/s, then the state-capitals answer degraded into invented cities. It is the clearest example in the Hub of why throughput and useful output cannot be treated as the same result. What ranW8A8 RKLLM ran through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode on Indiedroid Nova. Started fast, then invented cities around state seven. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Qwen 2.5 3B Instruct W8A8 RKLLM | RKLLM 1.2.1 NPU · 50 state-capitals prompt | 7 tok/s | n/a | n/a | Published result | Nova NPU models |
Open result detailsWhy this model is hereThe Nova NPU result reached 7.0 tok/s and returned all 50 state capitals correctly. It was slower than the smaller DeepSeek row but produced the useful answer on the shared prompt. What ranW8A8 RKLLM ran through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode on Indiedroid Nova. Returned all 50 state capitals correctly. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Llama 3.1 8B Instruct W8A8 RKLLM | RKLLM 1.2.1 NPU · 50 state-capitals prompt | 3.72 tok/s | n/a | n/a | Published result | Nova NPU models |
Open result detailsWhy this model is hereThe Nova NPU measured 3.72 tok/s. A separate 1.99 tok/s Pi 5 value remains labeled as an outside reference. It connects Trevor-measured Nova data with a clearly separated published Pi reference without presenting the two methods as a matched hardware contest. What ranW8A8 RKLLM ran through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode on Indiedroid Nova. Measured Nova NPU result. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| DeepSeek-R1 Distill Qwen 1.5B Q4 reference | llama.cpp CPU · Published reference range | about 6 to 8 tok/s | n/a | n/a | External reference | Pi 5 references |
Open result detailsWhy this model is hereThe Nova NPU run started at 11.5 tok/s, then the state-capitals answer degraded into invented cities. It is the clearest example in the Hub of why throughput and useful output cannot be treated as the same result. What ranQ4 reference ran through llama.cpp using CPU in Published reference range mode on Raspberry Pi 5. Approximate published range. Not Trevor-measured. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Qwen 2.5 3B Instruct Q4 reference | llama.cpp CPU · Published reference range | about 4 to 5 tok/s | n/a | n/a | External reference | Pi 5 references |
Open result detailsWhy this model is hereThe Nova NPU result reached 7.0 tok/s and returned all 50 state capitals correctly. It was slower than the smaller DeepSeek row but produced the useful answer on the shared prompt. What ranQ4 reference ran through llama.cpp using CPU in Published reference range mode on Raspberry Pi 5. Approximate published range. Not Trevor-measured. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
| Llama 3.1 8B Instruct Q4 reference | llama.cpp CPU · Jeff Geerling published result | 1.99 tok/s | n/a | n/a | External reference | Pi 5 references |
Open result detailsWhy this model is hereThe Nova NPU measured 3.72 tok/s. A separate 1.99 tok/s Pi 5 value remains labeled as an outside reference. It connects Trevor-measured Nova data with a clearly separated published Pi reference without presenting the two methods as a matched hardware contest. What ranQ4 reference ran through llama.cpp using CPU in Jeff Geerling published result mode on Raspberry Pi 5. Cited result. Not Trevor-measured. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | ||||||
Runs and methods
Four exact Q4_K_M GGUFs received the same raw prompt, context, thread count, batch settings, sampler, and 96-token output through Ollama 0.32.1 and native llama.cpp commit 9a286ac.
Qwen3 0.6B and 1.7B ran in three explicit modes with the same native binary and pp128/tg96 workload. Larger full-Vulkan models were retained as failure evidence.
The exact 8.8 GB Bartowski IQ4_XS file from pinned revision 8be84a0 was verified before testing at three and four CPU threads.
Qwen3.5 0.8B, Qwen3.5 2B, and Gemma 4 E2B ran with MTP off and on at every draft depth from one through four. Service logs proved the draft path was active.
Eight additional models were recorded with their actual averaging method, package peak, and artifact limitation where one applied.
The exact BitCPM TQ2_0 file exposed a runtime-support difference. A separate Gemma 3 text-only derivative documented a different conversion problem before the later Vulkan failure.
Six Ollama models received the same deterministic prompt with one cold and two warm requests during the original Kali and hardware test.
Three W8A8 RKLLM models ran on the RK3588S NPU with NPU utilization and memory monitored over SSH.
The values quoted in the Nova article remain ranges or cited results. They are kept separate from Trevor-measured rows and are not converted into invented exact midpoints.
Device record
The X1S has the deepest result set here: matched Ollama and llama.cpp CPU requests, explicit CPU and Vulkan modes, thermal data, MTP tests, compatibility failures, and retained i915 evidence.
View deviceDevice record
Three W8A8 models ran through the RK3588S NPU. The result set includes speed and the observed answer-quality outcome from the shared state-capitals prompt.
View deviceDevice record
The Pi 5 entries are external references, not Trevor’s matched rerun. They stay visible because they explain the original Nova comparison, but they do not belong on a controlled cross-device leaderboard.
View device