Original X1S CPU-Only Model Matrix
Six Ollama models received the same deterministic prompt with one cold and two warm requests during the original Kali and hardware test.
This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.
Comparison boundary
These warm Ollama results belong together. They should not be used as the other half of the later llama-bench runtime comparison.
Related pages
Recorded results
Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.
| Model and artifact | Runtime and mode | Generation | Prompt | Peak | State |
|---|---|---|---|---|---|
| Qwen3 0.6B Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 6.788 tok/s | n/a | 74 °C | Published result |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Fastest row, least complete response on the single prompt. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
| Qwen3 1.7B Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 3.129 tok/s | n/a | 77 °C | Published result |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Best interactive starting tier tested. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
| Qwen3 4B Instruct Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 1.824 tok/s | n/a | 77 °C | Published result |
Open result detailsWhy this model is hereThe Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan. It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Patient local batch use. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
| Phi-4 Mini Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 1.991 tok/s | n/a | 77 °C | Published result |
Open result detailsWhy this model is herePhi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang. It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Patient local batch use. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
| Gemma 3 4B Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 1.995 tok/s | n/a | 77 °C | Published result |
Open result detailsWhy this model is hereA separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss. It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Patient local batch use. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
| Qwen3 8B Q4_K_M | Ollama 0.32.1 CPU · Mean warm generation | 0.924 tok/s | n/a | 80 °C | Published result |
Open result detailsWhy this model is hereQwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang. It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S. Fit comfortably with sub-1 tok/s generation. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
Supporting evidence
Results and method
These are the few files that support this page directly. The repository holds the wider project history.
Controls
- Ollama 0.32.1 CPU backend
- 4,096-token context
- Temperature 0 and fixed seed
- One cold and two warm requests
- Normal 85 °C thermal guard
What this run established
- $Qwen3 0.6B was fastest but least complete on the single prompt.
- $Qwen3 1.7B was the best interactive starting tier tested.
- $Qwen3 8B fit comfortably, but generation stayed below 1 tok/s.