Additional Model Results on the X1S
Eight additional models were recorded with their actual averaging method, package peak, and artifact limitation where one applied.
This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.
Comparison boundary
The rows use different repetition counts and artifact states. They are useful operating points, not one controlled model-quality contest.
Related pages
Recorded results
Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.
| Model and artifact | Runtime and mode | Generation | Prompt | Peak | State |
|---|---|---|---|---|---|
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Two-run warm average | 10.44 tok/s | n/a | 70 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S. The fastest normal text-generation result in the X1S work. Throughput does not establish answer quality. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Two-run warm average | 4.89 tok/s | n/a | 75 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S. A clean middle-tier CPU result from the normal non-MTP Ollama path. Recorded setup
| |||||
| Qwen3.5 9B Q4_K_M | Ollama 0.32.1 CPU · Five-prompt average | 1.11 tok/s | n/a | 82 °C | Published result |
Open result detailsWhy this model is hereThe official Q4_K_M Ollama artifact fit in 16 GB and averaged 1.110 tok/s across five prompts. This was a fit-and-completion test for a much larger model on the X1S, not a claim that 1.110 tok/s is comfortable interactive speed. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Five-prompt average mode on Youyeetoo X1S. The 9B artifact fit in 16 GB and completed five prompts, but generation requires patience. Recorded setup
Artifact SHA-256 | |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Two-run warm average | 3.38 tok/s | n/a | 77 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S. The normal non-MTP Ollama result. The separate MTP sweep is recorded in its own test group. Recorded setup
| |||||
Gemma 4 E4B Q4_K_M | Ollama 0.32.1 CPU · Two-run warm average | 1.78 tok/s | n/a | 78 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E4B completed the original warm Ollama test at 1.78 tok/s. The row is useful as a recorded X1S operating point, but it does not have enough separate material for an indexable model page yet. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S. A clean normal Ollama row retained as an operating point, not a quality ranking. | |||||
Granite 4 Tiny-H Original Ollama tag | Ollama 0.32.1 CPU · Single original observation | 5.43 tok/s | n/a | 79 °C | Partial result |
Open result detailsWhy this model is hereThe X1S recorded one 5.43 tok/s Ollama observation. One observation is retained for completeness but is not promoted into a deeper standalone page. What ranOriginal Ollama tag ran through Ollama 0.32.1 using CPU in Single original observation mode on Youyeetoo X1S. One retained observation, not a repeated average. | |||||
LFM2.5 Mutable latest tag | Ollama 0.32.1 CPU · Single recorded tag state | 4.95 tok/s | n/a | 80 °C | Partial result |
Open result detailsWhy this model is hereThe X1S recorded one 4.95 tok/s result from a mutable latest tag. The tag was not content-pinned, so the row stays visible as a partial result without becoming its own search page. What ranMutable latest tag ran through Ollama 0.32.1 using CPU in Single recorded tag state mode on Youyeetoo X1S. The mutable latest tag was not content-pinned, so this remains a partial result. | |||||
Gemma 3n E2B gemma3n:e2b | Ollama 0.32.1 CPU · Single original observation | 3.24 tok/s | n/a | 80 °C | Partial result |
Open result detailsWhy this model is hereThe X1S recorded one 3.24 tok/s Ollama result. The row is kept separate from Gemma 4 E2B and does not yet have enough unique material for its own page. What rangemma3n:e2b ran through Ollama 0.32.1 using CPU in Single original observation mode on Youyeetoo X1S. One retained observation. This is Gemma 3n E2B, not Gemma 4 E2B. | |||||
Supporting evidence
Results and method
These are the few files that support this page directly. The repository holds the wider project history.
Controls
- Ollama CPU workflow
- Normal 85 °C thermal abort unless disclosed elsewhere
- Kernel and request timeout checks enabled
- Method recorded per row
What this run established
- $Qwen3.5 0.8B was the fastest text-generation result in the project.
- $Qwen3.5 9B fit and completed five prompts at 1.110 tok/s.
- $LFM2.5 used a mutable latest tag and is not content-pinned.