Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU · Depth 1, MTP off
11.34 tok/s n/a 74 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.
Paired MTP-off baseline for draft depth 1.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
6.043 tok/s n/a 75 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.
MTP was active and ran 46.73% slower than its paired baseline.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU · Depth 2, MTP off
11.49 tok/s n/a 77 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.
Paired MTP-off baseline for draft depth 2.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
4.019 tok/s n/a 79 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.
MTP was active and ran 65.03% slower than its paired baseline.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU · Depth 3, MTP off
11.21 tok/s n/a 76 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.
Paired MTP-off baseline for draft depth 3.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
3.133 tok/s n/a 77 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.
MTP was active and ran 72.06% slower than its paired baseline.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU · Depth 4, MTP off
11.16 tok/s n/a 78 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.
Paired MTP-off baseline for draft depth 4.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 5
Thermal ceiling 85 °C Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
2.624 tok/s n/a 79 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.
MTP was active and ran 76.48% slower than its paired baseline.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 5
Thermal ceiling 85 °C Qwen3.5 0.8B Q8_0
Ollama 0.32.1
CPU · Two-run warm average
10.44 tok/s n/a 70 °C Published result Additional models Open result detailsWhy this model is here Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.
It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.
The fastest normal text-generation result in the X1S work. Throughput does not establish answer quality.
Recorded setup
Context 4,096 tokens
Measured runs 2
Thermal ceiling 85 °C