Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU · Depth 1, MTP off
4.891 tok/s n/a 78 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.
Paired MTP-off baseline for draft depth 1.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
3.778 tok/s n/a 77 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.
MTP was active and ran 22.76% slower than its paired baseline.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU · Depth 2, MTP off
5.154 tok/s n/a 79 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.
Paired MTP-off baseline for draft depth 2.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
2.853 tok/s n/a 79 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.
MTP was active and ran 44.65% slower than its paired baseline.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU · Depth 3, MTP off
4.853 tok/s n/a 80 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.
Paired MTP-off baseline for draft depth 3.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
2.599 tok/s n/a 80 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.
MTP was active and ran 46.45% slower than its paired baseline.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 1
Thermal ceiling 85 °C Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU · Depth 4, MTP off
5.121 tok/s n/a 80 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.
Paired MTP-off baseline for draft depth 4.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 5
Thermal ceiling 85 °C Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
1.522 tok/s n/a 79 °C Published result MTP depth sweep Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.
MTP was active and ran 70.28% slower than its paired baseline.
Recorded setup
Threads 4
Output 96 tokens
Batch 512 / 512 microbatch
Measured runs 5
Thermal ceiling 85 °C Qwen3.5 2B Q8_0
Ollama 0.32.1
CPU · Two-run warm average
4.89 tok/s n/a 75 °C Published result Additional models Open result detailsWhy this model is here Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.
It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.
What ran Q8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.
A clean middle-tier CPU result from the normal non-MTP Ollama path.
Recorded setup
Context 4,096 tokens
Measured runs 2
Thermal ceiling 85 °C