Ollama MTP Depth Sweep on the N5095
Qwen3.5 0.8B, Qwen3.5 2B, and Gemma 4 E2B ran with MTP off and on at every draft depth from one through four. Service logs proved the draft path was active.
This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.
Comparison boundary
Each on/off pair belongs to its own requested draft depth. Do not compare the baseline from one pair as if it were the same request block as another.
Related pages
Useful comparisons from this test
Recorded results
Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.
| Model and artifact | Runtime and mode | Generation | Prompt | Peak | State |
|---|---|---|---|---|---|
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Depth 1, MTP off | 11.34 tok/s | n/a | 74 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 1. Recorded setup
| |||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 1, MTP on | 6.043 tok/s | n/a | 75 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S. MTP was active and ran 46.73% slower than its paired baseline. Recorded setup
| |||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Depth 2, MTP off | 11.49 tok/s | n/a | 77 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 2. Recorded setup
| |||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 2, MTP on | 4.019 tok/s | n/a | 79 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S. MTP was active and ran 65.03% slower than its paired baseline. Recorded setup
| |||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Depth 3, MTP off | 11.21 tok/s | n/a | 76 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 3. Recorded setup
| |||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 3, MTP on | 3.133 tok/s | n/a | 77 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S. MTP was active and ran 72.06% slower than its paired baseline. Recorded setup
| |||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU · Depth 4, MTP off | 11.16 tok/s | n/a | 78 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 4. Recorded setup
| |||||
| Qwen3.5 0.8B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 4, MTP on | 2.624 tok/s | n/a | 79 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair. It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S. MTP was active and ran 76.48% slower than its paired baseline. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Depth 1, MTP off | 4.891 tok/s | n/a | 78 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 1. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 1, MTP on | 3.778 tok/s | n/a | 77 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S. MTP was active and ran 22.76% slower than its paired baseline. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Depth 2, MTP off | 5.154 tok/s | n/a | 79 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 2. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 2, MTP on | 2.853 tok/s | n/a | 79 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S. MTP was active and ran 44.65% slower than its paired baseline. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Depth 3, MTP off | 4.853 tok/s | n/a | 80 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 3. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 3, MTP on | 2.599 tok/s | n/a | 80 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S. MTP was active and ran 46.45% slower than its paired baseline. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU · Depth 4, MTP off | 5.121 tok/s | n/a | 80 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 4. Recorded setup
| |||||
| Qwen3.5 2B Q8_0 | Ollama 0.32.1 CPU + draft MTP · Depth 4, MTP on | 1.522 tok/s | n/a | 79 °C | Published result |
Open result detailsWhy this model is hereQwen3.5 2B completed the normal Ollama run and MTP depths one through four. It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size. What ranQ8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S. MTP was active and ran 70.28% slower than its paired baseline. Recorded setup
| |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Depth 1, MTP off | 3.269 tok/s | n/a | 80 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 1. Recorded setup
| |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU + draft MTP · Depth 1, MTP on | 2.266 tok/s | n/a | 84 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S. MTP was active and ran 30.68% slower than its paired baseline. Recorded setup
| |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Depth 2, MTP off | 3.359 tok/s | n/a | 80 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 2. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard. Recorded setup
| |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU + draft MTP · Depth 2, MTP on | 1.96 tok/s | n/a | 84 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S. MTP was active and ran 41.63% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard. Recorded setup
| |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Depth 3, MTP off | 3.352 tok/s | n/a | 82 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 3. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard. Recorded setup
| |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU + draft MTP · Depth 3, MTP on | 1.474 tok/s | n/a | 86 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S. MTP was active and ran 56.03% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard. Recorded setup
| |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU · Depth 4, MTP off | 3.397 tok/s | n/a | 81 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S. Paired MTP-off baseline for draft depth 4. Recorded setup
| |||||
| Gemma 4 E2B Q4_K_M | Ollama 0.32.1 CPU + draft MTP · Depth 4, MTP on | 1.333 tok/s | n/a | 84 °C | Published result |
Open result detailsWhy this model is hereGemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model. It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active. What ranQ4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S. MTP was active and ran 60.75% slower than its paired baseline. Recorded setup
| |||||
Supporting evidence
Results and method
These are the few files that support this page directly. The repository holds the wider project history.
Controls
- Ollama 0.32.1
- Draft depths 1 through 4 were tested
- Service journal retained draft-MTP activation
- Paired output token counts and hashes matched
- Thermal and kernel guards remained enabled
What this run established
- $Every tested MTP depth was slower than MTP off.
- $Depth one reduced generation by 22.8% to 46.7%.
- $On this four-core N5095, draft overhead cost more than it saved.