/hub/runs/x1s-mtp-depth-sweep
Published result2026-08-31Youyeetoo X1S

Ollama MTP Depth Sweep on the N5095

Qwen3.5 0.8B, Qwen3.5 2B, and Gemma 4 E2B ran with MTP off and on at every draft depth from one through four. Service logs proved the draft path was active.

This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.

Comparison boundary

Each on/off pair belongs to its own requested draft depth. Do not compare the baseline from one pair as if it were the same request block as another.

Useful comparisons from this test

Recorded results

Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.

Model and artifactRuntime and modeGenerationPromptPeakState
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 1, MTP off
11.34 tok/sn/a74 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
6.043 tok/sn/a75 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 46.73% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 2, MTP off
11.49 tok/sn/a77 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
4.019 tok/sn/a79 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 65.03% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 3, MTP off
11.21 tok/sn/a76 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
3.133 tok/sn/a77 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 72.06% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 4, MTP off
11.16 tok/sn/a78 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
2.624 tok/sn/a79 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 76.48% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 1, MTP off
4.891 tok/sn/a78 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
3.778 tok/sn/a77 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 22.76% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 2, MTP off
5.154 tok/sn/a79 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
2.853 tok/sn/a79 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 44.65% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 3, MTP off
4.853 tok/sn/a80 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
2.599 tok/sn/a80 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 46.45% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 4, MTP off
5.121 tok/sn/a80 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
1.522 tok/sn/a79 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 70.28% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 1, MTP off
3.269 tok/sn/a80 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
2.266 tok/sn/a84 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 30.68% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 2, MTP off
3.359 tok/sn/a80 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
1.96 tok/sn/a84 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 41.63% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 3, MTP off
3.352 tok/sn/a82 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
1.474 tok/sn/a86 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 56.03% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 4, MTP off
3.397 tok/sn/a81 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
1.333 tok/sn/a84 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 60.75% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C

Supporting evidence

Results and method

These are the few files that support this page directly. The repository holds the wider project history.

Controls

  • Ollama 0.32.1
  • Draft depths 1 through 4 were tested
  • Service journal retained draft-MTP activation
  • Paired output token counts and hashes matched
  • Thermal and kernel guards remained enabled

What this run established

  • $Every tested MTP depth was slower than MTP off.
  • $Depth one reduced generation by 22.8% to 46.7%.
  • $On this four-core N5095, draft overhead cost more than it saved.

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"