/hub/models/qwen35-2b

Qwen3.5 · 2B

Qwen3.5 2B

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

9
result records
9
Trevor-measured rows
2
test groups
1
devices represented

Why Qwen3.5 2B is in The Hub

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

The rows on this page use Q8_0 through Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 5.154 tok/s through Ollama 0.32.1 using CPU in Depth 2, MTP off mode. Paired MTP-off baseline for draft depth 2.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 1, MTP off
4.891 tok/sn/a78 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
3.778 tok/sn/a77 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 22.76% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 2, MTP off
5.154 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
2.853 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 44.65% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 3, MTP off
4.853 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
2.599 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 46.45% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 4, MTP off
5.121 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
1.522 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 70.28% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Two-run warm average
4.89 tok/sn/a75 °CPublished resultAdditional models
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

A clean middle-tier CPU result from the normal non-MTP Ollama path.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C

Full tests containing Qwen3.5 2B

2026-08-31

Ollama MTP Depth Sweep on the N5095

Each on/off pair belongs to its own requested draft depth. Do not compare the baseline from one pair as if it were the same request block as another.

Open full test

2026-08-31

Additional Model Results on the X1S

The rows use different repetition counts and artifact states. They are useful operating points, not one controlled model-quality contest.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"