/hub/models/qwen35-08b

Qwen3.5 · 0.8B

Qwen3.5 0.8B

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

9
result records
9
Trevor-measured rows
2
test groups
1
devices represented

Why Qwen3.5 0.8B is in The Hub

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

The rows on this page use Q8_0 through Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 11.49 tok/s through Ollama 0.32.1 using CPU in Depth 2, MTP off mode. Paired MTP-off baseline for draft depth 2.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 1, MTP off
11.34 tok/sn/a74 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
6.043 tok/sn/a75 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 46.73% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 2, MTP off
11.49 tok/sn/a77 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
4.019 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 65.03% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 3, MTP off
11.21 tok/sn/a76 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
3.133 tok/sn/a77 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 72.06% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 4, MTP off
11.16 tok/sn/a78 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
2.624 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 76.48% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Two-run warm average
10.44 tok/sn/a70 °CPublished resultAdditional models
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

The fastest normal text-generation result in the X1S work. Throughput does not establish answer quality.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C

Full tests containing Qwen3.5 0.8B

2026-08-31

Ollama MTP Depth Sweep on the N5095

Each on/off pair belongs to its own requested draft depth. Do not compare the baseline from one pair as if it were the same request block as another.

Open full test

2026-08-31

Additional Model Results on the X1S

The rows use different repetition counts and artifact states. They are useful operating points, not one controlled model-quality contest.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"