/hub/models/qwen3-06b

Qwen3 · 0.6B

Qwen3 0.6B

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

6
result records
6
Trevor-measured rows
3
test groups
1
devices represented

Why Qwen3 0.6B is in The Hub

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

The rows on this page use Q4_K_M through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 8.593 tok/s through llama.cpp 9a286ac using Vulkan in Full Vulkan mode. Full Vulkan improved generation 16.2% over true CPU and cut the recorded package peak by 24 °C.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
6.725 tok/s10.48 tok/s81 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.014215
Thermal ceiling
85 °C

Artifact SHA-2567f4030143c1c477224c5434f8272c662a8b042079a0a584f0a27a1684fe2e1fa

Qwen3 0.6B
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
7.809 tok/s13.24 tok/s81 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 16.11% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.094756
Thermal ceiling
85 °C

Artifact SHA-2567f4030143c1c477224c5434f8272c662a8b042079a0a584f0a27a1684fe2e1fa

Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU · True CPU
7.397 tok/s10.67 tok/s81 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S.

Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU + GPU · Mixed host operations
7.425 tok/s34.59 tok/s80 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S.

Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
8.593 tok/s37.35 tok/s57 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Full Vulkan improved generation 16.2% over true CPU and cut the recorded package peak by 24 °C.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 0.6B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
6.788 tok/sn/a74 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Fastest row, least complete response on the single prompt.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing Qwen3 0.6B

2026-08-31

Matched Ollama and llama.cpp CPU Test

Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.

Open full test

2026-08-31

True CPU, Mixed Host Operations, and Full Vulkan

Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.

Open full test

2026-08-23

Original X1S CPU-Only Model Matrix

These warm Ollama results belong together. They should not be used as the other half of the later llama-bench runtime comparison.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"