/hub/models/qwen3-17b

Qwen3 · 1.7B

Qwen3 1.7B

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

6
result records
6
Trevor-measured rows
3
test groups
1
devices represented

Why Qwen3 1.7B is in The Hub

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

The rows on this page use Q4_K_M through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 3.462 tok/s through llama.cpp 9a286ac using Vulkan in Full Vulkan mode. Full Vulkan improved generation 15.9% over true CPU and cut the recorded package peak by 26 °C.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
2.852 tok/s3.849 tok/s83 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.002042
Thermal ceiling
85 °C

Artifact SHA-2563d0b790534fe4b79525fc3692950408dca41171676ed7e21db57af5c65ef6ab6

Qwen3 1.7B
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
3.321 tok/s5.075 tok/s83 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 16.46% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.012057
Thermal ceiling
85 °C

Artifact SHA-2563d0b790534fe4b79525fc3692950408dca41171676ed7e21db57af5c65ef6ab6

Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU · True CPU
2.986 tok/s3.893 tok/s84 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S.

Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU + GPU · Mixed host operations
2.997 tok/s12.37 tok/s83 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S.

Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
3.462 tok/s12.88 tok/s58 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Full Vulkan improved generation 15.9% over true CPU and cut the recorded package peak by 26 °C.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
3.129 tok/sn/a77 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Best interactive starting tier tested.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing Qwen3 1.7B

2026-08-31

Matched Ollama and llama.cpp CPU Test

Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.

Open full test

2026-08-31

True CPU, Mixed Host Operations, and Full Vulkan

Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.

Open full test

2026-08-23

Original X1S CPU-Only Model Matrix

These warm Ollama results belong together. They should not be used as the other half of the later llama-bench runtime comparison.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"