/hub/models/qwen3-4b-instruct

Qwen3 · 4B

Qwen3 4B Instruct

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

4
result records
3
Trevor-measured rows
3
test groups
1
devices represented

Why Qwen3 4B Instruct is in The Hub

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

The rows on this page use Q4_K_M through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 2.002 tok/s through Ollama 0.32.1 using CPU in Matched server request mode. Ollama reported 34.89% higher internal generation throughput than the paired llama.cpp row.

1 retained row is a failure: The run reached an i915 reset timeout. No speed score is published for it.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Qwen3 4B Instruct
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
1.484 tok/s1.899 tok/s81 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.000599
Thermal ceiling
85 °C

Artifact SHA-25685e4a5b7b8ef0e48af0e8658f5aaab9c2324c76c1641493f4d1e25fce54b18b9

Qwen3 4B Instruct
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
2.002 tok/s2.789 tok/s81 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 34.89% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.033289
Thermal ceiling
85 °C

Artifact SHA-25685e4a5b7b8ef0e48af0e8658f5aaab9c2324c76c1641493f4d1e25fce54b18b9

Qwen3 4B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed runCPU and Vulkan modes
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run reached an i915 reset timeout. No speed score is published for it.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 4B Instruct
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.824 tok/sn/a77 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing Qwen3 4B Instruct

2026-08-31

Matched Ollama and llama.cpp CPU Test

Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.

Open full test

2026-08-31

True CPU, Mixed Host Operations, and Full Vulkan

Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.

Open full test

2026-08-23

Original X1S CPU-Only Model Matrix

These warm Ollama results belong together. They should not be used as the other half of the later llama-bench runtime comparison.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"