/hub/models/qwen3-8b

Qwen3 · 8B

Qwen3 8B

Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.

2
result records
1
Trevor-measured rows
2
test groups
1
devices represented

Why Qwen3 8B is in The Hub

Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.

It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path.

The rows on this page use Q4_K_M through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 0.924 tok/s through Ollama 0.32.1 using CPU in Mean warm generation mode. Fit comfortably with sub-1 tok/s generation.

1 retained row is a failure: The run window contained an i915 GPU hang even though llama-bench returned zero. No clean performance row is claimed.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Qwen3 8B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed runCPU and Vulkan modes
Open result details

Why this model is here

Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.

It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run window contained an i915 GPU hang even though llama-bench returned zero. No clean performance row is claimed.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 8B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
0.924 tok/sn/a80 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.

It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Fit comfortably with sub-1 tok/s generation.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing Qwen3 8B

2026-08-31

True CPU, Mixed Host Operations, and Full Vulkan

Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.

Open full test

2026-08-23

Original X1S CPU-Only Model Matrix

These warm Ollama results belong together. They should not be used as the other half of the later llama-bench runtime comparison.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"