True CPU, Mixed Host Operations, and Full Vulkan
Qwen3 0.6B and 1.7B ran in three explicit modes with the same native binary and pp128/tg96 workload. Larger full-Vulkan models were retained as failure evidence.
This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.
Comparison boundary
Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.
Related pages
Useful comparisons from this test
Recorded results
Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.
| Model and artifact | Runtime and mode | Generation | Prompt | Peak | State |
|---|---|---|---|---|---|
| Qwen3 0.6B Q4_K_M | llama.cpp 9a286ac CPU · True CPU | 7.397 tok/s | 10.67 tok/s | 81 °C | Published result |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S. Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window. Recorded setup
| |||||
| Qwen3 0.6B Q4_K_M | llama.cpp 9a286ac CPU + GPU · Mixed host operations | 7.425 tok/s | 34.59 tok/s | 80 °C | Published result |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S. Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved. Recorded setup
| |||||
| Qwen3 0.6B Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | 8.593 tok/s | 37.35 tok/s | 57 °C | Published result |
Open result detailsWhy this model is hereThe smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes. It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. Full Vulkan improved generation 16.2% over true CPU and cut the recorded package peak by 24 °C. Recorded setup
| |||||
| Qwen3 1.7B Q4_K_M | llama.cpp 9a286ac CPU · True CPU | 2.986 tok/s | 3.893 tok/s | 84 °C | Published result |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S. Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window. Recorded setup
| |||||
| Qwen3 1.7B Q4_K_M | llama.cpp 9a286ac CPU + GPU · Mixed host operations | 2.997 tok/s | 12.37 tok/s | 83 °C | Published result |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S. Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved. Recorded setup
| |||||
| Qwen3 1.7B Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | 3.462 tok/s | 12.88 tok/s | 58 °C | Published result |
Open result detailsWhy this model is hereThis model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test. The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. Full Vulkan improved generation 15.9% over true CPU and cut the recorded package peak by 26 °C. Recorded setup
| |||||
| Qwen3 4B Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | n/a | n/a | n/a | Failed run |
Open result detailsWhy this model is hereThe Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan. It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. The run reached an i915 reset timeout. No speed score is published for it. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
| Phi-4 Mini Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | n/a | n/a | n/a | Failed run |
Open result detailsWhy this model is herePhi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang. It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. The run window contained an i915 GPU hang even though llama-bench returned zero. The kernel event overrides the process exit code. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
| Qwen3 8B Q4_K_M | llama.cpp 9a286ac Vulkan · Full Vulkan | n/a | n/a | n/a | Failed run |
Open result detailsWhy this model is hereQwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang. It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path. What ranQ4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. The run window contained an i915 GPU hang even though llama-bench returned zero. No clean performance row is claimed. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
| Gemma 3 4B Text-only derivative | llama.cpp 9a286ac Vulkan · Full Vulkan | n/a | n/a | n/a | Failed run |
Open result detailsWhy this model is hereA separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss. It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix. What ranText-only derivative ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S. Fence and preemption timeouts led to an i915 reset and vk::DeviceLostError. Recorded setupThe full test page contains the shared controls and comparison boundary for this row. | |||||
Supporting evidence
Results and method
These are the few files that support this page directly. The repository holds the wider project history.
Controls
- Native llama.cpp commit 9a286ac
- One pp128/tg96 workload
- True CPU hid Vulkan and disabled operation offload
- Mixed mode used zero GPU layers with host-operation offload
- Full Vulkan used the Intel GPU backend
- Kernel windows and package temperature recorded
What this run established
- $Full Vulkan raised prompt processing by 3.50x and 3.31x for the two clean models.
- $Generation improved 16.2% and 15.9%.
- $The larger full-offload path exposed i915 reset, hang, and device-loss failures.