/hub/runs/x1s-cpu-vulkan-modes
Published result2026-08-31Youyeetoo X1S

True CPU, Mixed Host Operations, and Full Vulkan

Qwen3 0.6B and 1.7B ran in three explicit modes with the same native binary and pp128/tg96 workload. Larger full-Vulkan models were retained as failure evidence.

This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.

Comparison boundary

Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.

Useful comparisons from this test

Recorded results

Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.

Model and artifactRuntime and modeGenerationPromptPeakState
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU · True CPU
7.397 tok/s10.67 tok/s81 °CPublished result
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S.

Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU + GPU · Mixed host operations
7.425 tok/s34.59 tok/s80 °CPublished result
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S.

Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
8.593 tok/s37.35 tok/s57 °CPublished result
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Full Vulkan improved generation 16.2% over true CPU and cut the recorded package peak by 24 °C.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU · True CPU
2.986 tok/s3.893 tok/s84 °CPublished result
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S.

Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU + GPU · Mixed host operations
2.997 tok/s12.37 tok/s83 °CPublished result
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S.

Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
3.462 tok/s12.88 tok/s58 °CPublished result
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Full Vulkan improved generation 15.9% over true CPU and cut the recorded package peak by 26 °C.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 4B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed run
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run reached an i915 reset timeout. No speed score is published for it.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Phi-4 Mini
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed run
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run window contained an i915 GPU hang even though llama-bench returned zero. The kernel event overrides the process exit code.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 8B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed run
Open result details

Why this model is here

Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.

It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run window contained an i915 GPU hang even though llama-bench returned zero. No clean performance row is claimed.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Gemma 3 4B
Text-only derivative
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed run
Open result details

Why this model is here

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

What ran

Text-only derivative ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Fence and preemption timeouts led to an i915 reset and vk::DeviceLostError.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Supporting evidence

Results and method

These are the few files that support this page directly. The repository holds the wider project history.

Controls

  • Native llama.cpp commit 9a286ac
  • One pp128/tg96 workload
  • True CPU hid Vulkan and disabled operation offload
  • Mixed mode used zero GPU layers with host-operation offload
  • Full Vulkan used the Intel GPU backend
  • Kernel windows and package temperature recorded

What this run established

  • $Full Vulkan raised prompt processing by 3.50x and 3.31x for the two clean models.
  • $Generation improved 16.2% and 15.9%.
  • $The larger full-offload path exposed i915 reset, hang, and device-loss failures.

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"