/hub/runs/x1s-matched-ollama-llamacpp
Published result2026-08-31Youyeetoo X1S

Matched Ollama and llama.cpp CPU Test

Four exact Q4_K_M GGUFs received the same raw prompt, context, thread count, batch settings, sampler, and 96-token output through Ollama 0.32.1 and native llama.cpp commit 9a286ac.

This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.

Comparison boundary

Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.

Useful comparisons from this test

Recorded results

Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.

Model and artifactRuntime and modeGenerationPromptPeakState
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
6.725 tok/s10.48 tok/s81 °CPublished result
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.014215
Thermal ceiling
85 °C

Artifact SHA-2567f4030143c1c477224c5434f8272c662a8b042079a0a584f0a27a1684fe2e1fa

Qwen3 0.6B
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
7.809 tok/s13.24 tok/s81 °CPublished result
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 16.11% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.094756
Thermal ceiling
85 °C

Artifact SHA-2567f4030143c1c477224c5434f8272c662a8b042079a0a584f0a27a1684fe2e1fa

Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
2.852 tok/s3.849 tok/s83 °CPublished result
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.002042
Thermal ceiling
85 °C

Artifact SHA-2563d0b790534fe4b79525fc3692950408dca41171676ed7e21db57af5c65ef6ab6

Qwen3 1.7B
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
3.321 tok/s5.075 tok/s83 °CPublished result
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 16.46% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.012057
Thermal ceiling
85 °C

Artifact SHA-2563d0b790534fe4b79525fc3692950408dca41171676ed7e21db57af5c65ef6ab6

Qwen3 4B Instruct
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
1.484 tok/s1.899 tok/s81 °CPublished result
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.000599
Thermal ceiling
85 °C

Artifact SHA-25685e4a5b7b8ef0e48af0e8658f5aaab9c2324c76c1641493f4d1e25fce54b18b9

Qwen3 4B Instruct
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
2.002 tok/s2.789 tok/s81 °CPublished result
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 34.89% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.033289
Thermal ceiling
85 °C

Artifact SHA-25685e4a5b7b8ef0e48af0e8658f5aaab9c2324c76c1641493f4d1e25fce54b18b9

Phi-4 Mini
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
1.57 tok/s2.115 tok/s82 °CPublished result
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
70 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.000531
Thermal ceiling
85 °C

Artifact SHA-2563c168af1dea0a414299c7d9077e100ac763370e5a98b3c53801a958a47f0a5db

Phi-4 Mini
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
2.072 tok/s3.131 tok/s82 °CPublished result
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 32.00% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
70 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.009724
Thermal ceiling
85 °C

Artifact SHA-2563c168af1dea0a414299c7d9077e100ac763370e5a98b3c53801a958a47f0a5db

Supporting evidence

Results and method

These are the few files that support this page directly. The repository holds the wider project history.

Controls

  • Same raw prompt and exact content-addressed GGUF
  • 4,096-token context
  • Four threads
  • Batch and microbatch 512
  • Temperature 0, seed 42, top-k 1
  • One warmup and three measured requests per model
  • Exactly 96 generated tokens

What this run established

  • $All 24 measured requests completed with clean kernel windows.
  • $Ollama reported 16.1% to 34.9% higher internal generation throughput in this configuration.
  • $The two runtimes produced different greedy text, which was retained as part of the record.

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"