/hub/runs/x1s-original-cpu-matrix
Published result2026-08-23Youyeetoo X1S

Original X1S CPU-Only Model Matrix

Six Ollama models received the same deterministic prompt with one cold and two warm requests during the original Kali and hardware test.

This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.

Comparison boundary

These warm Ollama results belong together. They should not be used as the other half of the later llama-bench runtime comparison.

Recorded results

Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.

Model and artifactRuntime and modeGenerationPromptPeakState
Qwen3 0.6B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
6.788 tok/sn/a74 °CPublished result
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Fastest row, least complete response on the single prompt.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 1.7B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
3.129 tok/sn/a77 °CPublished result
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Best interactive starting tier tested.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 4B Instruct
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.824 tok/sn/a77 °CPublished result
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Phi-4 Mini
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.991 tok/sn/a77 °CPublished result
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Gemma 3 4B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.995 tok/sn/a77 °CPublished result
Open result details

Why this model is here

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 8B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
0.924 tok/sn/a80 °CPublished result
Open result details

Why this model is here

Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.

It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Fit comfortably with sub-1 tok/s generation.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Supporting evidence

Results and method

These are the few files that support this page directly. The repository holds the wider project history.

Controls

  • Ollama 0.32.1 CPU backend
  • 4,096-token context
  • Temperature 0 and fixed seed
  • One cold and two warm requests
  • Normal 85 °C thermal guard

What this run established

  • $Qwen3 0.6B was fastest but least complete on the single prompt.
  • $Qwen3 1.7B was the best interactive starting tier tested.
  • $Qwen3 8B fit comfortably, but generation stayed below 1 tok/s.

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"