/hub/devices/youyeetoo-x1s

Youyeetoo X1S

The X1S has the deepest result set here: matched Ollama and llama.cpp CPU requests, explicit CPU and Vulkan modes, thermal data, MTP tests, compatibility failures, and retained i915 evidence.

61
records
53
published measurements
5
retained failures
7
test groups

System in this record

Processor
Intel Celeron N5095, 4 cores / 4 threads
Memory
16 GB installed, about 15 GiB visible
Storage
128 GB M.2 2280 NVMe supplied by Youyeetoo
Cooling
Supplied heatsink and active fan, open bench
Software
Kali Linux 2025.4 amd64, kernel 6.16.8+kali-amd64

Measured on Trevor’s X1S review unit.

Full tests

Published result2026-08-31

Matched Ollama and llama.cpp CPU Test

Four exact Q4_K_M GGUFs received the same raw prompt, context, thread count, batch settings, sampler, and 96-token output through Ollama 0.32.1 and native llama.cpp commit 9a286ac.

Open full test
Published result2026-08-31

True CPU, Mixed Host Operations, and Full Vulkan

Qwen3 0.6B and 1.7B ran in three explicit modes with the same native binary and pp128/tg96 workload. Larger full-Vulkan models were retained as failure evidence.

Open full test
Published result2026-08-31

Ling-mini-2.0 IQ4_XS CPU Thread Test

The exact 8.8 GB Bartowski IQ4_XS file from pinned revision 8be84a0 was verified before testing at three and four CPU threads.

Open full test
Published result2026-08-31

Ollama MTP Depth Sweep on the N5095

Qwen3.5 0.8B, Qwen3.5 2B, and Gemma 4 E2B ran with MTP off and on at every draft depth from one through four. Service logs proved the draft path was active.

Open full test
Published result2026-08-31

Additional Model Results on the X1S

Eight additional models were recorded with their actual averaging method, package peak, and artifact limitation where one applied.

Open full test
Published result2026-08-31

BitCPM and Gemma Runtime Compatibility

The exact BitCPM TQ2_0 file exposed a runtime-support difference. A separate Gemma 3 text-only derivative documented a different conversion problem before the later Vulkan failure.

Open full test
Published result2026-08-23

Original X1S CPU-Only Model Matrix

Six Ollama models received the same deterministic prompt with one cold and two warm requests during the original Kali and hardware test.

Open full test

Every retained row

Results on Youyeetoo X1S

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
6.725 tok/s10.48 tok/s81 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.014215
Thermal ceiling
85 °C

Artifact SHA-2567f4030143c1c477224c5434f8272c662a8b042079a0a584f0a27a1684fe2e1fa

Qwen3 0.6B
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
7.809 tok/s13.24 tok/s81 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 16.11% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.094756
Thermal ceiling
85 °C

Artifact SHA-2567f4030143c1c477224c5434f8272c662a8b042079a0a584f0a27a1684fe2e1fa

Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
2.852 tok/s3.849 tok/s83 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.002042
Thermal ceiling
85 °C

Artifact SHA-2563d0b790534fe4b79525fc3692950408dca41171676ed7e21db57af5c65ef6ab6

Qwen3 1.7B
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
3.321 tok/s5.075 tok/s83 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 16.46% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.012057
Thermal ceiling
85 °C

Artifact SHA-2563d0b790534fe4b79525fc3692950408dca41171676ed7e21db57af5c65ef6ab6

Qwen3 4B Instruct
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
1.484 tok/s1.899 tok/s81 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.000599
Thermal ceiling
85 °C

Artifact SHA-25685e4a5b7b8ef0e48af0e8658f5aaab9c2324c76c1641493f4d1e25fce54b18b9

Qwen3 4B Instruct
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
2.002 tok/s2.789 tok/s81 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 34.89% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
71 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.033289
Thermal ceiling
85 °C

Artifact SHA-25685e4a5b7b8ef0e48af0e8658f5aaab9c2324c76c1641493f4d1e25fce54b18b9

Phi-4 Mini
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
1.57 tok/s2.115 tok/s82 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
70 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.000531
Thermal ceiling
85 °C

Artifact SHA-2563c168af1dea0a414299c7d9077e100ac763370e5a98b3c53801a958a47f0a5db

Phi-4 Mini
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
2.072 tok/s3.131 tok/s82 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 32.00% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
70 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.009724
Thermal ceiling
85 °C

Artifact SHA-2563c168af1dea0a414299c7d9077e100ac763370e5a98b3c53801a958a47f0a5db

Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU · True CPU
7.397 tok/s10.67 tok/s81 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S.

Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
CPU + GPU · Mixed host operations
7.425 tok/s34.59 tok/s80 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S.

Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 0.6B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
8.593 tok/s37.35 tok/s57 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Full Vulkan improved generation 16.2% over true CPU and cut the recorded package peak by 24 °C.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU · True CPU
2.986 tok/s3.893 tok/s84 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in True CPU mode on Youyeetoo X1S.

Vulkan was hidden and host-operation offload was disabled. The five-repetition run had a clean kernel window.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
CPU + GPU · Mixed host operations
2.997 tok/s12.37 tok/s83 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU + GPU in Mixed host operations mode on Youyeetoo X1S.

Zero model layers still allowed host operations on the Intel GPU. Prompt processing rose sharply while generation barely moved.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 1.7B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
3.462 tok/s12.88 tok/s58 °CPublished resultCPU and Vulkan modes
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Full Vulkan improved generation 15.9% over true CPU and cut the recorded package peak by 26 °C.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3 4B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed runCPU and Vulkan modes
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run reached an i915 reset timeout. No speed score is published for it.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Phi-4 Mini
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed runCPU and Vulkan modes
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run window contained an i915 GPU hang even though llama-bench returned zero. The kernel event overrides the process exit code.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 8B
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed runCPU and Vulkan modes
Open result details

Why this model is here

Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.

It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run window contained an i915 GPU hang even though llama-bench returned zero. No clean performance row is claimed.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Gemma 3 4B
Text-only derivative
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed runCPU and Vulkan modes
Open result details

Why this model is here

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

What ran

Text-only derivative ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Fence and preemption timeouts led to an i915 reset and vk::DeviceLostError.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Ling-mini-2.0
IQ4_XS, pinned revision 8be84a0
llama.cpp 9a286ac
CPU · 3 threads, 85 °C guard
3.091 tok/s4.375 tok/s74 °CPublished resultLing-mini thread test
Open result details

Why this model is here

The exact pinned IQ4_XS file completed true-CPU tests at three and four threads.

The thread test shows a real speed and temperature tradeoff on the N5095 using one verified 8.8 GB artifact rather than a mutable model tag.

What ran

IQ4_XS, pinned revision 8be84a0 ran through llama.cpp 9a286ac using CPU in 3 threads, 85 °C guard mode on Youyeetoo X1S.

The clean three-thread row stayed well below the normal thermal guard.

Recorded setup

Threads
3
Prompt
128 tokens
Output
96 tokens
Measured runs
5
Generation stddev
0.001523
Thermal ceiling
85 °C
Source revision
8be84a0f4727

Artifact SHA-256a72d86d4cb4fedd940e34c08d008bb5cda42db80ce5c6bc5f9494e854a3d742d

Ling-mini-2.0
IQ4_XS, pinned revision 8be84a0
llama.cpp 9a286ac
CPU · 4 threads, 90 °C guard
4.083 tok/s5.817 tok/s85 °CPublished resultLing-mini thread test
Open result details

Why this model is here

The exact pinned IQ4_XS file completed true-CPU tests at three and four threads.

The thread test shows a real speed and temperature tradeoff on the N5095 using one verified 8.8 GB artifact rather than a mutable model tag.

What ran

IQ4_XS, pinned revision 8be84a0 ran through llama.cpp 9a286ac using CPU in 4 threads, 90 °C guard mode on Youyeetoo X1S.

The selected four-thread row improved generation 32.1% and completed under the disclosed 90 °C ceiling.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Measured runs
5
Generation stddev
0.009518
Thermal ceiling
90 °C
Source revision
8be84a0f4727

Artifact SHA-256a72d86d4cb4fedd940e34c08d008bb5cda42db80ce5c6bc5f9494e854a3d742d

Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 1, MTP off
11.34 tok/sn/a74 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
6.043 tok/sn/a75 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 46.73% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 2, MTP off
11.49 tok/sn/a77 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
4.019 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 65.03% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 3, MTP off
11.21 tok/sn/a76 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
3.133 tok/sn/a77 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 72.06% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Depth 4, MTP off
11.16 tok/sn/a78 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
2.624 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 76.48% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 1, MTP off
4.891 tok/sn/a78 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
3.778 tok/sn/a77 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 22.76% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 2, MTP off
5.154 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
2.853 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 44.65% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 3, MTP off
4.853 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
2.599 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 46.45% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Depth 4, MTP off
5.121 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
1.522 tok/sn/a79 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 70.28% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 1, MTP off
3.269 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
2.266 tok/sn/a84 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 30.68% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 2, MTP off
3.359 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
1.96 tok/sn/a84 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 41.63% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 3, MTP off
3.352 tok/sn/a82 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
1.474 tok/sn/a86 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 56.03% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 4, MTP off
3.397 tok/sn/a81 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
1.333 tok/sn/a84 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 60.75% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Two-run warm average
10.44 tok/sn/a70 °CPublished resultAdditional models
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

The fastest normal text-generation result in the X1S work. Throughput does not establish answer quality.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Two-run warm average
4.89 tok/sn/a75 °CPublished resultAdditional models
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

A clean middle-tier CPU result from the normal non-MTP Ollama path.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C
Qwen3.5 9B
Q4_K_M
Ollama 0.32.1
CPU · Five-prompt average
1.11 tok/sn/a82 °CPublished resultAdditional models
Open result details

Why this model is here

The official Q4_K_M Ollama artifact fit in 16 GB and averaged 1.110 tok/s across five prompts.

This was a fit-and-completion test for a much larger model on the X1S, not a claim that 1.110 tok/s is comfortable interactive speed.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Five-prompt average mode on Youyeetoo X1S.

The 9B artifact fit in 16 GB and completed five prompts, but generation requires patience.

Recorded setup

Measured runs
5
Thermal ceiling
85 °C

Artifact SHA-256dec52a44569a2a25341c4e4d3fee25846eed4f6f0b936278e3a3c900bb99d37c

Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Two-run warm average
3.38 tok/sn/a77 °CPublished resultAdditional models
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

The normal non-MTP Ollama result. The separate MTP sweep is recorded in its own test group.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C
Gemma 4 E4B
Q4_K_M
Ollama 0.32.1
CPU · Two-run warm average
1.78 tok/sn/a78 °CPublished resultAdditional models
Open result details

Why this model is here

Gemma 4 E4B completed the original warm Ollama test at 1.78 tok/s.

The row is useful as a recorded X1S operating point, but it does not have enough separate material for an indexable model page yet.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

A clean normal Ollama row retained as an operating point, not a quality ranking.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C
Granite 4 Tiny-H
Original Ollama tag
Ollama 0.32.1
CPU · Single original observation
5.43 tok/sn/a79 °CPartial resultAdditional models
Open result details

Why this model is here

The X1S recorded one 5.43 tok/s Ollama observation.

One observation is retained for completeness but is not promoted into a deeper standalone page.

What ran

Original Ollama tag ran through Ollama 0.32.1 using CPU in Single original observation mode on Youyeetoo X1S.

One retained observation, not a repeated average.

Recorded setup

Measured runs
1
Thermal ceiling
85 °C
LFM2.5
Mutable latest tag
Ollama 0.32.1
CPU · Single recorded tag state
4.95 tok/sn/a80 °CPartial resultAdditional models
Open result details

Why this model is here

The X1S recorded one 4.95 tok/s result from a mutable latest tag.

The tag was not content-pinned, so the row stays visible as a partial result without becoming its own search page.

What ran

Mutable latest tag ran through Ollama 0.32.1 using CPU in Single recorded tag state mode on Youyeetoo X1S.

The mutable latest tag was not content-pinned, so this remains a partial result.

Recorded setup

Measured runs
1
Thermal ceiling
85 °C
Gemma 3n E2B
gemma3n:e2b
Ollama 0.32.1
CPU · Single original observation
3.24 tok/sn/a80 °CPartial resultAdditional models
Open result details

Why this model is here

The X1S recorded one 3.24 tok/s Ollama result.

The row is kept separate from Gemma 4 E2B and does not yet have enough unique material for its own page.

What ran

gemma3n:e2b ran through Ollama 0.32.1 using CPU in Single original observation mode on Youyeetoo X1S.

One retained observation. This is Gemma 3n E2B, not Gemma 4 E2B.

Recorded setup

Measured runs
1
Thermal ceiling
85 °C
BitCPM-CANN 1B
Official TQ2_0 GGUF
llama.cpp 9a286ac
CPU · True CPU, five repetitions
8.584 tok/s13.01 tok/s76 °CPublished resultRuntime compatibility
Open result details

Why this model is here

The exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1.

This is a runtime-support result, not a speed contest. The same verified file produced inference in one runtime and a tensor-size overflow in the other.

What ran

Official TQ2_0 GGUF ran through llama.cpp 9a286ac using CPU in True CPU, five repetitions mode on Youyeetoo X1S.

Pinned llama.cpp accepted the verified file and completed the controlled true-CPU workload.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Generation stddev
0.029317
Thermal ceiling
85 °C

Artifact SHA-2562394c15cbea2181b72bfb4215d8417d8d1f2f6214069da2d01fde32ce3b13fce

BitCPM-CANN 1B
Official TQ2_0 GGUF
Ollama 0.32.1
CPU · Model load
n/an/an/aFailed runRuntime compatibility
Open result details

Why this model is here

The exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1.

This is a runtime-support result, not a speed contest. The same verified file produced inference in one runtime and a tensor-size overflow in the other.

What ran

Official TQ2_0 GGUF ran through Ollama 0.32.1 using CPU in Model load mode on Youyeetoo X1S.

Ollama rejected the same verified file with a tensor-size overflow before inference.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Artifact SHA-2562394c15cbea2181b72bfb4215d8417d8d1f2f6214069da2d01fde32ce3b13fce

Gemma 3 4B
Separate text-only derivative
llama.cpp 9a286ac
CPU · True CPU, five repetitions
1.624 tok/s2.172 tok/s81 °CPublished resultRuntime compatibility
Open result details

Why this model is here

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

What ran

Separate text-only derivative ran through llama.cpp 9a286ac using CPU in True CPU, five repetitions mode on Youyeetoo X1S.

The derived copy completed true CPU. The original multimodal source artifact was not overwritten.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Generation stddev
0.003291
Thermal ceiling
85 °C

Artifact SHA-256510408e8043ca1c741fe9a16088d47e8fa0d016033c6acf0c50c50c7c93b6530

Qwen3 0.6B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
6.788 tok/sn/a74 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

The smallest matched X1S model completed both CPU runtimes and all three explicit llama.cpp device modes.

It was small enough to show the CPU, mixed host-operation, and full-Vulkan differences without crossing into the larger-model i915 failures.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Fastest row, least complete response on the single prompt.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 1.7B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
3.129 tok/sn/a77 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

This model completed the matched CPU comparison and the full CPU, mixed, and Vulkan mode test.

The original X1S work put it in the usable interactive tier, so it became the second clean model for checking whether the Vulkan pattern held beyond 0.6B.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Best interactive starting tier tested.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 4B Instruct
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.824 tok/sn/a77 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

The Q4_K_M file completed the matched CPU test, then reached an i915 reset timeout during full Vulkan.

It marks the point where the clean small-model Vulkan result stopped scaling into a reliable larger-model run on this software stack.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Phi-4 Mini
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.991 tok/sn/a77 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Gemma 3 4B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.995 tok/sn/a77 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen3 8B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
0.924 tok/sn/a80 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.

It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Fit comfortably with sub-1 tok/s generation.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"