/hub/runs/x1s-requested-models
Published result2026-08-31Youyeetoo X1S

Additional Model Results on the X1S

Eight additional models were recorded with their actual averaging method, package peak, and artifact limitation where one applied.

This page keeps the shared controls, retained measurements, failures, and comparison limits together. The linked files are there when you need the underlying table or method.

Comparison boundary

The rows use different repetition counts and artifact states. They are useful operating points, not one controlled model-quality contest.

Recorded results

Open the details under any row for the specific setup, outcome, artifact identity, and supporting file.

Model and artifactRuntime and modeGenerationPromptPeakState
Qwen3.5 0.8B
Q8_0
Ollama 0.32.1
CPU · Two-run warm average
10.44 tok/sn/a70 °CPublished result
Open result details

Why this model is here

Qwen3.5 0.8B produced the fastest normal text-generation result in the X1S work and completed every MTP depth pair.

It is the speed-first CPU option from this result set and shows that active MTP still carried too much overhead on the four-core N5095.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

The fastest normal text-generation result in the X1S work. Throughput does not establish answer quality.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C
Qwen3.5 2B
Q8_0
Ollama 0.32.1
CPU · Two-run warm average
4.89 tok/sn/a75 °CPublished result
Open result details

Why this model is here

Qwen3.5 2B completed the normal Ollama run and MTP depths one through four.

It sits between the 0.8B speed result and the slower larger models, making it useful for checking whether MTP behavior changed with model size.

What ran

Q8_0 ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

A clean middle-tier CPU result from the normal non-MTP Ollama path.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C
Qwen3.5 9B
Q4_K_M
Ollama 0.32.1
CPU · Five-prompt average
1.11 tok/sn/a82 °CPublished result
Open result details

Why this model is here

The official Q4_K_M Ollama artifact fit in 16 GB and averaged 1.110 tok/s across five prompts.

This was a fit-and-completion test for a much larger model on the X1S, not a claim that 1.110 tok/s is comfortable interactive speed.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Five-prompt average mode on Youyeetoo X1S.

The 9B artifact fit in 16 GB and completed five prompts, but generation requires patience.

Recorded setup

Measured runs
5
Thermal ceiling
85 °C

Artifact SHA-256dec52a44569a2a25341c4e4d3fee25846eed4f6f0b936278e3a3c900bb99d37c

Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Two-run warm average
3.38 tok/sn/a77 °CPublished result
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

The normal non-MTP Ollama result. The separate MTP sweep is recorded in its own test group.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C
Gemma 4 E4B
Q4_K_M
Ollama 0.32.1
CPU · Two-run warm average
1.78 tok/sn/a78 °CPublished result
Open result details

Why this model is here

Gemma 4 E4B completed the original warm Ollama test at 1.78 tok/s.

The row is useful as a recorded X1S operating point, but it does not have enough separate material for an indexable model page yet.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

A clean normal Ollama row retained as an operating point, not a quality ranking.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C
Granite 4 Tiny-H
Original Ollama tag
Ollama 0.32.1
CPU · Single original observation
5.43 tok/sn/a79 °CPartial result
Open result details

Why this model is here

The X1S recorded one 5.43 tok/s Ollama observation.

One observation is retained for completeness but is not promoted into a deeper standalone page.

What ran

Original Ollama tag ran through Ollama 0.32.1 using CPU in Single original observation mode on Youyeetoo X1S.

One retained observation, not a repeated average.

Recorded setup

Measured runs
1
Thermal ceiling
85 °C
LFM2.5
Mutable latest tag
Ollama 0.32.1
CPU · Single recorded tag state
4.95 tok/sn/a80 °CPartial result
Open result details

Why this model is here

The X1S recorded one 4.95 tok/s result from a mutable latest tag.

The tag was not content-pinned, so the row stays visible as a partial result without becoming its own search page.

What ran

Mutable latest tag ran through Ollama 0.32.1 using CPU in Single recorded tag state mode on Youyeetoo X1S.

The mutable latest tag was not content-pinned, so this remains a partial result.

Recorded setup

Measured runs
1
Thermal ceiling
85 °C
Gemma 3n E2B
gemma3n:e2b
Ollama 0.32.1
CPU · Single original observation
3.24 tok/sn/a80 °CPartial result
Open result details

Why this model is here

The X1S recorded one 3.24 tok/s Ollama result.

The row is kept separate from Gemma 4 E2B and does not yet have enough unique material for its own page.

What ran

gemma3n:e2b ran through Ollama 0.32.1 using CPU in Single original observation mode on Youyeetoo X1S.

One retained observation. This is Gemma 3n E2B, not Gemma 4 E2B.

Recorded setup

Measured runs
1
Thermal ceiling
85 °C

Supporting evidence

Results and method

These are the few files that support this page directly. The repository holds the wider project history.

Controls

  • Ollama CPU workflow
  • Normal 85 °C thermal abort unless disclosed elsewhere
  • Kernel and request timeout checks enabled
  • Method recorded per row

What this run established

  • $Qwen3.5 0.8B was the fastest text-generation result in the project.
  • $Qwen3.5 9B fit and completed five prompts at 1.110 tok/s.
  • $LFM2.5 used a mutable latest tag and is not content-pinned.

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"