Youyeetoo X1S
The deepest dataset here: Ollama CPU requests, llama.cpp microbenchmarks, two clean Vulkan pairs, thermal readings, and retained GPU failure evidence.
Open the write-upThis pulls my local inference results into one place. Search by device, model, runtime, or test, then open the original write-up for the setup, limits, failures, and source material behind the number.
Comparison boundary: only rows in the same test group should be compared directly. Ollama requests, synthetic llama-bench runs, RKLLM NPU tests, and published Pi references are different setups.
| Device and model | Runtime | Test group | Generation | Prompt | Peak | Evidence | Source |
|---|---|---|---|---|---|---|---|
Qwen3 0.6B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S original Ollama matrix Fastest in the original matrix, but the single response was the least complete. | 6.79 tok/s | n/a | 74 °C | Measured | Read test |
Qwen3 1.7B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S original Ollama matrix Best interactive starting tier in the original test. | 3.13 tok/s | n/a | 77 °C | Measured | Read test |
Qwen3 4B Instruct Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S original Ollama matrix Usable for patient local batch work. | 1.82 tok/s | n/a | 77 °C | Measured | Read test |
Phi-4 Mini Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S original Ollama matrix Usable for patient local batch work. | 1.99 tok/s | n/a | 77 °C | Measured | Read test |
Gemma 3 4B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S original Ollama matrix Usable for patient local batch work. | 2.00 tok/s | n/a | 77 °C | Measured | Read test |
Qwen3 8B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S original Ollama matrix Fit in memory and stopped naturally, but generation was below 1 tok/s. | 0.92 tok/s | n/a | 80 °C | Measured | Read test |
Qwen3 0.6B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S round 2 Ollama API Warm generation. The response stopped naturally at 54 tokens. | 7.81 tok/s | n/a | n/a | Measured | Read test |
Qwen3 1.7B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S round 2 Ollama API Warm generation. The response stopped naturally at 56 tokens. | 3.47 tok/s | n/a | n/a | Measured | Read test |
Qwen3 4B Instruct Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S round 2 Ollama API Warm generation through the Ollama API. | 1.98 tok/s | n/a | n/a | Measured | Read test |
Phi-4 Mini Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S round 2 Ollama API Warm generation through the Ollama API. | 2.10 tok/s | n/a | n/a | Measured | Read test |
Gemma 3 4B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S round 2 Ollama API Warm generation through the Ollama API. | 2.12 tok/s | n/a | n/a | Measured | Read test |
Qwen3 8B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S round 2 Ollama API Warm generation. The response stopped naturally at 84 tokens. | 0.98 tok/s | n/a | n/a | Measured | Read test |
Granite 4 Tiny-H Youyeetoo X1S · Intel Celeron N5095 · Ollama tag | Ollama CPU | X1S candidate Ollama models Same guarded Ollama request used for the candidate model pass. | 5.43 tok/s | n/a | 79 °C | Measured | Read test |
LFM2.5 Youyeetoo X1S · Intel Celeron N5095 · Ollama tag | Ollama CPU | X1S candidate Ollama models Same guarded Ollama request used for the candidate model pass. | 4.95 tok/s | n/a | 80 °C | Measured | Read test |
Gemma 3n E2B Youyeetoo X1S · Intel Celeron N5095 · Ollama tag | Ollama CPU | X1S candidate Ollama models This is Gemma 3n E2B, not Gemma 4 E2B. | 3.24 tok/s | n/a | 80 °C | Measured | Read test |
Qwen3.5 0.8B Youyeetoo X1S · Intel Celeron N5095 · Q8_0 | Ollama CPU | X1S official Ollama follow-up Two-run warm average. The cold result was 10.43 tok/s. | 10.4 tok/s | n/a | 70 °C | Measured | Read test |
Qwen3.5 2B Youyeetoo X1S · Intel Celeron N5095 · Q8_0 | Ollama CPU | X1S official Ollama follow-up Two-run warm average. The cold result was 4.97 tok/s. | 4.89 tok/s | n/a | 75 °C | Measured | Read test |
Gemma 4 E2B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S official Ollama follow-up Two-run warm average. The cold result was 3.36 tok/s. | 3.38 tok/s | n/a | 77 °C | Measured | Read test |
Gemma 4 E4B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | Ollama CPU | X1S official Ollama follow-up Cold and two-run warm average both recorded 1.78 tok/s. | 1.78 tok/s | n/a | 78 °C | Measured | Read test |
BitCPM-CANN 1B Youyeetoo X1S · Intel Celeron N5095 · TQ2_0 | llama.cpp CPU | X1S llama-bench CPU Two repetitions with zero GPU layers. Ollama 0.32.1 could not load the same file. | 8.63 tok/s | 13.30 tok/s | 68 °C | Measured | Read test |
BitCPM-CANN 1B Youyeetoo X1S · Intel Celeron N5095 · TQ2_0 | Ollama 0.32.1 CPU | X1S runtime compatibility Ollama rejected the verified file during model loading with a tensor size overflow. | n/a | n/a | n/a | Diagnostic | Read test |
Qwen3 0.6B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp CPU | X1S llama-bench CPU Synthetic CPU microbenchmark. Compare directly only with the matched llama-bench Vulkan row. | 7.42 tok/s | n/a | 80 °C | Measured | Read test |
Qwen3 1.7B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp CPU | X1S llama-bench CPU Synthetic CPU microbenchmark. Compare directly only with the matched llama-bench Vulkan row. | 2.98 tok/s | n/a | 82 °C | Measured | Read test |
Qwen3 4B Instruct Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp CPU | X1S llama-bench CPU Synthetic CPU microbenchmark. This is not the same harness as the Ollama API request. | 1.58 tok/s | n/a | n/a | Measured | Read test |
Phi-4 Mini Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp CPU | X1S llama-bench CPU Synthetic CPU microbenchmark. This is not the same harness as the Ollama API request. | 1.64 tok/s | n/a | n/a | Measured | Read test |
Gemma 3 4B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M text-only copy | llama.cpp CPU | X1S llama-bench CPU Current llama.cpp required a derived text-only compatibility copy. | 1.64 tok/s | n/a | n/a | Measured | Read test |
Qwen3 8B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp CPU | X1S llama-bench CPU The CPU test hit the ten-minute cap before a result was accepted. | n/a | n/a | n/a | Diagnostic | Read test |
Qwen3 0.6B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp Vulkan | X1S clean Vulkan pairs Clean Vulkan run. Matched CPU llama-bench generation was 7.42 tok/s. | 8.56 tok/s | 37.29 tok/s | 53 °C | Measured | Read test |
Qwen3 1.7B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp Vulkan | X1S clean Vulkan pairs Clean Vulkan run. Matched CPU llama-bench generation was 2.98 tok/s. | 3.48 tok/s | 12.86 tok/s | 56 °C | Measured | Read test |
Qwen3 4B Instruct Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp Vulkan | X1S unstable Vulkan diagnostics Kernel reset timeout occurred during the run window. Do not use as a clean performance result. | 1.46 tok/s | n/a | n/a | Diagnostic | Read test |
Phi-4 Mini Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp Vulkan | X1S unstable Vulkan diagnostics An i915 GPU hang was logged. The process return code did not make this a clean result. | 1.48 tok/s | n/a | n/a | Diagnostic | Read test |
Qwen3 8B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M | llama.cpp Vulkan | X1S unstable Vulkan diagnostics An i915 GPU hang was logged. The CPU test hit the ten-minute cap. | 0.84 tok/s | n/a | n/a | Diagnostic | Read test |
Gemma 3 4B Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M text-only copy | llama.cpp Vulkan | X1S unstable Vulkan diagnostics Fence timeouts, an i915 preemption reset, and vk::DeviceLostError ended the run. | n/a | n/a | n/a | Diagnostic | Read test |
DeepSeek-R1 Distill Qwen 1.5B Indiedroid Nova · Rockchip RK3588S · W8A8 | RKLLM 1.2.1 NPU | Nova RKLLM NPU test Fastest Nova result, but the output began inventing cities around the seventh state. | 11.5 tok/s | n/a | n/a | Measured | Read test |
Qwen 2.5 3B Instruct Indiedroid Nova · Rockchip RK3588S · W8A8 | RKLLM 1.2.1 NPU | Nova RKLLM NPU test Returned all 50 state capitals correctly in the recorded test. | 7.00 tok/s | n/a | n/a | Measured | Read test |
Llama 3.1 8B Instruct Indiedroid Nova · Rockchip RK3588S · W8A8 | RKLLM 1.2.1 NPU | Nova RKLLM NPU test Nova measurement. Keep separate from the Pi reference because runtime and quantization differ. | 3.72 tok/s | n/a | n/a | Measured | Read test |
DeepSeek-R1 Distill Qwen 1.5B Raspberry Pi 5 · Broadcom BCM2712 · Q4 | llama.cpp CPU | Pi 5 external reference Published range was approximately 6 to 8 tok/s. The midpoint is shown only for filtering and is not Trevor-measured. | 7.00 tok/s | n/a | n/a | External reference | Read test |
Qwen 2.5 3B Instruct Raspberry Pi 5 · Broadcom BCM2712 · Q4 | llama.cpp CPU | Pi 5 external reference Published range was approximately 4 to 5 tok/s. The midpoint is shown only for filtering and is not Trevor-measured. | 4.50 tok/s | n/a | n/a | External reference | Read test |
Llama 3.1 8B Instruct Raspberry Pi 5 · Broadcom BCM2712 · Q4 | llama.cpp CPU | Pi 5 external reference Jeff Geerling result cited in the Nova article. It is not Trevor-measured or a matched Nova comparison. | 1.99 tok/s | n/a | n/a | External reference | Read test |
The deepest dataset here: Ollama CPU requests, llama.cpp microbenchmarks, two clean Vulkan pairs, thermal readings, and retained GPU failure evidence.
Open the write-upThree W8A8 models on the RK3588S NPU through RKLLM. The recorded prompt also caught the difference between raw speed and useful output.
Open the write-upThese are clearly marked external references from the Nova article, not my own matched Pi 5 rerun. A controlled board-to-board test is still open work.
Open the write-upGeneration speed is shown in tokens per second. Prompt speed appears only where the original test recorded it. Temperature is the peak package reading from that run or test group, not a claimed maximum for the device. Every row links back to the article that explains the hardware, software, cooling, prompt, and limits.
A Q4 Ollama request at a 4,096-token context is not the same test as a synthetic pp128/tg96 microbenchmark. An RK3588S NPU running W8A8 RKLLM is not a matched comparison with a Raspberry Pi 5 running Q4 llama.cpp. Keeping those boundaries visible makes the database more useful than sorting unrelated numbers from fastest to slowest.
The next useful step is a repeatable cross-device protocol with the same model artifact, prompt, context, token limit, cooling disclosure, and failure checks on every board. That will make the device dropdown useful for real head-to-head comparisons, while the existing results stay available as the tests they actually were.