/benchmarks

Local AI Benchmark Hub

This pulls my local inference results into one place. Search by device, model, runtime, or test, then open the original write-up for the setup, limits, failures, and source material behind the number.

39
published rows
30
my measured rows
3
devices represented
39
matching this filter

Comparison boundary: only rows in the same test group should be compared directly. Ollama requests, synthetic llama-bench runs, RKLLM NPU tests, and published Pi references are different setups.

Benchmark results

Device and modelRuntimeTest groupGenerationPromptPeakEvidenceSource
Qwen3 0.6B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S original Ollama matrix
Fastest in the original matrix, but the single response was the least complete.
6.79 tok/sn/a74 °CMeasuredRead test
Qwen3 1.7B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S original Ollama matrix
Best interactive starting tier in the original test.
3.13 tok/sn/a77 °CMeasuredRead test
Qwen3 4B Instruct
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S original Ollama matrix
Usable for patient local batch work.
1.82 tok/sn/a77 °CMeasuredRead test
Phi-4 Mini
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S original Ollama matrix
Usable for patient local batch work.
1.99 tok/sn/a77 °CMeasuredRead test
Gemma 3 4B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S original Ollama matrix
Usable for patient local batch work.
2.00 tok/sn/a77 °CMeasuredRead test
Qwen3 8B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S original Ollama matrix
Fit in memory and stopped naturally, but generation was below 1 tok/s.
0.92 tok/sn/a80 °CMeasuredRead test
Qwen3 0.6B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S round 2 Ollama API
Warm generation. The response stopped naturally at 54 tokens.
7.81 tok/sn/an/aMeasuredRead test
Qwen3 1.7B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S round 2 Ollama API
Warm generation. The response stopped naturally at 56 tokens.
3.47 tok/sn/an/aMeasuredRead test
Qwen3 4B Instruct
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S round 2 Ollama API
Warm generation through the Ollama API.
1.98 tok/sn/an/aMeasuredRead test
Phi-4 Mini
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S round 2 Ollama API
Warm generation through the Ollama API.
2.10 tok/sn/an/aMeasuredRead test
Gemma 3 4B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S round 2 Ollama API
Warm generation through the Ollama API.
2.12 tok/sn/an/aMeasuredRead test
Qwen3 8B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S round 2 Ollama API
Warm generation. The response stopped naturally at 84 tokens.
0.98 tok/sn/an/aMeasuredRead test
Granite 4 Tiny-H
Youyeetoo X1S · Intel Celeron N5095 · Ollama tag
Ollama
CPU
X1S candidate Ollama models
Same guarded Ollama request used for the candidate model pass.
5.43 tok/sn/a79 °CMeasuredRead test
LFM2.5
Youyeetoo X1S · Intel Celeron N5095 · Ollama tag
Ollama
CPU
X1S candidate Ollama models
Same guarded Ollama request used for the candidate model pass.
4.95 tok/sn/a80 °CMeasuredRead test
Gemma 3n E2B
Youyeetoo X1S · Intel Celeron N5095 · Ollama tag
Ollama
CPU
X1S candidate Ollama models
This is Gemma 3n E2B, not Gemma 4 E2B.
3.24 tok/sn/a80 °CMeasuredRead test
Qwen3.5 0.8B
Youyeetoo X1S · Intel Celeron N5095 · Q8_0
Ollama
CPU
X1S official Ollama follow-up
Two-run warm average. The cold result was 10.43 tok/s.
10.4 tok/sn/a70 °CMeasuredRead test
Qwen3.5 2B
Youyeetoo X1S · Intel Celeron N5095 · Q8_0
Ollama
CPU
X1S official Ollama follow-up
Two-run warm average. The cold result was 4.97 tok/s.
4.89 tok/sn/a75 °CMeasuredRead test
Gemma 4 E2B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S official Ollama follow-up
Two-run warm average. The cold result was 3.36 tok/s.
3.38 tok/sn/a77 °CMeasuredRead test
Gemma 4 E4B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
Ollama
CPU
X1S official Ollama follow-up
Cold and two-run warm average both recorded 1.78 tok/s.
1.78 tok/sn/a78 °CMeasuredRead test
BitCPM-CANN 1B
Youyeetoo X1S · Intel Celeron N5095 · TQ2_0
llama.cpp
CPU
X1S llama-bench CPU
Two repetitions with zero GPU layers. Ollama 0.32.1 could not load the same file.
8.63 tok/s13.30 tok/s68 °CMeasuredRead test
BitCPM-CANN 1B
Youyeetoo X1S · Intel Celeron N5095 · TQ2_0
Ollama 0.32.1
CPU
X1S runtime compatibility
Ollama rejected the verified file during model loading with a tensor size overflow.
n/an/an/aDiagnosticRead test
Qwen3 0.6B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
CPU
X1S llama-bench CPU
Synthetic CPU microbenchmark. Compare directly only with the matched llama-bench Vulkan row.
7.42 tok/sn/a80 °CMeasuredRead test
Qwen3 1.7B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
CPU
X1S llama-bench CPU
Synthetic CPU microbenchmark. Compare directly only with the matched llama-bench Vulkan row.
2.98 tok/sn/a82 °CMeasuredRead test
Qwen3 4B Instruct
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
CPU
X1S llama-bench CPU
Synthetic CPU microbenchmark. This is not the same harness as the Ollama API request.
1.58 tok/sn/an/aMeasuredRead test
Phi-4 Mini
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
CPU
X1S llama-bench CPU
Synthetic CPU microbenchmark. This is not the same harness as the Ollama API request.
1.64 tok/sn/an/aMeasuredRead test
Gemma 3 4B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M text-only copy
llama.cpp
CPU
X1S llama-bench CPU
Current llama.cpp required a derived text-only compatibility copy.
1.64 tok/sn/an/aMeasuredRead test
Qwen3 8B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
CPU
X1S llama-bench CPU
The CPU test hit the ten-minute cap before a result was accepted.
n/an/an/aDiagnosticRead test
Qwen3 0.6B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
Vulkan
X1S clean Vulkan pairs
Clean Vulkan run. Matched CPU llama-bench generation was 7.42 tok/s.
8.56 tok/s37.29 tok/s53 °CMeasuredRead test
Qwen3 1.7B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
Vulkan
X1S clean Vulkan pairs
Clean Vulkan run. Matched CPU llama-bench generation was 2.98 tok/s.
3.48 tok/s12.86 tok/s56 °CMeasuredRead test
Qwen3 4B Instruct
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
Vulkan
X1S unstable Vulkan diagnostics
Kernel reset timeout occurred during the run window. Do not use as a clean performance result.
1.46 tok/sn/an/aDiagnosticRead test
Phi-4 Mini
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
Vulkan
X1S unstable Vulkan diagnostics
An i915 GPU hang was logged. The process return code did not make this a clean result.
1.48 tok/sn/an/aDiagnosticRead test
Qwen3 8B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M
llama.cpp
Vulkan
X1S unstable Vulkan diagnostics
An i915 GPU hang was logged. The CPU test hit the ten-minute cap.
0.84 tok/sn/an/aDiagnosticRead test
Gemma 3 4B
Youyeetoo X1S · Intel Celeron N5095 · Q4_K_M text-only copy
llama.cpp
Vulkan
X1S unstable Vulkan diagnostics
Fence timeouts, an i915 preemption reset, and vk::DeviceLostError ended the run.
n/an/an/aDiagnosticRead test
DeepSeek-R1 Distill Qwen 1.5B
Indiedroid Nova · Rockchip RK3588S · W8A8
RKLLM 1.2.1
NPU
Nova RKLLM NPU test
Fastest Nova result, but the output began inventing cities around the seventh state.
11.5 tok/sn/an/aMeasuredRead test
Qwen 2.5 3B Instruct
Indiedroid Nova · Rockchip RK3588S · W8A8
RKLLM 1.2.1
NPU
Nova RKLLM NPU test
Returned all 50 state capitals correctly in the recorded test.
7.00 tok/sn/an/aMeasuredRead test
Llama 3.1 8B Instruct
Indiedroid Nova · Rockchip RK3588S · W8A8
RKLLM 1.2.1
NPU
Nova RKLLM NPU test
Nova measurement. Keep separate from the Pi reference because runtime and quantization differ.
3.72 tok/sn/an/aMeasuredRead test
DeepSeek-R1 Distill Qwen 1.5B
Raspberry Pi 5 · Broadcom BCM2712 · Q4
llama.cpp
CPU
Pi 5 external reference
Published range was approximately 6 to 8 tok/s. The midpoint is shown only for filtering and is not Trevor-measured.
7.00 tok/sn/an/aExternal referenceRead test
Qwen 2.5 3B Instruct
Raspberry Pi 5 · Broadcom BCM2712 · Q4
llama.cpp
CPU
Pi 5 external reference
Published range was approximately 4 to 5 tok/s. The midpoint is shown only for filtering and is not Trevor-measured.
4.50 tok/sn/an/aExternal referenceRead test
Llama 3.1 8B Instruct
Raspberry Pi 5 · Broadcom BCM2712 · Q4
llama.cpp
CPU
Pi 5 external reference
Jeff Geerling result cited in the Nova article. It is not Trevor-measured or a matched Nova comparison.
1.99 tok/sn/an/aExternal referenceRead test

Youyeetoo X1S

The deepest dataset here: Ollama CPU requests, llama.cpp microbenchmarks, two clean Vulkan pairs, thermal readings, and retained GPU failure evidence.

Open the write-up

Indiedroid Nova

Three W8A8 models on the RK3588S NPU through RKLLM. The recorded prompt also caught the difference between raw speed and useful output.

Open the write-up

Raspberry Pi 5

These are clearly marked external references from the Nova article, not my own matched Pi 5 rerun. A controlled board-to-board test is still open work.

Open the write-up

How to read these local AI results

Generation speed is shown in tokens per second. Prompt speed appears only where the original test recorded it. Temperature is the peak package reading from that run or test group, not a claimed maximum for the device. Every row links back to the article that explains the hardware, software, cooling, prompt, and limits.

Why this is not one giant leaderboard

A Q4 Ollama request at a 4,096-token context is not the same test as a synthetic pp128/tg96 microbenchmark. An RK3588S NPU running W8A8 RKLLM is not a matched comparison with a Raspberry Pi 5 running Q4 llama.cpp. Keeping those boundaries visible makes the database more useful than sorting unrelated numbers from fastest to slowest.

What I want to add next

The next useful step is a repeatable cross-device protocol with the same model artifact, prompt, context, token limit, cooling disclosure, and failure checks on every board. That will make the device dropdown useful for real head-to-head comparisons, while the existing results stay available as the tests they actually were.

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"