/hub/models/deepseek-r1-15b

DeepSeek-R1 · 1.5B

DeepSeek-R1 Distill Qwen 1.5B

The Nova NPU run started at 11.5 tok/s, then the state-capitals answer degraded into invented cities.

2
result records
1
Trevor-measured rows
2
test groups
2
devices represented

Why DeepSeek-R1 Distill Qwen 1.5B is in The Hub

The Nova NPU run started at 11.5 tok/s, then the state-capitals answer degraded into invented cities.

It is the clearest example in the Hub of why throughput and useful output cannot be treated as the same result.

The rows on this page use W8A8 RKLLM and Q4 reference through RKLLM 1.2.1 and llama.cpp. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 11.5 tok/s through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode. Started fast, then invented cities around state seven.

1 row is an outside reference, kept separate from Trevor-measured results.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
DeepSeek-R1 Distill Qwen 1.5B
W8A8 RKLLM
RKLLM 1.2.1
NPU · 50 state-capitals prompt
11.5 tok/sn/an/aPublished resultNova NPU models
Open result details

Why this model is here

The Nova NPU run started at 11.5 tok/s, then the state-capitals answer degraded into invented cities.

It is the clearest example in the Hub of why throughput and useful output cannot be treated as the same result.

What ran

W8A8 RKLLM ran through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode on Indiedroid Nova.

Started fast, then invented cities around state seven.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

DeepSeek-R1 Distill Qwen 1.5B
Q4 reference
llama.cpp
CPU · Published reference range
about 6 to 8 tok/sn/an/aExternal referencePi 5 references
Open result details

Why this model is here

The Nova NPU run started at 11.5 tok/s, then the state-capitals answer degraded into invented cities.

It is the clearest example in the Hub of why throughput and useful output cannot be treated as the same result.

What ran

Q4 reference ran through llama.cpp using CPU in Published reference range mode on Raspberry Pi 5.

Approximate published range. Not Trevor-measured.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing DeepSeek-R1 Distill Qwen 1.5B

2026-04-10

Indiedroid Nova RKLLM NPU Test

The three Nova rows share the local test. Raspberry Pi numbers in the article came from outside references with different runtime and quantization.

Open full test

2026-04-10

Raspberry Pi 5 Published Reference Rows

These are external references with different runtime, quantization, and hardware approach. A matched Trevor-run Pi 5 test has not been completed.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"