/hub/models/qwen25-3b

Qwen 2.5 · 3B

Qwen 2.5 3B Instruct

The Nova NPU result reached 7.0 tok/s and returned all 50 state capitals correctly.

2
result records
1
Trevor-measured rows
2
test groups
2
devices represented

Why Qwen 2.5 3B Instruct is in The Hub

The Nova NPU result reached 7.0 tok/s and returned all 50 state capitals correctly.

It was slower than the smaller DeepSeek row but produced the useful answer on the shared prompt.

The rows on this page use W8A8 RKLLM and Q4 reference through RKLLM 1.2.1 and llama.cpp. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 7 tok/s through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode. Returned all 50 state capitals correctly.

1 row is an outside reference, kept separate from Trevor-measured results.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Qwen 2.5 3B Instruct
W8A8 RKLLM
RKLLM 1.2.1
NPU · 50 state-capitals prompt
7 tok/sn/an/aPublished resultNova NPU models
Open result details

Why this model is here

The Nova NPU result reached 7.0 tok/s and returned all 50 state capitals correctly.

It was slower than the smaller DeepSeek row but produced the useful answer on the shared prompt.

What ran

W8A8 RKLLM ran through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode on Indiedroid Nova.

Returned all 50 state capitals correctly.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Qwen 2.5 3B Instruct
Q4 reference
llama.cpp
CPU · Published reference range
about 4 to 5 tok/sn/an/aExternal referencePi 5 references
Open result details

Why this model is here

The Nova NPU result reached 7.0 tok/s and returned all 50 state capitals correctly.

It was slower than the smaller DeepSeek row but produced the useful answer on the shared prompt.

What ran

Q4 reference ran through llama.cpp using CPU in Published reference range mode on Raspberry Pi 5.

Approximate published range. Not Trevor-measured.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing Qwen 2.5 3B Instruct

2026-04-10

Indiedroid Nova RKLLM NPU Test

The three Nova rows share the local test. Raspberry Pi numbers in the article came from outside references with different runtime and quantization.

Open full test

2026-04-10

Raspberry Pi 5 Published Reference Rows

These are external references with different runtime, quantization, and hardware approach. A matched Trevor-run Pi 5 test has not been completed.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"