/hub/models/llama31-8b

Llama 3.1 · 8B

Llama 3.1 8B Instruct

The Nova NPU measured 3.72 tok/s. A separate 1.99 tok/s Pi 5 value remains labeled as an outside reference.

2
result records
1
Trevor-measured rows
2
test groups
2
devices represented

Why Llama 3.1 8B Instruct is in The Hub

The Nova NPU measured 3.72 tok/s. A separate 1.99 tok/s Pi 5 value remains labeled as an outside reference.

It connects Trevor-measured Nova data with a clearly separated published Pi reference without presenting the two methods as a matched hardware contest.

The rows on this page use W8A8 RKLLM and Q4 reference through RKLLM 1.2.1 and llama.cpp. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 3.72 tok/s through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode. Measured Nova NPU result.

1 row is an outside reference, kept separate from Trevor-measured results.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Llama 3.1 8B Instruct
W8A8 RKLLM
RKLLM 1.2.1
NPU · 50 state-capitals prompt
3.72 tok/sn/an/aPublished resultNova NPU models
Open result details

Why this model is here

The Nova NPU measured 3.72 tok/s. A separate 1.99 tok/s Pi 5 value remains labeled as an outside reference.

It connects Trevor-measured Nova data with a clearly separated published Pi reference without presenting the two methods as a matched hardware contest.

What ran

W8A8 RKLLM ran through RKLLM 1.2.1 using NPU in 50 state-capitals prompt mode on Indiedroid Nova.

Measured Nova NPU result.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Llama 3.1 8B Instruct
Q4 reference
llama.cpp
CPU · Jeff Geerling published result
1.99 tok/sn/an/aExternal referencePi 5 references
Open result details

Why this model is here

The Nova NPU measured 3.72 tok/s. A separate 1.99 tok/s Pi 5 value remains labeled as an outside reference.

It connects Trevor-measured Nova data with a clearly separated published Pi reference without presenting the two methods as a matched hardware contest.

What ran

Q4 reference ran through llama.cpp using CPU in Jeff Geerling published result mode on Raspberry Pi 5.

Cited result. Not Trevor-measured.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing Llama 3.1 8B Instruct

2026-04-10

Indiedroid Nova RKLLM NPU Test

The three Nova rows share the local test. Raspberry Pi numbers in the article came from outside references with different runtime and quantization.

Open full test

2026-04-10

Raspberry Pi 5 Published Reference Rows

These are external references with different runtime, quantization, and hardware approach. A matched Trevor-run Pi 5 test has not been completed.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"