/hub/compare

Compare two local AI results

Result A is the starting point. Result B is the runtime, backend, thread count, or MTP setting you want to check against it. The presets below only pair rows that belong together.

Choose what you want to compare

Each option is a prepared pair from one test group. Result A is the baseline. Result B is the setting being checked against it.

Published resultYouyeetoo X1S · 2026-08-31

Matched Ollama and llama.cpp CPU Test

Four exact Q4_K_M GGUFs received the same raw prompt, context, thread count, batch settings, sampler, and 96-token output through Ollama 0.32.1 and native llama.cpp commit 9a286ac.

Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.

Result A

Published resultQ4_K_M

Qwen3 0.6B

llama.cpp 9a286ac · CPU · Matched server request

Generation
6.725 tok/s
Prompt
10.48 tok/s
Peak
81 °C

Three measured requests completed with a clean kernel window.

Open Qwen3 0.6B results

Result B

Published resultQ4_K_M

Qwen3 0.6B

Ollama 0.32.1 · CPU · Matched server request

Generation
7.809 tok/s
Prompt
13.24 tok/s
Peak
81 °C

Ollama reported 16.11% higher internal generation throughput than the paired llama.cpp row.

Open Qwen3 0.6B results

What changed

Result B generated 16.1% faster than Result A in this test group. The cards below keep the artifact, runtime, backend, mode, temperature, and outcome next to that difference.

Metric
Result A
Result B
Difference
Generation
6.725 tok/s
7.809 tok/s
+16.1% B vs A
Prompt processing
10.48 tok/s
13.24 tok/s
+26.3% B vs A
Peak package
81 °C
81 °C
+0 °C B vs A
Open the full test, controls, and files

Tests with valid comparison pairs

Matched Ollama and llama.cpp CPU Test

4 prepared comparisons. Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.

Open full test

True CPU, Mixed Host Operations, and Full Vulkan

2 prepared comparisons. Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.

Open full test

Ling-mini-2.0 IQ4_XS CPU Thread Test

1 prepared comparisons. The four-thread row used a disclosed 90 °C abort after the first attempt reached the normal 85 °C guard.

Open full test

Ollama MTP Depth Sweep on the N5095

12 prepared comparisons. Each on/off pair belongs to its own requested draft depth. Do not compare the baseline from one pair as if it were the same request block as another.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"