/hub/models/ling-mini-2

Ling-mini · 16B total / 3B active

Ling-mini-2.0

The exact pinned IQ4_XS file completed true-CPU tests at three and four threads.

2
result records
2
Trevor-measured rows
1
test groups
1
devices represented

Why Ling-mini-2.0 is in The Hub

The exact pinned IQ4_XS file completed true-CPU tests at three and four threads.

The thread test shows a real speed and temperature tradeoff on the N5095 using one verified 8.8 GB artifact rather than a mutable model tag.

The rows on this page use IQ4_XS, pinned revision 8be84a0 through llama.cpp 9a286ac. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 4.083 tok/s through llama.cpp 9a286ac using CPU in 4 threads, 90 °C guard mode. The selected four-thread row improved generation 32.1% and completed under the disclosed 90 °C ceiling.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Ling-mini-2.0
IQ4_XS, pinned revision 8be84a0
llama.cpp 9a286ac
CPU · 3 threads, 85 °C guard
3.091 tok/s4.375 tok/s74 °CPublished resultLing-mini thread test
Open result details

Why this model is here

The exact pinned IQ4_XS file completed true-CPU tests at three and four threads.

The thread test shows a real speed and temperature tradeoff on the N5095 using one verified 8.8 GB artifact rather than a mutable model tag.

What ran

IQ4_XS, pinned revision 8be84a0 ran through llama.cpp 9a286ac using CPU in 3 threads, 85 °C guard mode on Youyeetoo X1S.

The clean three-thread row stayed well below the normal thermal guard.

Recorded setup

Threads
3
Prompt
128 tokens
Output
96 tokens
Measured runs
5
Generation stddev
0.001523
Thermal ceiling
85 °C
Source revision
8be84a0f4727

Artifact SHA-256a72d86d4cb4fedd940e34c08d008bb5cda42db80ce5c6bc5f9494e854a3d742d

Ling-mini-2.0
IQ4_XS, pinned revision 8be84a0
llama.cpp 9a286ac
CPU · 4 threads, 90 °C guard
4.083 tok/s5.817 tok/s85 °CPublished resultLing-mini thread test
Open result details

Why this model is here

The exact pinned IQ4_XS file completed true-CPU tests at three and four threads.

The thread test shows a real speed and temperature tradeoff on the N5095 using one verified 8.8 GB artifact rather than a mutable model tag.

What ran

IQ4_XS, pinned revision 8be84a0 ran through llama.cpp 9a286ac using CPU in 4 threads, 90 °C guard mode on Youyeetoo X1S.

The selected four-thread row improved generation 32.1% and completed under the disclosed 90 °C ceiling.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Measured runs
5
Generation stddev
0.009518
Thermal ceiling
90 °C
Source revision
8be84a0f4727

Artifact SHA-256a72d86d4cb4fedd940e34c08d008bb5cda42db80ce5c6bc5f9494e854a3d742d

Full tests containing Ling-mini-2.0

2026-08-31

Ling-mini-2.0 IQ4_XS CPU Thread Test

The four-thread row used a disclosed 90 °C abort after the first attempt reached the normal 85 °C guard.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"