/hub/models/bitcpm-cann-1b

BitCPM · 1B

BitCPM-CANN 1B

The exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1.

2
result records
1
Trevor-measured rows
1
test groups
1
devices represented

Why BitCPM-CANN 1B is in The Hub

The exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1.

This is a runtime-support result, not a speed contest. The same verified file produced inference in one runtime and a tensor-size overflow in the other.

The rows on this page use Official TQ2_0 GGUF through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 8.584 tok/s through llama.cpp 9a286ac using CPU in True CPU, five repetitions mode. Pinned llama.cpp accepted the verified file and completed the controlled true-CPU workload.

1 retained row is a failure: Ollama rejected the same verified file with a tensor-size overflow before inference.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
BitCPM-CANN 1B
Official TQ2_0 GGUF
llama.cpp 9a286ac
CPU · True CPU, five repetitions
8.584 tok/s13.01 tok/s76 °CPublished resultRuntime compatibility
Open result details

Why this model is here

The exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1.

This is a runtime-support result, not a speed contest. The same verified file produced inference in one runtime and a tensor-size overflow in the other.

What ran

Official TQ2_0 GGUF ran through llama.cpp 9a286ac using CPU in True CPU, five repetitions mode on Youyeetoo X1S.

Pinned llama.cpp accepted the verified file and completed the controlled true-CPU workload.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Generation stddev
0.029317
Thermal ceiling
85 °C

Artifact SHA-2562394c15cbea2181b72bfb4215d8417d8d1f2f6214069da2d01fde32ce3b13fce

BitCPM-CANN 1B
Official TQ2_0 GGUF
Ollama 0.32.1
CPU · Model load
n/an/an/aFailed runRuntime compatibility
Open result details

Why this model is here

The exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1.

This is a runtime-support result, not a speed contest. The same verified file produced inference in one runtime and a tensor-size overflow in the other.

What ran

Official TQ2_0 GGUF ran through Ollama 0.32.1 using CPU in Model load mode on Youyeetoo X1S.

Ollama rejected the same verified file with a tensor-size overflow before inference.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Artifact SHA-2562394c15cbea2181b72bfb4215d8417d8d1f2f6214069da2d01fde32ce3b13fce

Full tests containing BitCPM-CANN 1B

2026-08-31

BitCPM and Gemma Runtime Compatibility

BitCPM is not a matched runtime speed contest because Ollama never loaded the model. The Gemma derivative is inspectable but does not repair the Vulkan device-loss path.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"