Why BitCPM-CANN 1B is in The Hub
The exact official TQ2_0 GGUF ran in pinned llama.cpp and failed to load in Ollama 0.32.1.
This is a runtime-support result, not a speed contest. The same verified file produced inference in one runtime and a tensor-size overflow in the other.
The rows on this page use Official TQ2_0 GGUF through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.