Why Qwen3 8B is in The Hub
Qwen3 8B fit during the original CPU matrix at 0.924 tok/s. A later full-Vulkan attempt coincided with an i915 GPU hang.
It establishes that fitting in 16 GB and completing on CPU did not guarantee a stable full-GPU path.
The rows on this page use Q4_K_M through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.