Why DeepSeek-R1 Distill Qwen 1.5B is in The Hub
The Nova NPU run started at 11.5 tok/s, then the state-capitals answer degraded into invented cities.
It is the clearest example in the Hub of why throughput and useful output cannot be treated as the same result.
The rows on this page use W8A8 RKLLM and Q4 reference through RKLLM 1.2.1 and llama.cpp. They describe those exact artifacts and runs, not every version that shares the model name.