Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.
9
result records
9
Trevor-measured rows
2
test groups
1
devices represented
Why Gemma 4 E2B is in The Hub
Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.
It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.
The rows on this page use Q4_K_M through Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.
What the recorded rows show
The fastest retained generation row is 3.397 tok/s through Ollama 0.32.1 using CPU in Depth 4, MTP off mode. Paired MTP-off baseline for draft depth 4.
Compare inside the full test
The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.
Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.
It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.
What ran
Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.
MTP was active and ran 41.63% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.
Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.
It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.
What ran
Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.
MTP was active and ran 56.03% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.