/hub/models/gemma-4-e2b

Gemma 4 · E2B

Gemma 4 E2B

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

9
result records
9
Trevor-measured rows
2
test groups
1
devices represented

Why Gemma 4 E2B is in The Hub

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

The rows on this page use Q4_K_M through Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 3.397 tok/s through Ollama 0.32.1 using CPU in Depth 4, MTP off mode. Paired MTP-off baseline for draft depth 4.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 1, MTP off
3.269 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 1, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 1.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 1, MTP on
2.266 tok/sn/a84 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 1, MTP on mode on Youyeetoo X1S.

MTP was active and ran 30.68% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 2, MTP off
3.359 tok/sn/a80 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 2, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 2. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 2, MTP on
1.96 tok/sn/a84 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 2, MTP on mode on Youyeetoo X1S.

MTP was active and ran 41.63% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 3, MTP off
3.352 tok/sn/a82 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 3, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 3. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 3, MTP on
1.474 tok/sn/a86 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 3, MTP on mode on Youyeetoo X1S.

MTP was active and ran 56.03% slower than its paired baseline. This pair includes a 90 °C rerun after the first attempt stopped at the normal 85 °C guard.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
1
Thermal ceiling
90 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Depth 4, MTP off
3.397 tok/sn/a81 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Depth 4, MTP off mode on Youyeetoo X1S.

Paired MTP-off baseline for draft depth 4.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU + draft MTP · Depth 4, MTP on
1.333 tok/sn/a84 °CPublished resultMTP depth sweep
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU + draft MTP in Depth 4, MTP on mode on Youyeetoo X1S.

MTP was active and ran 60.75% slower than its paired baseline.

Recorded setup

Threads
4
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Thermal ceiling
85 °C
Gemma 4 E2B
Q4_K_M
Ollama 0.32.1
CPU · Two-run warm average
3.38 tok/sn/a77 °CPublished resultAdditional models
Open result details

Why this model is here

Gemma 4 E2B completed the normal Ollama run and a proven MTP sweep with its separate assistant model.

It tested the MTP path that required an external assistant instead of assuming the presence of model tensors meant speculative decoding was active.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Two-run warm average mode on Youyeetoo X1S.

The normal non-MTP Ollama result. The separate MTP sweep is recorded in its own test group.

Recorded setup

Context
4,096 tokens
Measured runs
2
Thermal ceiling
85 °C

Full tests containing Gemma 4 E2B

2026-08-31

Ollama MTP Depth Sweep on the N5095

Each on/off pair belongs to its own requested draft depth. Do not compare the baseline from one pair as if it were the same request block as another.

Open full test

2026-08-31

Additional Model Results on the X1S

The rows use different repetition counts and artifact states. They are useful operating points, not one controlled model-quality contest.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"