/hub/models/gemma-3-4b

Gemma 3 · 4B

Gemma 3 4B

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

3
result records
2
Trevor-measured rows
3
test groups
1
devices represented

Why Gemma 3 4B is in The Hub

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

The rows on this page use Text-only derivative, Separate text-only derivative, and Q4_K_M through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 1.995 tok/s through Ollama 0.32.1 using CPU in Mean warm generation mode. Patient local batch use.

1 retained row is a failure: Fence and preemption timeouts led to an i915 reset and vk::DeviceLostError.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Gemma 3 4B
Text-only derivative
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed runCPU and Vulkan modes
Open result details

Why this model is here

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

What ran

Text-only derivative ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

Fence and preemption timeouts led to an i915 reset and vk::DeviceLostError.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Gemma 3 4B
Separate text-only derivative
llama.cpp 9a286ac
CPU · True CPU, five repetitions
1.624 tok/s2.172 tok/s81 °CPublished resultRuntime compatibility
Open result details

Why this model is here

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

What ran

Separate text-only derivative ran through llama.cpp 9a286ac using CPU in True CPU, five repetitions mode on Youyeetoo X1S.

The derived copy completed true CPU. The original multimodal source artifact was not overwritten.

Recorded setup

Threads
4
Prompt
128 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
5
Generation stddev
0.003291
Thermal ceiling
85 °C

Artifact SHA-256510408e8043ca1c741fe9a16088d47e8fa0d016033c6acf0c50c50c7c93b6530

Gemma 3 4B
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.995 tok/sn/a77 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

A separate text-only derivative completed true CPU. The later Vulkan path ended with timeouts, an i915 reset, and device loss.

It records two different compatibility boundaries: converting the older multimodal artifact for current llama.cpp and the separate Vulkan failure that conversion did not fix.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing Gemma 3 4B

2026-08-31

True CPU, Mixed Host Operations, and Full Vulkan

Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.

Open full test

2026-08-31

BitCPM and Gemma Runtime Compatibility

BitCPM is not a matched runtime speed contest because Ollama never loaded the model. The Gemma derivative is inspectable but does not repair the Vulkan device-loss path.

Open full test

2026-08-23

Original X1S CPU-Only Model Matrix

These warm Ollama results belong together. They should not be used as the other half of the later llama-bench runtime comparison.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"