/hub/models/phi-4-mini

Phi-4 · Mini

Phi-4 Mini

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

4
result records
3
Trevor-measured rows
3
test groups
1
devices represented

Why Phi-4 Mini is in The Hub

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

The rows on this page use Q4_K_M through llama.cpp 9a286ac and Ollama 0.32.1. They describe those exact artifacts and runs, not every version that shares the model name.

What the recorded rows show

The fastest retained generation row is 2.072 tok/s through Ollama 0.32.1 using CPU in Matched server request mode. Ollama reported 32.00% higher internal generation throughput than the paired llama.cpp row.

1 retained row is a failure: The run window contained an i915 GPU hang even though llama-bench returned zero. The kernel event overrides the process exit code.

Compare inside the full test

The rows below can come from different prompts and methods. Open the full test before treating two numbers as a direct comparison.

Recorded results

Model and artifactRuntime and modeGenerationPromptPeakStateFull test
Phi-4 Mini
Q4_K_M
llama.cpp 9a286ac
CPU · Matched server request
1.57 tok/s2.115 tok/s82 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through llama.cpp 9a286ac using CPU in Matched server request mode on Youyeetoo X1S.

Three measured requests completed with a clean kernel window.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
70 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.000531
Thermal ceiling
85 °C

Artifact SHA-2563c168af1dea0a414299c7d9077e100ac763370e5a98b3c53801a958a47f0a5db

Phi-4 Mini
Q4_K_M
Ollama 0.32.1
CPU · Matched server request
2.072 tok/s3.131 tok/s82 °CPublished resultMatched CPU runtimes
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Matched server request mode on Youyeetoo X1S.

Ollama reported 32.00% higher internal generation throughput than the paired llama.cpp row.

Recorded setup

Threads
4
Context
4,096 tokens
Prompt
70 tokens
Output
96 tokens
Batch
512 / 512 microbatch
Measured runs
3
Generation stddev
0.009724
Thermal ceiling
85 °C

Artifact SHA-2563c168af1dea0a414299c7d9077e100ac763370e5a98b3c53801a958a47f0a5db

Phi-4 Mini
Q4_K_M
llama.cpp 9a286ac
Vulkan · Full Vulkan
n/an/an/aFailed runCPU and Vulkan modes
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through llama.cpp 9a286ac using Vulkan in Full Vulkan mode on Youyeetoo X1S.

The run window contained an i915 GPU hang even though llama-bench returned zero. The kernel event overrides the process exit code.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Phi-4 Mini
Q4_K_M
Ollama 0.32.1
CPU · Mean warm generation
1.991 tok/sn/a77 °CPublished resultOriginal CPU matrix
Open result details

Why this model is here

Phi-4 Mini completed both matched CPU servers. Its full-Vulkan window coincided with an i915 GPU hang.

It provides a second model family inside the matched runtime result and a retained failure outside the clean Vulkan pair.

What ran

Q4_K_M ran through Ollama 0.32.1 using CPU in Mean warm generation mode on Youyeetoo X1S.

Patient local batch use.

Recorded setup

The full test page contains the shared controls and comparison boundary for this row.

Full tests containing Phi-4 Mini

2026-08-31

Matched Ollama and llama.cpp CPU Test

Internal generation rates are comparable inside this eight-row test. Model load time was not an identical part of the timed path.

Open full test

2026-08-31

True CPU, Mixed Host Operations, and Full Vulkan

Compare the clean Qwen rows inside this group. The larger-model entries are failure outcomes, not performance rows.

Open full test

2026-08-23

Original X1S CPU-Only Model Matrix

These warm Ollama results belong together. They should not be used as the other half of the later llama-bench runtime comparison.

Open full test

© 2026 Trevor Unland. All rights reserved.

RSS Feed

$ echo "Built with React + TypeScript"