Verified Mac LLM benchmarks
Tokens per second on Qwen3-8B Q4_K_M, measured on the standard llama-bench harness. Every row was signed by the Secure Enclave of the Mac that produced it, so a number here is attached to a specific machine and a specific set of conditions rather than to a claim.
The table fills as sellers publish reports. Runs on battery, in low power mode, or on a hot machine are excluded.
No measured rows yet. Nothing estimated will ever appear here, so this section stays empty until real machines have reported. The methodology below is fixed and is what those rows will be measured against.
What a number here means, exactly.
One pinned harness
Qwen3-8B Q4_K_M on llama.cpp, all layers on the GPU via Metal, flash attention on, 5 repetitions. The engine build is pinned and recorded on every row, because the ruler itself drifts: the same Mac can read three times the prompt-processing speed on a newer build with generation almost unchanged.
Two numbers, never one
Prompt processing is prefill: compute-bound, scaling with GPU cores. Token generation is decode: bound by memory bandwidth, which is why Ultra chips lead it and why it is the more useful number for chat-shaped work. Averaging them together would produce a figure that means nothing.
Conditions recorded
Power source, low power mode, and thermal state at the start and end of each run are stored with the measurement. A Mac on battery has its GPU power capped, so those runs still appear on their own report but are kept out of this table.
What this table is not
It is not a claim about your Mac's maximum speed. A different runtime on the same machine may well be faster, and on the newest silicon the gap between engines is structural rather than incidental. Every row says which engine and which build produced it, and the honest reading is always “on this harness”. There is no Macfax score and no invented metric: the number is llama-bench’s number.
What exactly is measured?
llama-bench's own pp512 and tg128 on Qwen3-8B Q4_K_M, 5 repetitions, flash attention on, all layers on the GPU. That is the community's own ruler, adopted verbatim rather than reinvented, so a Macfax row can be quoted in a reference thread without translation.
Why are prompt and generation separate numbers?
They measure different limits. Prompt processing is compute-bound and scales with GPU cores; token generation is memory-bandwidth-bound, which is why Ultra chips lead it. A single blended tokens-per-second figure is undefined, so this table never emits one.
How is this different from the tables people post already?
Every row here was produced by a Mac that signed it with its own Secure Enclave key, on one pinned engine build and one pinned model file, with power source and thermal state recorded. Self-reported rows carry none of that, which is why they scatter. Rows measured on battery or in low power mode are excluded from this table.
Can I get my Mac on this table?
Run a free Macfax report. The benchmark runs as part of it, and the row lands on your report and here. There is no paid tier for the number.
Does a faster number mean a better Mac?
For running local models, largely yes, and memory is usually the binding constraint before speed is: a model has to fit before it can run. What this table cannot tell you is whether one particular unit is healthy. That is what the report the number comes from is for.
Put your Mac on the table.
The benchmark runs as part of a free Macfax report. Your number lands on the report and here, signed by your own machine.