Local AI on Apple silicon

Verified Mac LLM benchmarks

Tokens per second on Qwen3-8B Q4_K_M, measured on the standard llama-bench harness. Every row was signed by the Secure Enclave of the Mac that produced it, so a number here is attached to a specific machine and a specific set of conditions rather than to a claim.

11 measured runs across 7 configurations. Runs on battery, in low power mode, on a hot machine, or on one that was busy with something else are excluded.

The table
ConfigurationGenerationPromptRunsRange · gen
Apple M3 Ultra · 80-core GPU · 512 GB95.8 tok/s1524.8 tok/s1
Apple M5 Max · 40-core GPU · 64 GB93.2 tok/s2777.4 tok/s393.1 – 98.5
Apple M4 Max · 40-core GPU · 128 GB83.9 tok/s891.1 tok/s1
Apple M3 Max · 40-core GPU · 48 GB65.8 tok/s789.2 tok/s1
Apple M2 Max · 38-core GPU · 64 GB61.0 tok/s638.6 tok/s1
Apple M1 Pro · 16-core GPU · 32 GB25.4 tok/s253.4 tok/s1
Apple M4 · 10-core GPU · 16 – 32 GB20.0 tok/s226.9 tok/s319.7 – 20.5

Median of measured runs per configuration. Peak throughput is set by memory bandwidth and GPU core count rather than by an individual Mac, and on Apple silicon the core count identifies the bandwidth bin exactly, so rows are grouped on chip and GPU cores. Memory capacity decides which models fit, not how fast they run, so it is shown as the range the measured machines had rather than used to split the row. A single-run row is one machine, not a consensus.

Methodology

What a number here means, exactly.

One pinned harness

Qwen3-8B Q4_K_M on llama.cpp, all layers on the GPU via Metal, flash attention on, 5 repetitions. The engine build is pinned and recorded on every row, because the ruler itself drifts: the same Mac can read three times the prompt-processing speed on a newer build with generation almost unchanged.

Two numbers, never one

Prompt processing is prefill: compute-bound, scaling with GPU cores. Token generation is decode: bound by memory bandwidth, which is why Ultra chips lead it and why it is the more useful number for chat-shaped work. Averaging them together would produce a figure that means nothing.

Conditions recorded

Power source, low power mode, thermal state at both ends of the run, and how much of the machine the benchmark actually had to itself are all stored with the measurement. Runs made on battery, in low power mode, on a hot Mac, or on one busy with something else still appear on their own report, and are kept out of this table so that what is averaged here is comparable.

What this table is not

It is not a claim about your Mac's maximum speed. A different runtime on the same machine may well be faster, and on the newest silicon the gap between engines is structural rather than incidental. Every row says which engine and which build produced it, and the honest reading is always “on this harness”. There is no Macfax score and no invented metric: the number is llama-bench’s number.

Questions

What exactly is measured?

llama-bench's own pp512 and tg128 on Qwen3-8B Q4_K_M, 5 repetitions, flash attention on, all layers on the GPU. That is the community's own ruler, adopted verbatim rather than reinvented, so a Macfax row can be quoted in a reference thread without translation.

Why are prompt and generation separate numbers?

They measure different limits. Prompt processing is compute-bound and scales with GPU cores; token generation is memory-bandwidth-bound, which is why Ultra chips lead it. A single blended tokens-per-second figure is undefined, so this table never emits one.

How is this different from the tables people post already?

Every row here was produced by a Mac that signed it with its own Secure Enclave key, on one pinned engine build and one pinned model file, with power source and thermal state recorded. Self-reported rows carry none of that, which is why they scatter. Rows measured on battery or in low power mode are excluded from this table.

Can I get my Mac on this table?

Run a free Macfax report. The benchmark runs as part of it, and the row lands on your report and here. There is no paid tier for the number.

Does a faster number mean a better Mac?

For running local models, largely yes, and memory is usually the binding constraint before speed is: a model has to fit before it can run. What this table cannot tell you is whether one particular unit is healthy. That is what the report the number comes from is for.

Put your Mac on the table.

The benchmark runs as part of a free Macfax report. Your number lands on the report and here, signed by your own machine.