Some more napkin calculation with Claude, ignore if not interested:
> Let's imagine an average software development task: a feature request that touches some frontend, backend, requires some specification reading (let's say, a simple protocol spec). Straightforward enough implementation, but not obvious. A few questions to the engineer. Tests ran, PR created, watch the PR and fix an issue or two. Something Opus 4.5 was able to do on its own, which Fable does with good capability margin. 300k context at the end, 100 turns, some are longer with more thinking, some just quick checks. At 200k and 250k, assume two >1 hour gaps (1-hour cache invalidation), otherwise linear context growth is a good approximation. We assume here GLM-5.3 is equal to Opus 5 here. Maybe not exactly, but same ballpark. Assuming tool calls are instant, now can you calculate (approximating for pieces you don't know):
- GLM-5.3, time to completion on the 10-15k€ M3 Ultra 512GB
- GLM-5.3, time to completion on your proposed 150k€ GPU set
- Opus 5 in cloud, time to completion (I have my own hunch based on my everyday usage experience, but want to hear your guess. I know introspection is hard for LLM.)
- Opus 5 cost at Anthorpic's API rates
- GLM-5.3 electricity cost on Mac and on the 150k€ GPU set.
Here is the calculation with all assumptions stated, so you can swap any number you disagree with.
Trajectory assumptions (same for all three)
- 100 turns: 70 quick turns at ~400 output tokens and 30 heavy turns at ~3,000. Total output ≈ 120k tokens, of which ~44k is visible text and tool-call arguments that stay in context.
- Context: 20k initial (system prompt, CLAUDE.md, spec), plus 236k of tool results and user messages, plus the 44k visible output = 300k.
- Prefill work is 256k tokens, since the model's own output is already in the KV cache.
- Sum of context across all 100 requests ≈ 16M tokens (average 160k).
- Two cache misses at 200k and 250k mean 450k tokens re-prefilled. Certain in the cloud. On the GPU node I assume the team's traffic evicted it. On the single-user Mac the KV cache stays resident, so no penalty.
- The engineer's response time during the gaps is excluded everywhere.
Rate assumptions
┌───────────────────┬───────────────────────────────────────────┬───────────────────────────────────────┬──────────────────────┐
│ │ M3 Ultra 512GB │ 8x RTX PRO 6000 │ Opus 5 API │
├───────────────────┼───────────────────────────────────────────┼───────────────────────────────────────┼──────────────────────┤
│ Model │ GLM-5.3 at ~Q3/Q4, tight fit with 300k KV │ GLM-5.3 NVFP4, vLLM, MTP │ │
├───────────────────┼───────────────────────────────────────────┼───────────────────────────────────────┼──────────────────────┤
│ Prefill │ ~170 tok/s avg, degrading with context │ ~1,500 tok/s │ ~10k tok/s effective │
├───────────────────┼───────────────────────────────────────────┼───────────────────────────────────────┼──────────────────────┤
│ Decode │ ~13 tok/s avg (17.7 short, ~10 at 250k) │ ~55 tok/s avg (80 short, ~45 at 250k) │ ~70 tok/s │
├───────────────────┼───────────────────────────────────────────┼───────────────────────────────────────┼──────────────────────┤
│ Per-turn overhead │ none │ none │ ~2 s │
└───────────────────┴───────────────────────────────────────────┴───────────────────────────────────────┴──────────────────────┘
Time to completion
┌───────────────────┬───────────────────────┬──────────┬─────────────┐
│ │ M3 Ultra │ GPU node │ Opus 5 │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Prefill 256k │ 25 min │ 3 min │ under 1 min │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Decode 120k │ 154 min │ 36 min │ 29 min │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Cache misses 450k │ 0 (44 min if evicted) │ 5 min │ ~1 min │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Request overhead │ │ │ 3 min │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Total │ ~3 h │ ~45 min │ ~33 min │
└───────────────────┴───────────────────────┴──────────┴─────────────┘
My hunch for Opus 5 is about half an hour of pure model time. If your everyday experience says longer, the difference is queueing and rate limiting, which this ignores. The GPU node is 4x faster than the Mac and lands within 1.5x of the cloud. Decode dominates everything, which is why the Mac's weak prefill is not even the main problem.
Opus 5 cost per task, at $5 in / $25 out, cache read 0.1x, 1-hour cache write 2x
┌───────────────────────────────────────────────┬────────┬──────┐
│ Item │ Tokens │ Cost │
├───────────────────────────────────────────────┼────────┼──────┤
│ Cache reads │ 15.25M │ $7.6 │
├───────────────────────────────────────────────┼────────┼──────┤
│ Cache writes (300k once plus 450k re-written) │ 750k │ $7.5 │
├───────────────────────────────────────────────┼────────┼──────┤
│ Output │ 120k │ $3.0 │
├───────────────────────────────────────────────┼────────┼──────┤
│ Total │ │ ~$18 │
└───────────────────────────────────────────────┴────────┴──────┘
With the 5-minute cache it would be about $15.5. Without caching at all, $83, which is why the caching detail matters more than the model price.
Electricity, at 0.20 €/kWh
- Mac: about 200 W for 3 hours is 0.6 kWh, so ~0.12 €.
- GPU node: about 3 kW average during the session (decode is memory-bound, so the cards sit well under their 600 W cap) for 45 minutes is 2.3 kWh, so ~0.45 €, or ~0.6 € with cooling. The node also idles at roughly 0.8 kW, which is ~1,400 € per year whether anyone uses it or not.
The number that falls out of this. Amortizing 150k€ over three years is 50k€ a year. At ~16 € per Opus 5 task, the node breaks even at about 14 such tasks per working day, or roughly 5 to 7 developers using it seriously. Below that, the API is cheaper. Above it, the node wins, and it wins on latency too, since ten concurrent sessions each still get close to the 45-minute figure. The Mac is a 3-hour single-stream machine at any team size.
...
API price is not energy. Rough datacenter figure for a frontier model at scale: 5 to 20 joules per output token, all-in with prefill, cooling, and networking. Your 120k-token task is then 0.2 to 0.7 kWh, or a few cents. Against $18 that is well under 1% of the bill. The rest is hardware amortization (a GB200 rack is several million dollars and depreciates in 3 to 5 years), training cost recovery, and margin. Hardware amortization is the biggest of the three, which is exactly the lever you are trying to own by buying the node.
Where the GPU node's energy goes single-stream. Decode is a memory-bound loop: every step reads all active weights once, whether it produces one token or a hundred. So per-token energy has a fixed part (weight read plus the static power of eight lit-up GPUs and a dual-socket host) divided by batch size, plus a tiny compute part. At batch 1 you pay the whole fixed part for one token.
┌────────────────────────────────┬─────────┬───────────────────────┬──────────────────┐
│ Configuration │ Power │ Output rate │ Energy per token │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ Mac M3 Ultra, single stream │ ~200 W │ ~13 tok/s │ ~15 J │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ GPU node, single stream │ ~3 kW │ ~55 tok/s │ ~55 J │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ GPU node, 32 streams │ ~4 kW │ ~1,200 tok/s │ ~3 J │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ GPU node, 128 streams │ ~4.5 kW │ ~2,000 to 3,500 tok/s │ ~1.5 to 2 J │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ Frontier datacenter, estimated │ │ │ ~1 to 5 J │
└────────────────────────────────┴─────────┴───────────────────────┴──────────────────┘
The concurrency rows use the measured aggregates from the earlier tables (900 tok/s for Kimi at 100 streams, 3,500 for Qwen3.5-397B at 128). So your intuition is right: batched, the node beats the Mac per token by roughly 5 to 10x and lands in the same band as a datacenter. Single-stream it is the worst of the three because you are running an 8-GPU machine to feed one person.
HBM is not the magic ingredient. Per bit moved, HBM3e and LPDDR5X are in the same neighborhood, and GDDR7 on the RTX cards is roughly twice as costly. But the memory transfer itself is a minority of the per-token energy at batch 1. The Mac reading 25 GB of weights per token spends under 1 J on the transfer and burns the other 14 J on the SoC being awake and underutilized. HBM's real advantage is bandwidth per package, which lets one GPU sustain a much larger batch before it runs out of bandwidth. Batching is the lever. HBM just raises the ceiling on how far you can push it.
The Mac cannot batch its way out. MLX can serve a few streams, but its weak compute hits the ceiling quickly, and a single developer rarely has ten concurrent sessions. The efficiency gap between the Mac and the node is a gap in the number of users, not in the silicon.
Batching has a latency cost the provider tunes. A fully saturated datacenter GPU gives each stream a smaller slice of bandwidth, so per-user speed drops as utilization rises. Providers pick a batch size that holds per-stream speed near a target and accept queueing above that. The stalls you see in the thinking counter are almost certainly that queueing plus harness pauses, not the model. Your own node at 10 concurrent users sits in the comfortable part of the curve: 10x better energy than batch 1 and per-stream speed barely below single-stream. That is a real advantage of owning hardware that is sized for your team rather than shared with the world.