Author Topic: Local AI and hardware to run it  (Read 4860 times)

SpacedCowboy and 9 Guests are viewing this topic.

Offline tom66Topic starter

  • Super Contributor
  • ***
  • Posts: 8840
  • Country: gb
  • Professional HW / FPGA / Embedded Engr. & Hobbyist
Local AI and hardware to run it
« on: September 08, 2026, 10:47:18 am »
I'm spec'ing out a system for developers at the company to run local AI models.  I'd like to run large models, above the standard 27B/31B, so need large unified memory.  The system ideally needs to be able to host at least four users concurrently with acceptable performance of 15-20 tok/s, as well as support code completion.

Budget is not unlimited.  But it's not a consumer application so if the hardware meets the requirements and the cost can be justified, then we can buy it.

So far it seems possibly the best option is the Mac Pro M5 with either 96GB or 256GB of RAM. This will cost between £6-10k.  I'm not clear if it's worth going for the 64-core or 80-core GPU.  From what I can see the Mac Pro has 1200GB/s of memory bandwidth so at 32GB model size (after quants) I'm estimating 37 tok/s for one user as I imagine memory bandwidth will be the main limit here.  I'd assume scaling for parallel ops is non linear, as long as llama.cpp can keep the model reads in sync with each other, but I'm not sure.  Having 256GB of RAM would allow for both an expert model and a lighter model to be loaded concurrently.

I did consider building out a system with two RTX 5090's in it, but the cards have increased considerably in cost.  I'd have to spend £9k on the cards, and still build a system that could support them.  And that would limit me to models that split well across cards - not something all models are designed for.

The awkward thing about the Mac Pro is that it's got to run Mac OS, which will be an issue from an IT management perspective, given they're set up for Windows only.  It would be nice to run Linux on it, but the information I can find online suggests there isn't good support for the Apple NPUs on Linux yet (and this might not ever happen given Apple's historic resistance towards Linux support.)

Welcome everyone's thoughts here.
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6464
  • Country: nz
Re: Local AI and hardware to run it
« Reply #1 on: September 08, 2026, 10:59:58 am »
So far it seems possibly the best option is the Mac Pro M5 with either 96GB or 256GB of RAM.

There is no such computer.

The Mac Pro line was discontinued in March, and never used anything newer than the M2 Ultra.
 

Offline tom66Topic starter

  • Super Contributor
  • ***
  • Posts: 8840
  • Country: gb
  • Professional HW / FPGA / Embedded Engr. & Hobbyist
Re: Local AI and hardware to run it
« Reply #2 on: September 08, 2026, 11:17:13 am »
So far it seems possibly the best option is the Mac Pro M5 with either 96GB or 256GB of RAM.

There is no such computer.

The Mac Pro line was discontinued in March, and never used anything newer than the M2 Ultra.

Sorry, mean the Mac Studio!
https://www.apple.com/uk/shop/buy-mac/mac-studio
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6464
  • Country: nz
Re: Local AI and hardware to run it
« Reply #3 on: September 08, 2026, 11:50:44 am »
Sorry, mean the Mac Studio!
https://www.apple.com/uk/shop/buy-mac/mac-studio

Ideal is the Mac Studio with 512GB of unified RAM.

In early March 512GB RAM was a $4000 option on top of the base $2499 price (so total $6499) for a 36GB RAM machine (with the minimum 512GB SSD). But that was discontinued on IIRC March 8 as RAM prices and availability started to go crazy.

At the moment an M5 Ultra with 256GB RAM and 1 TB SSD is $9499, but NB "512GB memory option for M5 Ultra coming late October".

People are offering 512GB Mac Studios on eBay (and Trademe here in NZ), but at large markups from the original price.
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: Local AI and hardware to run it
« Reply #4 on: September 08, 2026, 11:52:03 am »
The whole discussion gets interesting when we get beyond the "personal programming assistant" idea to "company-wide AI cluster" - suddenly the cost range of 2k (what an individual can usually actually afford) - 10k (what small minority can afford) transitions to 10k (what almost any company, even small startups, can easily afford) - 500k (what larger companies easily could).

For those interested - Claude's take on how much it takes to get as close to frontier performance as possible:
Code: [Select]
The best open-weight models sit roughly one generation behind. On the Artificial Analysis index, GLM-5.3 and Kimi K3 score 44 against 53 for Fable 5.1. On BenchLM's coding leaderboard the gap is wider (Fable 5.1 at 84, Kimi K3 at 68, GLM-5.2 at 61). In practice that means very good interactive coding and short agentic tasks, and a noticeable drop on multi-hour autonomous work. It is far better than a 27B or 31B model, though Qwen3.6-27B at 77% SWE-bench Verified is now surprisingly strong for a single card.

The candidates, and what fits in a 100 to 200k€ box

┌───────────────────┬────────────────────────┬────────────────────────────────────────────────────────┬───────────────────────────────────────────────────────────────────────────────────────────────────────┐
│       Model       │          Size          │                Fits on 8x 96GB (768GB)?                │                                                 Notes                                                 │
├───────────────────┼────────────────────────┼────────────────────────────────────────────────────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ GLM-5.3 / GLM-5.2 │ 744B total, 40B active │ Yes at INT4/NVFP4 (~400-475GB). FP8 (~810GB) does not. │ Best agentic coder you can actually run. 5.2 is MIT; 5.3 is MIT-equivalent under 10B$ revenue.        │
├───────────────────┼────────────────────────┼────────────────────────────────────────────────────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ Kimi K2.7-Code    │ ~1T, ~32B active       │ Yes at INT4 (~640GB)                                   │ K2.5 measured at 90 tok/s single stream, ~900 tok/s aggregate at 100 concurrent on this exact config. │
├───────────────────┼────────────────────────┼────────────────────────────────────────────────────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ DeepSeek V4 Flash │ 284B, 13B active       │ Yes, easily (~170GB)                                   │ ~79% SWE-bench Verified, fast, lots of KV headroom. Best throughput per euro.                         │
├───────────────────┼────────────────────────┼────────────────────────────────────────────────────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ DeepSeek V4 Pro   │ 1.6T, 49B active       │ No (~800GB at Q4, no room for KV cache)                │ Needs 2 nodes or H200 class, 300k€+.                                                                  │
├───────────────────┼────────────────────────┼────────────────────────────────────────────────────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ Kimi K3           │ 2.8T, 104B active      │ No (1.56TB at MXFP4)                                   │ Needs ~2TB+ GPU memory.                                                                               │
├───────────────────┼────────────────────────┼────────────────────────────────────────────────────────┼───────────────────────────────────────────────────────────────────────────────────────────────────────┤
│ Qwen3.8-2.4T      │ 2.4T, 95B active       │ No (1.2TB at Q4)                                       │ Same story.                                                                                           │
└───────────────────┴────────────────────────┴────────────────────────────────────────────────────────┴───────────────────────────────────────────────────────────────────────────────────────────────────────┘

So the sweet spot is GLM-5.3 as the "smart" model plus DeepSeek V4 Flash as the fast one, on the same node.

Recommended build (about 150-180k€ incl. VAT)

- Supermicro SYS-522GA-NRT or AS-5126GS-TNRT, both validated for 8x RTX PRO 6000 Blackwell Server Edition.
- 8x RTX PRO 6000 Blackwell 96GB. Cards are ~16k$ list now, up from 8.5k$ at launch, so expect 13-17k€ each. This is ~120k€ of the total.
- 2x EPYC 9575F, 1-2TB DDR5, a few TB NVMe. The community reference rig for Kimi on this GPU count is exactly this.
- vLLM or SGLang with an OpenAI-compatible endpoint, fronted by whatever agent harness you like (OpenCode, Cline, Qwen Code). The harness matters nearly as much as the model.
- Budget 5-6kW of power and cooling. This lives in a server room, not under a desk.

Cheaper alternatives

- 4x RTX PRO 6000 in a tower (Supermicro SYS-741GE-TNRT or Exxact's validated 4x Max-Q build), around 80-100k€. Runs GLM-5.2 mixed precision at ~80 tok/s single stream with 400k context, or V4 Flash with room to spare. Fewer concurrent users.
- NVIDIA DGX Station GB300, 85-100k$. One quiet box with 288GB HBM3e plus 496GB LPDDR5X. Runs GLM-5.x at INT4, but performance drops when KV cache spills out of HBM.

Why not Macs

Four M5 Ultra Studios with 512GB (shipping late October, likely 12-14k€ each) pool 2TB over Thunderbolt 5 RDMA and could in theory hold Kimi K3. But the cluster serves essentially one stream at 25-30 tok/s. That's a superb single-developer machine and a poor company server. A GPU node batches, which is what a team needs.

One caveat on timing: RTX PRO 6000 pricing has nearly doubled in 18 months and moved 20% over the summer. Get a fixed quote before you commit to a budget line.

The 512GB Mac doesn't sound like a bad deal once it's out:
https://pinggy.io/blog/self_hosting_llms_on_512gb_m5_ultra_mac_studio/
« Last Edit: September 08, 2026, 12:08:08 pm by Siwastaja »
 

Offline tom66Topic starter

  • Super Contributor
  • ***
  • Posts: 8840
  • Country: gb
  • Professional HW / FPGA / Embedded Engr. & Hobbyist
Re: Local AI and hardware to run it
« Reply #5 on: September 08, 2026, 12:12:43 pm »
I'm surprised Claude/etc. haven't been nerfed when it comes to asking about open-weights models.  Anthropic has been passively anti-open-weightss.  They suggest that they are not safe because they can be abliterated (restrictions removed).  But obviously they're also competitive with Claude.

I think I'll probably be able to argue for somewhere up to £20k.  We're currently paying about £100/m for AI credits per engineer and have five engineers using that.  A Mac Studio runs at 450W peak, assume that's running 24/7 then that's about 20p/hour in electricity & air con in the server room plus other costs.  So on a costs basis, assuming the server never sleeps, it'll be more cost effective.  The issue I think is scaling... if we want to run different models, at the same time, the Studio is not going to work; its memory bandwidth will kill it for parallel workloads.  So we'd want to build a cluster of these.  At what point does it make more sense to build a rack with A100s in it?
 

Offline kite31

  • Frequent Contributor
  • **
  • Posts: 266
  • Country: au
Re: Local AI and hardware to run it
« Reply #6 on: September 08, 2026, 12:20:17 pm »
The system ideally needs to be able to host at least four users concurrently with acceptable performance of 15-20 tok/s, as well as support code completion.
Are you assuming a Studio per person for your 4+ users? An alternative is to use smaller machines with Thunderbolt 5 for RDMA. The gain is more for use by a group than an individual, more efficient use of the resources. It also bumps your effective available memory.

I am talking about possible technology options. I do not know your real requirements or priorities.
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: Local AI and hardware to run it
« Reply #7 on: September 08, 2026, 12:44:52 pm »
The 512GB Mac doesn't sound like a bad deal once it's out:
https://pinggy.io/blog/self_hosting_llms_on_512gb_m5_ultra_mac_studio/

Code: [Select]
Independent measurement from the oMLX benchmark database on an M3 Ultra 512GB puts it at 213.6 tok/s prefill and 17.7 tok/s generation
Are you guys sure this local AI is going to fly?

What I do every day, all the time with Claude is agentic loop which ends up solving a problem with 200k-500k tokens of context. Every turn is easily 2-5k tokens of thinking. Same for open-weight models, the value is in thinking, then tool call, then more thinking. Thinking tokens grow fast and so does turns. Even a relatively small task is at least 30 turns.

Quick ballpark running an open weight model on this Mac transitions trivially small 3-minute feature request into a multi-hour territory. It's not a question of if these models are capable enough for multi-step workloads - even if the model can handle it, the user likely won't - a 30-minute autonomous run (solving a real problem, adding a significant new feature and testing it) becomes a full day's wait. It's still somewhat faster than a human being, but eats an order of magnitude of the "AI edge".

I think people make a mistake when they think 10-30 tok/s "sounds good" because they can't read the output faster than that. But what they miss is the reasoning between every step.

It seems the 150k€ GPU build Claude suggested would be significantly faster, coming closer to the cloud speed. So this is still interesting in a company scale, where a 150k€ investment is not that big. CEO's car is already more expensive than that. But the 10-15k€ Mac will be quickly costing a lot in work not being done because it's not fast enough to run serious AI for serious agentic development/research applications.

I'm surprised Claude/etc. haven't been nerfed when it comes to asking about open-weights models.

It would be idiotic to start controlling models to support your own business. Not only it's quite hard to control LLMs that way, and if you do it, you instantly get caught and very bad publicity. The "honesty" principle seems to pay and Anthropic's customers pay for it. OpenAI's situation is more dire as their users are not as willing to pay for the service (even when it's on par with Anthropic's quality).

Quote
Anthropic has been passively anti-open-weightss.  They suggest that they are not safe because they can be abliterated (restrictions removed).

I mean, it's not "passive", they define the whole company's purpose as AI safety research company. And their concern is very well proven - we did discuss the case where an individual was able to use open-weight models to hack into and root a phone Claude declined to do. It was their phone, sure, but the bottom line is, the LLM couldn't know that, it performed the attack no-questions-asked, and would have performed it to a third party's phone. It's very concerning they do that. But oh well, this seems unavoidable path. Security has always been a race between the "bad guys" and "good guys" using the latest tools to their advantage, and a tool supplier is not a bad guy just for releasing the tool in public.
 

Offline Swake

  • Super Contributor
  • ***
  • Posts: 1081
  • Country: be
Re: Local AI and hardware to run it
« Reply #8 on: September 08, 2026, 03:07:15 pm »
Here are some real world numbers.

Config:
Ollama server with RTX5090 32G VRAM
LLM is Qwen 3.8-27B Q4, with 128K context (oversized to 150000 as compression happens at about 130000 so effective context is 128K)
This system is not specifically optimized for speed but for convenience, so it is possible to squeeze out some more context or a slightly bigger model (Q5 or Q6).

Client running the dev harness is a virtual machine with 16 cores and 64GB system memory (total overkill, 4 cores and 16GB is more than enough, till you start compiling something but that doesn't happen frequently here)

Performance graph is attached. So yes that looks nice 80 tok/s on average with caveman thinking approach (=reasoning in telegraphic fragments instead of full human style language).

How it feels in reality is this: I give it some task, wait a couple minutes for the LLM to return with some questions about the deliverables, then I answer these and approve the work and.... I wait.... and wait.... anywhere from 5 minutes to a couple hours depending on the size of the task. What should I do while I'm waiting?? at first I took another drink, walked the dog, visited the hairdresser, cooked dinner,.....  not joking here! Of course got annoyed of this situation very quickly because this project can make so much more speed if I add more 'robots'.... I can manage at least 5 robots at ones, maybe more. We'll see, it is a serious budget of course, even if I go cloud.


Now my 2 cents on your situation:
We're currently paying about £100/m for AI credits per engineer and have five engineers using that.
Something tells me that your engineers are still at their very first steps when it comes to using AI. 100 bucks is only a couple hours of consumption. Metaphorically they are at the stage of 'addition' while they could be at the stage of 'multiplication and exponents'. Using LLM's for dev requires another way of thinking. Your engineers can learn it, it takes some time but with the right setup and maybe a good trainer to learn the right method it can be accelerated. Because of the specific workflow setup here I can now predict how valuable the output will be and this made me start thrusting the LLM more and more. That made a huge jump in valuable output rate.
When it fits, stop using the hammer
 

Offline Swake

  • Super Contributor
  • ***
  • Posts: 1081
  • Country: be
Re: Local AI and hardware to run it
« Reply #9 on: September 08, 2026, 03:14:49 pm »
Something to keep in mind is that a single RTX5090 blows 575W of heat. Not counting the rest of the machine.
It will not be a cold winter in my office....
When it fits, stop using the hammer
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: Local AI and hardware to run it
« Reply #10 on: September 08, 2026, 04:07:22 pm »
Some more napkin calculation with Claude, ignore if not interested:

Code: [Select]
> Let's imagine an average software development task: a feature request that touches some frontend, backend, requires some specification reading (let's say, a simple protocol spec). Straightforward enough implementation, but not obvious. A few questions to the engineer. Tests ran, PR created, watch the PR and fix an issue or two. Something Opus 4.5 was able to do on its own, which Fable does with good capability margin. 300k context at the end, 100 turns, some are longer with more thinking, some just quick checks. At 200k and 250k, assume two >1 hour gaps (1-hour cache invalidation), otherwise linear context growth is a good approximation. We assume here GLM-5.3 is equal to Opus 5 here. Maybe not exactly, but same ballpark. Assuming tool calls are instant, now can you calculate (approximating for pieces you don't know):
  - GLM-5.3, time to completion on the 10-15k€ M3 Ultra 512GB
  - GLM-5.3, time to completion on your proposed 150k€ GPU set
  - Opus 5 in cloud, time to completion (I have my own hunch based on my everyday usage experience, but want to hear your guess. I know introspection is hard for LLM.)
  - Opus 5 cost at Anthorpic's API rates
  - GLM-5.3 electricity cost on Mac and on the 150k€ GPU set.

Here is the calculation with all assumptions stated, so you can swap any number you disagree with.

Trajectory assumptions (same for all three)

- 100 turns: 70 quick turns at ~400 output tokens and 30 heavy turns at ~3,000. Total output ≈ 120k tokens, of which ~44k is visible text and tool-call arguments that stay in context.
- Context: 20k initial (system prompt, CLAUDE.md, spec), plus 236k of tool results and user messages, plus the 44k visible output = 300k.
- Prefill work is 256k tokens, since the model's own output is already in the KV cache.
- Sum of context across all 100 requests ≈ 16M tokens (average 160k).
- Two cache misses at 200k and 250k mean 450k tokens re-prefilled. Certain in the cloud. On the GPU node I assume the team's traffic evicted it. On the single-user Mac the KV cache stays resident, so no penalty.
- The engineer's response time during the gaps is excluded everywhere.

Rate assumptions

┌───────────────────┬───────────────────────────────────────────┬───────────────────────────────────────┬──────────────────────┐
│                   │              M3 Ultra 512GB               │            8x RTX PRO 6000            │      Opus 5 API      │
├───────────────────┼───────────────────────────────────────────┼───────────────────────────────────────┼──────────────────────┤
│ Model             │ GLM-5.3 at ~Q3/Q4, tight fit with 300k KV │ GLM-5.3 NVFP4, vLLM, MTP              │                      │
├───────────────────┼───────────────────────────────────────────┼───────────────────────────────────────┼──────────────────────┤
│ Prefill           │ ~170 tok/s avg, degrading with context    │ ~1,500 tok/s                          │ ~10k tok/s effective │
├───────────────────┼───────────────────────────────────────────┼───────────────────────────────────────┼──────────────────────┤
│ Decode            │ ~13 tok/s avg (17.7 short, ~10 at 250k)   │ ~55 tok/s avg (80 short, ~45 at 250k) │ ~70 tok/s            │
├───────────────────┼───────────────────────────────────────────┼───────────────────────────────────────┼──────────────────────┤
│ Per-turn overhead │ none                                      │ none                                  │ ~2 s                 │
└───────────────────┴───────────────────────────────────────────┴───────────────────────────────────────┴──────────────────────┘

Time to completion

┌───────────────────┬───────────────────────┬──────────┬─────────────┐
│                   │       M3 Ultra        │ GPU node │   Opus 5    │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Prefill 256k      │ 25 min                │ 3 min    │ under 1 min │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Decode 120k       │ 154 min               │ 36 min   │ 29 min      │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Cache misses 450k │ 0 (44 min if evicted) │ 5 min    │ ~1 min      │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Request overhead  │                       │          │ 3 min       │
├───────────────────┼───────────────────────┼──────────┼─────────────┤
│ Total             │ ~3 h                  │ ~45 min  │ ~33 min     │
└───────────────────┴───────────────────────┴──────────┴─────────────┘

My hunch for Opus 5 is about half an hour of pure model time. If your everyday experience says longer, the difference is queueing and rate limiting, which this ignores. The GPU node is 4x faster than the Mac and lands within 1.5x of the cloud. Decode dominates everything, which is why the Mac's weak prefill is not even the main problem.

Opus 5 cost per task, at $5 in / $25 out, cache read 0.1x, 1-hour cache write 2x

┌───────────────────────────────────────────────┬────────┬──────┐
│                     Item                      │ Tokens │ Cost │
├───────────────────────────────────────────────┼────────┼──────┤
│ Cache reads                                   │ 15.25M │ $7.6 │
├───────────────────────────────────────────────┼────────┼──────┤
│ Cache writes (300k once plus 450k re-written) │ 750k   │ $7.5 │
├───────────────────────────────────────────────┼────────┼──────┤
│ Output                                        │ 120k   │ $3.0 │
├───────────────────────────────────────────────┼────────┼──────┤
│ Total                                         │        │ ~$18 │
└───────────────────────────────────────────────┴────────┴──────┘

With the 5-minute cache it would be about $15.5. Without caching at all, $83, which is why the caching detail matters more than the model price.

Electricity, at 0.20 €/kWh

- Mac: about 200 W for 3 hours is 0.6 kWh, so ~0.12 €.
- GPU node: about 3 kW average during the session (decode is memory-bound, so the cards sit well under their 600 W cap) for 45 minutes is 2.3 kWh, so ~0.45 €, or ~0.6 € with cooling. The node also idles at roughly 0.8 kW, which is ~1,400 € per year whether anyone uses it or not.

The number that falls out of this. Amortizing 150k€ over three years is 50k€ a year. At ~16 € per Opus 5 task, the node breaks even at about 14 such tasks per working day, or roughly 5 to 7 developers using it seriously. Below that, the API is cheaper. Above it, the node wins, and it wins on latency too, since ten concurrent sessions each still get close to the 45-minute figure. The Mac is a 3-hour single-stream machine at any team size.

...

API price is not energy. Rough datacenter figure for a frontier model at scale: 5 to 20 joules per output token, all-in with prefill, cooling, and networking. Your 120k-token task is then 0.2 to 0.7 kWh, or a few cents. Against $18 that is well under 1% of the bill. The rest is hardware amortization (a GB200 rack is several million dollars and depreciates in 3 to 5 years), training cost recovery, and margin. Hardware amortization is the biggest of the three, which is exactly the lever you are trying to own by buying the node.

Where the GPU node's energy goes single-stream. Decode is a memory-bound loop: every step reads all active weights once, whether it produces one token or a hundred. So per-token energy has a fixed part (weight read plus the static power of eight lit-up GPUs and a dual-socket host) divided by batch size, plus a tiny compute part. At batch 1 you pay the whole fixed part for one token.

┌────────────────────────────────┬─────────┬───────────────────────┬──────────────────┐
│         Configuration          │  Power  │      Output rate      │ Energy per token │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ Mac M3 Ultra, single stream    │ ~200 W  │ ~13 tok/s             │ ~15 J            │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ GPU node, single stream        │ ~3 kW   │ ~55 tok/s             │ ~55 J            │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ GPU node, 32 streams           │ ~4 kW   │ ~1,200 tok/s          │ ~3 J             │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ GPU node, 128 streams          │ ~4.5 kW │ ~2,000 to 3,500 tok/s │ ~1.5 to 2 J      │
├────────────────────────────────┼─────────┼───────────────────────┼──────────────────┤
│ Frontier datacenter, estimated │         │                       │ ~1 to 5 J        │
└────────────────────────────────┴─────────┴───────────────────────┴──────────────────┘

The concurrency rows use the measured aggregates from the earlier tables (900 tok/s for Kimi at 100 streams, 3,500 for Qwen3.5-397B at 128). So your intuition is right: batched, the node beats the Mac per token by roughly 5 to 10x and lands in the same band as a datacenter. Single-stream it is the worst of the three because you are running an 8-GPU machine to feed one person.

HBM is not the magic ingredient. Per bit moved, HBM3e and LPDDR5X are in the same neighborhood, and GDDR7 on the RTX cards is roughly twice as costly. But the memory transfer itself is a minority of the per-token energy at batch 1. The Mac reading 25 GB of weights per token spends under 1 J on the transfer and burns the other 14 J on the SoC being awake and underutilized. HBM's real advantage is bandwidth per package, which lets one GPU sustain a much larger batch before it runs out of bandwidth. Batching is the lever. HBM just raises the ceiling on how far you can push it.

The Mac cannot batch its way out. MLX can serve a few streams, but its weak compute hits the ceiling quickly, and a single developer rarely has ten concurrent sessions. The efficiency gap between the Mac and the node is a gap in the number of users, not in the silicon.

Batching has a latency cost the provider tunes. A fully saturated datacenter GPU gives each stream a smaller slice of bandwidth, so per-user speed drops as utilization rises. Providers pick a batch size that holds per-stream speed near a target and accept queueing above that. The stalls you see in the thinking counter are almost certainly that queueing plus harness pauses, not the model. Your own node at 10 concurrent users sits in the comfortable part of the curve: 10x better energy than batch 1 and per-stream speed barely below single-stream. That is a real advantage of owning hardware that is sized for your team rather than shared with the world.

 
The following users thanked this post: KE5FX

Offline tom66Topic starter

  • Super Contributor
  • ***
  • Posts: 8840
  • Country: gb
  • Professional HW / FPGA / Embedded Engr. & Hobbyist
Re: Local AI and hardware to run it
« Reply #11 on: September 08, 2026, 04:30:01 pm »
The system ideally needs to be able to host at least four users concurrently with acceptable performance of 15-20 tok/s, as well as support code completion.
Are you assuming a Studio per person for your 4+ users? An alternative is to use smaller machines with Thunderbolt 5 for RDMA. The gain is more for use by a group than an individual, more efficient use of the resources. It also bumps your effective available memory.

I am talking about possible technology options. I do not know your real requirements or priorities.

No, goal would be one Studio per four engineers, as a remote resource sitting in the server room in air conditioning. It would be accessible over a network connection for code completion and chatbot functions.

We aren't (yet) using a lot of agentic AI, but that may change in the near future.  In which case we'd likely buy a second machine for agentic workloads.
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: Local AI and hardware to run it
« Reply #12 on: September 08, 2026, 05:58:07 pm »
No, goal would be one Studio per four engineers, as a remote resource sitting in the server room in air conditioning.

That's exactly what it would struggle with, unless the engineers stagger their working times. Which would be quite difficult, since on baseline (1 engineer, 1 task at a time) it's already slow compared to datacenter HW, something the engineers are probably used to already. The fact you would be waiting 1-2 hours for a simple job to finish increases the pressure to run multiple in parallel, and increases the pressure for stuff to run overnight, for all those 4 engineers, so the box would sit at high utilization. Great investment value for the HW - it isn't that expensive, and you utilize it 100% - but not so great if you calculate any value for the human time, as you probably should.

I think if the company has four engineers working full-time with software, then the 150k option that really runs 3-4 parallel agent flow from every 4 engineers (so ~16 streams total concurrently) at nearly cloud speeds, does not sound bad at all. That's what I would seriously consider. The Mac box would be a great introduction for an engineer or two, they can evaluate the idea and then there's an upgrade path. The appeal is exactly that you don't need to put it in a server room and think about how it shares the resources. You just buy it and put it on your table and start playing around.

Of course, financially this doesn't make much sense because there is plenty of relatively cheap per-token billed inference you can buy in the cloud, so I'm assuming the self-sustainability is a value here in itself. But if you give the Mac Studio to 4 real software engineers who have worked with modern-day AI workflows you are going to have 4 very disappointed software engineers.

And then again, if you give them Claude API with $500 spend limit per month, they will be even more disappointed. Apparently this is also happening in real companies, so it's very hard to compare the options  :-//

I spent $7000 worth of Claude's API usage last month, for the $200 fixed fee. Anthropic's API rates truly are horribly expensive, so if you want to rationalize a HW investment, that's a good comparison. And nothing wrong in that, we rationalize buying toys we want all the time - fine. Four engineers Claude'ing at my rate is $28000 / month, so the Mac pays for itself in 2 weeks (except it doesn't because it can't run that workload), the $150k set pays itself in half a year.
« Last Edit: September 08, 2026, 06:08:06 pm by Siwastaja »
 
The following users thanked this post: thm_w

Offline tom66Topic starter

  • Super Contributor
  • ***
  • Posts: 8840
  • Country: gb
  • Professional HW / FPGA / Embedded Engr. & Hobbyist
Re: Local AI and hardware to run it
« Reply #13 on: September 08, 2026, 06:38:18 pm »
I think we're doing quite different things with AI.  Most of our work is writing code, planning, and finding bugs.  There, 50 tok/sec is more than fast enough - Claude via GitHub Copilot is somewhere around 100tok/sec.  We might be able to justify 1 Studio per engineer, but it's a much larger investment.  But, as I said to my boss already, I think Github Copilot itself is truly worth at least £1k/m in improved productivity, so a £10k machine might be justifiable on that level.

We have found in our limited tests so far, as our application doesn't work well with conventional testing pipelines (it's very bare-to-the-metal), you can't make an agent that designs a feature and then tests it.  Because it can't test it.  So you end up with the human in the loop still.  The goal is to make the human as efficient as possible.  (We're working on a CI/CD pipeline for this stuff but we need to implement huge parts of our hardware into a simulation harness, which is a lot of work that hasn't been done yet.)

I think that given we've already seen one significant uplift in API prices, it won't be that long before we see another.  I could easily see our £100-200/m increase by 10x in the next few years without any change in workload, simply due to the enormous cost of datacenter infrastructure.  There's apparently a $3tn investment committed so far into this stuff; that's a scale that just boggles the mind.  Investors are going to start getting antsy for a return on that.   Also, we don't exclusively use Claude, so we don't want to only use Claude's subscription, as generous as it might be for that application.  We're using a mix of Gemini, GPT-4 for code completion, and Claude Opus/Fable as a kind of "expert support".  This might mean that a computer like the Mac Studio becomes something like a £50k purchase in a few years time.  It's hard to say; predicting the future is tricky.  But I doubt it is getting any cheaper any time soon, at least.
 

Offline ejeffrey

  • Super Contributor
  • ***
  • Posts: 4842
  • Country: us
Re: Local AI and hardware to run it
« Reply #14 on: September 09, 2026, 04:21:32 am »
That's exactly what it would struggle with, unless the engineers stagger their working times. Which would be quite difficult, since on baseline (1 engineer, 1 task at a time) it's already slow compared to datacenter HW, something the engineers are probably used to already

One thing I don't know about is how well the Mac scales with batch parallelism.  It's got nowhere near the FP4/FP8 FLOPs of the high end NVidia GPUs, but it may be able to run half a dozen users in parallel without much hit, as long as you provision the additional RAM for multiple user KV caches.  If your engineers were happy with something like DeepSeek V4 Flash or Qwen 3.8 Flash or FLM 5.3 flash, a single mac studio shared by 4 engineers might actually be pretty good.  And honestly, if mostly what you want is code completion and chat, this might be good enough.

But if your engineers want to do large codebase analysis, complex development where you give it the requirements and analyzes the existing code base, writes a multi step plan, creates tests, and executes, they aren't going to be happy with the flash models exclusively.  And scaling up the local model to support one of the "max" type models (2T+ parameters) is going to get expensive.  And then you will be wanting to amortize the costs over a couple dozen users, and then you will probably start to hit the FLOPs limit on the apple silicon and want a real GPU.  And pretty soon you are at a $500K+ nvidia cluster that needs to support hundreds of users to justify the cost and you have a dedicate IT staff to run it.

I actually have reasonable hope for local LLMs in the long run. for a couple of reasons.  For one, eventually RAM is going to come back down in price.  It may take a couple of years, but unless Skynet kills us first all, RAM is going to eventually drop in price, and then a computer with 1 TB+ of RAM is not going to be crazy.  The other reason is that the model power trajectory is pretty good.  I've been playing with the open weight models for a while now, and the flash models have gotten *much* better in just the last 3-6 months.  It's not just that the field as a whole has advanced, my definite perception is that the flash models are partly closing the gap with frontier models.  I'm also doing a lot more with Claude Sonnet and Gemini Flash -- although I don't know what size those models are, and Google has apparently been having serious problems with their Pro model that was expected to be released back in April.    Finally, eventually other people are going to copy Apple and start building affordable unified memory systems with good bandwidth and decent iGPUs.  And I'm sure apple is looking at the landscale as they design their next generation M processors and considering beefing up the GPU to close some of the gap with Nvidia.

All this together makes me think that it's going to be -- if not strictly cost effective -- at least financially feasible to buy an LLM server for your personal use or to share with a handful of developers.  And there are plenty of reasons to want that even if it's not cost effective.  I'm sure lots of others would like to reduce their dependency on big tech and keep control of their own data.  But it has to be cheaper and better before that's going to be practical.
 

Offline KE5FX

  • Super Contributor
  • ***
  • Posts: 2638
  • Country: us
    • KE5FX.COM
Re: Local AI and hardware to run it
« Reply #15 on: September 09, 2026, 06:19:54 am »
We don't know much about how the new Mac Studio is going to perform yet.  I would not go that direction unless/until someone figures out how to make it talk to GPUs.  I have a feeling that if you go with 4x RTX 6000s on a Linux PC or server, each of those machines will support twice as many developers as one Studio. 

And you'll have the ability to change your mind later and reconfigure them as 8x nodes if desired.  With Apple you are locked into whatever you bring home from the store on day 1.

4x RTX 6000s works especially well with GLM 5.3 Flash and DeepSeek 4 Flash, as well as the latest/greatest Qwen 27B edition.

It would be a very good idea to spend some time on the Local Inference Lab Discord before (and after) committing a lot of cash to a particular setup.  Using the right quant and the right inference recipe on the hardware you have makes all the difference in the world, and the people there are building up some substantial expertise.
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: Local AI and hardware to run it
« Reply #16 on: September 09, 2026, 06:47:37 am »
Honestly, I don't understand the "autocompletion" use case. I gave it a shot ~2 years ago when it was all the rage (today's agentic software development was still at its infancy) and found it totally useless. I did ask Copilot to review functions in VS code and it suggested making good code worse (e.g., thought that sizeof is a function and started adding random () where it doesn't belong - classic beginner mistake), was not able to do simple refactors like renaming a variable. It was also the same time the study came out showing that developers expect 30% improvement but get slowed down by 20% instead.

Maybe it can be made work somewhat better today, e.g. I guess the variable renaming would work today, but it's still in the 30% improvement territory. To me, all the value comes from human-like (but 100x faster) capabilities of understand large context and actually do human-like contributions autonomously. This is the 100x productivity territory and there is no doubt about it, stuff you never imagined you would ever have time to deal with just starts to happen. And for that, the 512GB Mac + GLM-5.3 or similar is awfully close, which makes it interesting. It seems though that quantizing something like GLM-5.3 to 4 bits might exactly kill the long-context agentic programming capabilities, and for significant work on meaningful features around 150ktokens of context feels like bare minimum, so that means you can't run it on 512GB of memory, but it's awfully close as the 768GB set Claude suggested to me will run it very nicely.

Now the question is what happens if you take smaller models (e.g. the Flash version KE5FX mentions above) and design a very good harness which divides the tasks. Some smaller models would easily run on that Mac, even a few in parallel, even for 4 engineers. But the track record for using complex harnesses and multi-agent farms is mixed. People have high expectations but mostly they fail. A complicated workflow trends up, and then comes the next frontier model which manages the complexity through its actual cognitive capabilities in a simpler harness.

But I'm sure the Mac would run pretty smoothly for the "analyze the flow of this function" or "refactor this function into pieces" or autocompletion, but this again I believe is the 30% productivity increase territory (or, as MS themselves marketed, 50%, which might be true in some cases). I find it even more uninteresting today than it was a few years ago. And it would still have more latency than in cloud (remember, simple models can be run on the cloud as well and get massive speed there); 3-second prefill probably doesn't matter for "refactor this function" but kills autocompletion.
« Last Edit: September 09, 2026, 06:53:45 am by Siwastaja »
 

Offline Swake

  • Super Contributor
  • ***
  • Posts: 1081
  • Country: be
Re: Local AI and hardware to run it
« Reply #17 on: September 09, 2026, 07:24:18 am »
The progress made with those LLM models goes at a speed never seen before. Last month's model is outdated compared to today's. Whatever experience you had 2 years ago is of no relevance today. 6 months ago I still gave instructions on how to structure things and so on, not anymore. Not writing a single line of code any longer, I'm not even peeking at the code. Nearing the stage of not caring what coding language is being used. The job of classic application dev engineering has changed entirely.

Stay away from the Mac ecosystem, it is too proprietary, you control nothing about it. Maybe good enough for a one person operation, but not enterprise worthy in any sense. Not even talking about the price/performance ratio.
When it fits, stop using the hammer
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: Local AI and hardware to run it
« Reply #18 on: September 09, 2026, 08:17:50 am »
The progress made with those LLM models goes at a speed never seen before. Last month's model is outdated compared to today's. Whatever experience you had 2 years ago is of no relevance today. 6 months ago I still gave instructions on how to structure things and so on, not anymore. Not writing a single line of code any longer, I'm not even peeking at the code.

Exactly, so you went the same way I did, but note that tom66 wants to use AI for autocompletion. That's the same thing I tested 2 years ago and found useless. How much has that advanced, and how much it can even advance is a different question. I don't believe there is much room for improvement. You want autocomplete, you get autocomplete. It's what it is. Limited capability even if it works perfectly. Useful, but mildly at best. In the meantime, as you say, agentic AI has grown into a very capable entire "engineer", and that's totally different compared to the early prototypes 2 years ago.

(I disagree somewhat with "last month's model is outdated" idea - this is somewhat marketing hype. For example, Fable is a significant improvement, in a limited sense some sort of gamechanger even, but then again even Opus 4.5 was very capable in February, and "a new Fable" doesn't happen every month (every year, maybe). Capabilities do fundamentally improve every year, which is already a very short time, but not every month really.)
« Last Edit: September 09, 2026, 08:23:46 am by Siwastaja »
 

Offline Swake

  • Super Contributor
  • ***
  • Posts: 1081
  • Country: be
Re: Local AI and hardware to run it
« Reply #19 on: September 09, 2026, 09:44:08 am »
Indeed, autocomplete is passé, we have this since long before the current LLM thing became a hype. There is no point anymore to write the code yourself if an automate can do it for you. Unless you're a nerdy programmer and that is what makes you kick, but that is an exception case and has another reason d'être.

I have to use local LLM (reason is not important), and was astonished about what Qwen 3.5 could do without given detailed instructions about what and how, but still it needed some guidance. Now, only months later, Qwen 3.8 has flabbergasted me again. I still force it in a workflow with discover/plan/build/verify/review/recap but the instructions therein have shrunk to things particular to my use case only and I feel like it can still be reduced.

What really shocked me this time is that before it was me being in charge of the details and telling the LLM what to do, now that has changed around and it is the LLM saying 'what about you evaluate this before I continue' and pointing things out that I really did not see coming but are spot on.

Firmware stuff and other 'hardware' code probably has still room for improvement, but that too is not going to take very long anymore.
When it fits, stop using the hammer
 

Offline KE5FX

  • Super Contributor
  • ***
  • Posts: 2638
  • Country: us
    • KE5FX.COM
Re: Local AI and hardware to run it
« Reply #20 on: September 09, 2026, 05:58:12 pm »
We have found in our limited tests so far, as our application doesn't work well with conventional testing pipelines (it's very bare-to-the-metal), you can't make an agent that designs a feature and then tests it.  Because it can't test it.

One interesting thing I recently tried was hooking up a counter and a coax switch matrix to a device under development, all connected to the PC via USB.  I asked Fable to compose a test plan for the DUT (DUD?), giving it a copy of the user manual and telling to to validate both the performance and functional aspects, along with instructions for how to use the counter and switch matrix.  Results were gratifying.  I was able to just walk away for a few hours and come back to a nicely-formatted test report with plots and everything. 

So, don't assume that it "can't test it."  Unless you are working on something seriously weird and/or dangerous, if people can test it, so can the clanker.
 
The following users thanked this post: Siwastaja

Offline ejeffrey

  • Super Contributor
  • ***
  • Posts: 4842
  • Country: us
Re: Local AI and hardware to run it
« Reply #21 on: September 09, 2026, 06:14:25 pm »
We don't know much about how the new Mac Studio is going to perform yet.  I would not go that direction unless/until someone figures out how to make it talk to GPUs.  I have a feeling that if you go with 4x RTX 6000s on a Linux PC or server, each of those machines will support twice as many developers as one Studio. 

At more than 4x the cost you would need to support more than 2x the users for that to make sense.  There is definitely a point where the GPUs win because they have higher FLOPS and you can use bigger batches to amortize the RAM cost.  But I don't know what that point is.  As long as you are in the realm of "I barely have enough RAM to support the models I want at all, and I want it to be as fast as possible" the Mac studios, of whatever size, are going to be dramatically cheaper than anything from nvidia, and far higher performance than anything at a similar price point.

[/quote]
4x RTX 6000s works especially well with GLM 5.3 Flash and DeepSeek 4 Flash, as well as the latest/greatest Qwen 27B edition.
[/quote]

Sure, but those are also things that the Mac Studio is going to be pretty good at.  It's not really true that "we don't know how it's going to perform", it's a pretty well understood architecture. 

Quote
It would be a very good idea to spend some time on the Local Inference Lab Discord before (and after) committing a lot of cash to a particular setup.  Using the right quant and the right inference recipe on the hardware you have makes all the difference in the world, and the people there are building up some substantial expertise.

I would say it's a better idea to test out the models you want on OpenRouter or a cloud service and make sure you are happy with the actual real world results of the model(s) you are interested in.  People in chat rooms and message boards can and will tell you almost anything, but a couple hundred dollars worth of cloud computing will tell you much more about what matters to you.  Then figure out if a given hardware option can actually run it at a useful speed.
 

Offline KE5FX

  • Super Contributor
  • ***
  • Posts: 2638
  • Country: us
    • KE5FX.COM
Re: Local AI and hardware to run it
« Reply #22 on: September 09, 2026, 06:31:41 pm »
Quote
I would say it's a better idea to test out the models you want on OpenRouter or a cloud service and make sure you are happy with the actual real world results of the model(s) you are interested in.  People in chat rooms and message boards can and will tell you almost anything, but a couple hundred dollars worth of cloud computing will tell you much more about what matters to you.  Then figure out if a given hardware option can actually run it at a useful speed.

Yes and no.  They maintain a large archive of quants and recipes, most with reproducible benchmarks of both decode and prefill tokens/second, often with KLD measurements as well.  The Discord is just where the sausage is made. 

When you use OpenRouter, you are using who-knows-what quant on who-knows-what hardware.  If you're lucky the provider is telling the truth, but even then it's unlikely that they will be running the exact setup that will end up being optimal in your situation.

Agreed that there's reason to be optimistic about the new Macs, especially since Apple doesn't seem to be overpricing the RAM this time around.  But I would still wait for real-world reports, as well as for actual pricing on the 512 GB model.  Prefill is a big deal in real-world use and they are not going to be great at that. 

Not that the RTX 6000 is a world-beater, but it is the best hardware most of us can put our hands on, and the upcoming Mac Studios won't change that fact by themselves.
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11224
  • Country: fi
Re: Local AI and hardware to run it
« Reply #23 on: September 09, 2026, 06:35:48 pm »
One "benefit" of the current situation is that HW keeps it value better than ever. In 1997 or 2010 you would have been a fool to buy a "too powerful" set of computer hardware, you would rather buy "just enough" for the job because if you find out you need more RAM / more CPU / more disk / whatever after just 6 months, you get the upgrade for much cheaper. Any investment lost half of its value in a year.

Now, you can buy the Mac, or the GPUs, or whatever you want if you have the cash. They'll keep their value like you were buying gold, good chances the value could even increase. Feels ridiculous but that what it is. So why not just buy the Mac, if it turns out to be a wrong move, just sell it for ~ the same money, or in a good case, for more.
 
The following users thanked this post: tom66

Offline ejeffrey

  • Super Contributor
  • ***
  • Posts: 4842
  • Country: us
Re: Local AI and hardware to run it
« Reply #24 on: September 09, 2026, 07:09:41 pm »
Yes and no.  They maintain a large archive of quants and recipes, most with reproducible benchmarks of both decode and prefill tokens/second, often with KLD measurements as well.  The Discord is just where the sausage is made. 

When you use OpenRouter, you are using who-knows-what quant on who-knows-what hardware.  If you're lucky the provider is telling the truth, but even then it's unlikely that they will be running the exact setup that will end up being optimal in your situation.

That's all fine to figure out if it's going to be fast, but much more important is to find out if you are actually going to be happy with the output.  That's the reason I suggest using a cloud service.  OpenRouter is cheap and easy, and if you have satisfactory output then it's likely that you will get similarly good results when you run it on your own hardware.  The thing I am more trying to caution against is people buying hardware sized specifically to run a single model and then being unhappy with the results, especially if they are used to something like Claude Code.

 
The following users thanked this post: Siwastaja


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->