Author Topic: Local AI and hardware to run it  (Read 4098 times)

0 Members and 2 Guests are viewing this topic.

Offline ejeffrey

  • Super Contributor
  • ***
  • Posts: 4841
  • Country: us
Re: Local AI and hardware to run it
« Reply #50 on: September 28, 2026, 03:08:22 pm »
I think in the long run the hope would be for high bandwidth flash to make it's way into the consumer space.  It doesn't even exist for datacenter applications yet so that is probably at least 3-5 years away.  It is probably going to have similar packaging costs as HBM so it may or may not ever be affordable.  But if you could take a mid range consumer gpu and add 1 TB of in-package flash that would make a pretty powerful inference device for a single user or a handful of users.  it still wont be cheap, but maybe at least affordable for professional use.

Alternately if HBF displaces some of the HBM in datacenters maybe that will help dram prices come back down to earth and people will be able to actually buy computers again.
« Last Edit: September 28, 2026, 03:10:47 pm by ejeffrey »
 

Offline Randy222

  • Super Contributor
  • ***
  • Posts: 1803
  • Country: ca
Re: Local AI and hardware to run it
« Reply #51 on: September 28, 2026, 05:12:15 pm »
AI needs a cpu with 900,000 cores. ;)
AI needs ultrabandwidth memory, HBM is too slow.

Three ways to do it, processing in memory, interleaving compute inside the memory stack, or going 2D rather than 3D and hybrid bonding a little compute chiplet on each memory chip and putting them all next to each other (bad for interconnect bandwidth. but it's of lesser importance and big stacks are terrible for yields).
:-//

WSE3 is 900,000 cores w/ 44GB of on-chip SRAM @ 21PB/s

Is that not "ultrabandwidth". Serious question because this tech is moving fast.

 

Offline booscrawl

  • Regular Contributor
  • *
  • Posts: 189
  • Country: us
Re: Local AI and hardware to run it
« Reply #52 on: September 29, 2026, 01:26:22 am »
I would not go that direction unless/until someone figures out how to make it talk to GPUs.

Any Mac with USB4/Thunderbolt can talk to GPUs at x4 PCIe speeds using an adapter and tinygrad today.
 

Offline KE5FX

  • Super Contributor
  • ***
  • Posts: 2638
  • Country: us
    • KE5FX.COM
Re: Local AI and hardware to run it
« Reply #53 on: September 29, 2026, 03:05:24 am »
I would not go that direction unless/until someone figures out how to make it talk to GPUs.

Any Mac with USB4/Thunderbolt can talk to GPUs at x4 PCIe speeds using an adapter and tinygrad today.

Interesting.  So Macs are OK with loading the nvidia driver these days?
 

Offline SpacedCowboy

  • Frequent Contributor
  • **
  • Posts: 435
  • Country: gb
  • Aging physicist
Re: Local AI and hardware to run it
« Reply #54 on: September 29, 2026, 06:57:13 am »
I would not go that direction unless/until someone figures out how to make it talk to GPUs.

Any Mac with USB4/Thunderbolt can talk to GPUs at x4 PCIe speeds using an adapter and tinygrad today.

Interesting.  So Macs are OK with loading the nvidia driver these days?

Tinygrad proves it’s possible to bypass the “thou shalt not connect an eGPU to Apple silicon”, but it does it by not registering as a GPU, it’s a kernel extension that maps the GPU hardware over a custom BAR range. That gives you access to the hardware but it doesn’t let you drive displays with it because it doesn’t match what the OS expects for a GPU card. This is just for AI inference/training.

The bandwidth issues (you’re limited by the TB bandwidth to and from the computer) mean this is really only useful for connecting a high-end large-vram GPU to a modest Mac mini, as effectively VRAM expansion to run larger models in (that fit in the eGPU VRAM) IMHO. You’re also out on a bit of a limb wrt software and community support. I think it’s a great technical achievement but not so sure on the utility unless you happen to want to link up a bunch of RTX cards you have lying around, and you also only have a Mac mini :)

I used to work in platform architecture at Apple. One of the projects I came up with to try (it’s run very much like postgrad research, come up with ideas or others will give you a pre-generated one) was to network, via thunderbolt, a bunch of macs for opencl to work with (I’m dating myself here :) ). It turned out to be just too inefficient for *almost* every use-case. If you were running a long-running weather simulation it would be awesome, but almost everyone wanted results from their kernels “right now, dammit”, not in a few hundred to a thousand ms. AI training/inference might actually make that viable again. I wonder if anyone has picked up my old project…

Edit: *slaps head*, yep RDMA… Officially supported even…
« Last Edit: September 29, 2026, 07:02:53 am by SpacedCowboy »
 

Offline Swake

  • Super Contributor
  • ***
  • Posts: 1081
  • Country: be
Re: Local AI and hardware to run it
« Reply #55 on: September 29, 2026, 07:10:02 am »
lol, a Mac is more of a status symbol than it will ever be a computing device. Good enough for the computer illiterate connecting very occasionally with other macs, if you start doing real computing things just ignore them.

The software options are limited and often entirely proprietary. The announced computing power is available for seconds only, after that it throttles down due to heat and/or power limits. Forget upgrading it and make sure to take your adapters if you want to connect about anything. Ok, this is slightly better with the last generation but still nowhere near what it should be. It is a very expensive lock-in ecosystem with low price/performance ratio. If it needs repairs you're f...... because it is going to cost you another leg.


When it fits, stop using the hammer
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6464
  • Country: nz
Re: Local AI and hardware to run it
« Reply #56 on: September 29, 2026, 08:44:35 am »
lol, a Mac is more of a status symbol than it will ever be a computing device. Good enough for the computer illiterate connecting very occasionally with other macs, if you start doing real computing things just ignore them.

The software options are limited and often entirely proprietary. The announced computing power is available for seconds only, after that it throttles down due to heat and/or power limits. Forget upgrading it and make sure to take your adapters if you want to connect about anything. Ok, this is slightly better with the last generation but still nowhere near what it should be. It is a very expensive lock-in ecosystem with low price/performance ratio. If it needs repairs you're f...... because it is going to cost you another leg.

There must be a midwit meme for this somewhere.

NB the above is the opinion in the middle of the bell curve, and has been for 40 years.
 

Offline Swake

  • Super Contributor
  • ***
  • Posts: 1081
  • Country: be
Re: Local AI and hardware to run it
« Reply #57 on: September 29, 2026, 12:32:51 pm »
NB the above is the opinion in the middle of the bell curve, and has been for 40 years.
Haha, yes, 40+ years. They have a very good marketing department!
When it fits, stop using the hammer
 

Online Marco

  • Super Contributor
  • ***
  • Posts: 7751
  • Country: nl
Re: Local AI and hardware to run it
« Reply #58 on: September 29, 2026, 03:29:46 pm »
WSE3 is 900,000 cores w/ 44GB of on-chip SRAM @ 21PB/s

44GB is a bit tight though. Too much compute, not enough memory.
 

Offline KE5FX

  • Super Contributor
  • ***
  • Posts: 2638
  • Country: us
    • KE5FX.COM
Re: Local AI and hardware to run it
« Reply #59 on: September 29, 2026, 05:20:47 pm »
Edit: *slaps head*, yep RDMA… Officially supported even…

Right, I get the part about being able to allocate a BAR that's accessible to code running on the Mac and read/write to it via RDMA.  But what I don't understand is what happens after that.  Great, you can talk to the GPU hardware, but we're not talking about the EGA CRTC here, right?  These things are insanely complex. 

How do you actually know how to program the GPU hardware at the register level?  Nobody except nvidia knows how to do that, do they?  If the Mac won't load their driver, it seems like exposing a BAR is as far as you can go.
 

Offline SpacedCowboy

  • Frequent Contributor
  • **
  • Posts: 435
  • Country: gb
  • Aging physicist
Re: Local AI and hardware to run it
« Reply #60 on: September 29, 2026, 06:09:30 pm »
How do you actually know how to program the GPU hardware at the register level?  Nobody except nvidia knows how to do that, do they?  If the Mac won't load their driver, it seems like exposing a BAR is as far as you can go.

Well, given that this is an AI-focused company, and given that there's a very clear {try, measure, compare} automatable script you could do with this, perhaps they got an LLM to do it :)

Start off by setting up a PCIe monitor, watching what goes over the bus from the real driver, watch the effect on the screen (or capture it and feed it back to the automated system). Capturing PCIe isn't trivial, but it's certainly not impossible either.

But, more likely, they just got some ex-Nvidia engineers who know what they're doing, and reverse-engineered it. Or even just some very clever people - Asahi Linux happily uses the Mac GPUs, and Apple are equally forthcoming about their own GPU architecture internals as Nvidia...

lol, a Mac is more of a status symbol than it will ever be a computing device. Good enough for the computer illiterate connecting very occasionally with other macs, if you start doing real computing things just ignore them.

The software options are limited and often entirely proprietary. The announced computing power is available for seconds only, after that it throttles down due to heat and/or power limits. Forget upgrading it and make sure to take your adapters if you want to connect about anything. Ok, this is slightly better with the last generation but still nowhere near what it should be. It is a very expensive lock-in ecosystem with low price/performance ratio. If it needs repairs you're f...... because it is going to cost you another leg.

Hmm, let's think:

  • Fastest single-core computer on the planet ? check.
  • Largest memory of any desktop computer on the planet ? check.
  • Most profitable computer-company on the planet ? check.
  • The best hardware on any portable computer on the planet ? check.
  • 40,000 employees all inter-linked with Mac computers ("connecting very occasionally" my hairy backside...) ? check.
  • The computer line that still holds its value after years of operation on the second-hand market? check.

And you question the demand for them ... I mean, I could go on and on. For someone to call the Mac computers "A status symbol" is akin to being the blind man in a sighted society, except that you can explain colour to a blind man, you say that blue is like the feel of water on your skin, that green is the texture of the grass you run your fingers through, that red is the heat of the sun at midday. You can't argue with wilful ignorance - it's the "being ignorant and proud of that fact" that is the saddest part of it all.

I use a Mac because it's the best damn unix machine (bar SGI, but they went bust) I've ever used. Linux boxes are fantastic at server-side deployments, Windows boxes are great for day-to-day business interoperability, Macs are fantastic for programming, for creative work, and now for AI. To state otherwise is to piss into the wind, and all that happens then is you get a wet face. I prefer the non-wet face...
 

Offline kite31

  • Frequent Contributor
  • **
  • Posts: 263
  • Country: au
Re: Local AI and hardware to run it
« Reply #61 on: September 29, 2026, 10:14:54 pm »
Any portable computer has trouble with heat dissipation under intensive loads such as AI. A couple of minutes at full tilt then some slowdown, depending on size of the portable and weight of the work. A phone running local AI gets hot in seconds, being passively cooled. A Mac Studio runs flat out all day, as it must given its graphics-intensive uses in film production, 3D animation, and game development not to mention local AI.

If I see senior research engineers at Google choosing Macs as their most productive tool, should I assume they are clueless? If I saw that Australia's largest Cisco/MS WAN at the time was designed, traffic-simulated and later performance-analysed on a Mac, should I tell its architect that was folly?
 

Online Marco

  • Super Contributor
  • ***
  • Posts: 7751
  • Country: nl
Re: Local AI and hardware to run it
« Reply #62 on: September 29, 2026, 11:24:31 pm »
What's the point in hobbling a GPU with such a constrained interconnect? They aren't cheap, make the most of them.
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6464
  • Country: nz
Re: Local AI and hardware to run it
« Reply #63 on: September 29, 2026, 11:39:00 pm »
NB the above is the opinion in the middle of the bell curve, and has been for 40 years.
Haha, yes, 40+ years. They have a very good marketing department!

Microsoft certainly do. They've probably bought a lot of managers a lot of dinners.
 

Offline Swake

  • Super Contributor
  • ***
  • Posts: 1081
  • Country: be
Re: Local AI and hardware to run it
« Reply #64 on: September 30, 2026, 07:53:50 am »
Fastest single-core computer on the planet ? check.
Fastest as in clockspeed? -> Intel Ultra 9, comparing clockspeed is worthless in terms of comparing system performance.
Fastest compute for a CPU with only one core? -> Likely Pentium 4, maybe an old AMD or some exotic CPU for whatever supercomputer, anyhow it is going to be Y2k technology that is not really relevant today.
Fastest compute on a multi-core CPU but with only one core active? AMD Ryzen 9 9950X3D2, btw, this would be a very wasteful use case these days.
Fastest single thread? Likely the M3 ultra 32 with a very specific memory config. Again, why on earth would you buy such a chip to run only a single thread.

Largest memory of any desktop computer on the planet ? check.
A Dell Precision 7960 can hold 4TB of DDR5 ECC RAM. Other brands have similar offerings. Is there an Apple computer that comes close? Upgradable?

Most profitable computer-company on the planet ? check.
I would not compare compute performance based on financial numbers of the company?
If you do, at least compare apples with apples, (oops...  ;) yeah that one was an easy one ). Extract the numbers of the computer hardware department and compare it to HP, Dell, Lenovo, etc...

The best hardware on any portable computer on the planet ? check.
Very subjective subject.

40,000 employees all inter-linked with Mac computers ("connecting very occasionally" my hairy backside...) ? check.
How and to do what? I'm not curious about physical connection but about how this is managed and what trade-offs / compromises were made to get it working? For example, if that organization is running 'only Macs', well.... you're missing out on a couple things, at least that. Everyone is free to do how they prefer it.

The computer line that still holds its value after years of operation on the second-hand market? check.
True, that said, it is again very much 'marketing' related. For those users the perceived performance on the personal image is more important than the effective compute performance.


A Mac Studio runs flat out all day, as it must given its graphics-intensive uses in film production, 3D animation, and game development not to mention local AI.
Since the late '80 the graphics people have been 'worked to love the Macs'. And that is fine. I'm not saying that it is not working, I'm saying it is not an objective choice. Price / performance is plain bad. There are real use examples with QWEN3.8B27 on the Mac studio M5 Ultra (30 to 45 tk/s depending on context size) compared to a RTX5090 (98 to 156 tk/ same workload as on the Mac). That is 3x slower for 3x as much money. This rate cannot be justified.
When it fits, stop using the hammer
 

Offline brucehoult

  • Super Contributor
  • ***
  • Posts: 6464
  • Country: nz
Re: Local AI and hardware to run it
« Reply #65 on: September 30, 2026, 09:32:44 am »
Price / performance is plain bad. There are real use examples with QWEN3.8B27 on the Mac studio M5 Ultra (30 to 45 tk/s depending on context size) compared to a RTX5090 (98 to 156 tk/ same workload as on the Mac). That is 3x slower for 3x as much money. This rate cannot be justified.

That's a very strange model example to choose after all your "you don't buy a Ryzen 9 9950X3D2 to use only one core" examples.

You don't buy a Mac Studio with 512GB of unified in-package RAM to run Qwen 3.8 27B!!!   On that you'll want to run Qwen 3.8-Flash-Next in full BF16 precision, which needs around 355 GB of RAM.

See how your RTX 5090 likes that.
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11222
  • Country: fi
Re: Local AI and hardware to run it
« Reply #66 on: September 30, 2026, 11:01:44 am »
On that you'll want to run Qwen 3.8-Flash-Next in full BF16 precision, which needs around 355 GB of RAM.

The interesting question no one seems to be willing to answer, though, will it run 5 parallel agent, each averaging 300ktokens of context, and still fit the RAM at all, and if yes, what's the speed of those agents? This would be considered "light" use pattern by today's terms. A lot of RAM is useful, but performance, not just amount of RAM, is important too.

The comparison, 512GB worth of GPU's, is nearly 10x more expensive though, so it's Apples to oranges comparison, yes. You are right that the Mac Studio hits a certain combination of specs / prices that can't be had with the "buy GPUs" solutions - they are either smaller-RAM, faster, similar cost; or larger-RAM, significantly faster, but also significantly more expensive.

For a lot of RAM and mediocre performance, the Mac seems to nail it, with decent price point - it's pretty OK value for the amount of RAM, RAM being so expensive now. The hard question is really, how mediocre the mediocre performance is, and is it good enough or not.
« Last Edit: September 30, 2026, 11:03:53 am by Siwastaja »
 

Offline SpacedCowboy

  • Frequent Contributor
  • **
  • Posts: 435
  • Country: gb
  • Aging physicist
Re: Local AI and hardware to run it
« Reply #67 on: September 30, 2026, 11:29:59 am »
Fastest single-core computer on the planet ? check.
Fastest as in clockspeed? -> Intel Ultra 9, comparing clockspeed is worthless in terms of comparing system performance.
Fastest compute for a CPU with only one core? -> Likely Pentium 4, maybe an old AMD or some exotic CPU for whatever supercomputer, anyhow it is going to be Y2k technology that is not really relevant today.
Fastest compute on a multi-core CPU but with only one core active? AMD Ryzen 9 9950X3D2, btw, this would be a very wasteful use case these days.
Fastest single thread? Likely the M3 ultra 32 with a very specific memory config. Again, why on earth would you buy such a chip to run only a single thread.
If you don't understand that single-threaded compute is just as important as multi-threaded compute, I'm not sure there's even any point in discussion. Lots of tasks are irreducably serial in nature and cannot be composed into a parallel load. Conceptually, something as simple as:
Code: [Select]
void traverse_list(struct Node* head) {
    struct Node* current = head;
    int total = 0;

    while (current != NULL) {
        total += current->data;
        current = current->next; // Strict loop-carried dependency
    }

    printf("Total: %d\n", total);
}

... linked-list traversal will run on a single thread until completion, in general any loop where the result is dependent on the previous state is not something a compiler can parallelise. Oh, and it'd be the M6, not the M3 that's the fastest on the planet. The M6 is about 50% faster than the M3.


Largest memory of any desktop computer on the planet ? check.
A Dell Precision 7960 can hold 4TB of DDR5 ECC RAM. Other brands have similar offerings. Is there an Apple computer that comes close? Upgradable?

That's not a desktop, dude, that's a floor-standing full-tower. The studio *is* a desktop, *is* quiet even under load so you don't need headphones on to use it (50dB!), and *is* about 1/15 the size of that behemoth.

Most profitable computer-company on the planet ? check.
I would not compare compute performance based on financial numbers of the company?
If you do, at least compare apples with apples, (oops...  ;) yeah that one was an easy one ). Extract the numbers of the computer hardware department and compare it to HP, Dell, Lenovo, etc...

My point there is not to do with performance, it's to do with popularity. When a company is the largest computer company on the planet, they're doing something right. Apple extract revenue from computer users for more than just hardware sales, just like others (including dell, hp). Just using computers and computer-related income (no ipads, phones, tv, headphones, watches etc) the breakdown looks something like

Apple: 143 Billion
Dell: 92 Billion
HP: 55 Billion

Apple is pretty much both of the other two put together.

The best hardware on any portable computer on the planet ? check.
Very subjective subject.

Ok, show me a portable that has comparable computing power and battery life, that is as light, and that has hardware you actually want to use (eg: a touchpad that ... works in the corner)
 
40,000 employees all inter-linked with Mac computers ("connecting very occasionally" my hairy backside...) ? check.
How and to do what? I'm not curious about physical connection but about how this is managed and what trade-offs / compromises were made to get it working? For example, if that organization is running 'only Macs', well.... you're missing out on a couple things, at least that. Everyone is free to do how they prefer it.

Apple run their entire business off Apple computers - everyone from a secretary through engineers, legal, finance, execs has an Apple mac, generally a portable, even if its permanently docked. Engineers get the bigger toys to play with.

Let's see, if I wanted to arrange a meeting, I'd use Apple Directory, which shows me rooms (and their convenience features), people, manages the overlapping schedules pulled from people's Calendar, I can search, click on a name, launch video-chat directly, or group-chat. It shows me a map of the building and the optimal route to get there, rendered in 3D if necessary. It can transfer that route to my phone, and update it as I go even if just walking around Apple Park, or Infinite loop.

Documents are universally PDF, which is built-in to the OS for both editing and viewing, Apple Mail handles all internal email, and slack.apple.com is for IM. Shared storage is available via my dept. group and externally via (IIRC) Dropbox, I can put files into my Public folder and send people a link, they can read the file but not see the directory contents, I have CI processes for my code, all integrated into slack and email, both phone and desktop. I could just go on and on here. Literally everything from expenses to company-issued device-management has an application.

What's missing ? Nothing. Nothing is missing.

The computer line that still holds its value after years of operation on the second-hand market? check.
True, that said, it is again very much 'marketing' related. For those users the perceived performance on the personal image is more important than the effective compute performance.

You seem to have this mental image of a Mac user that it's all about image. I don't know anyone like that, I've only read about these mythical people in PC-focused computer magazines and websites. I know that the most popular computer at Google (softeware engineers, designers, managers) is a MacBook Pro. I know that the most popular computer in college education (ie: when you first *choose* your own computer) is a Macbook (of some type, chosen by the vast majority of humanities, business, communications, pre-med, computer science, film/video production, engineering, and architecture majors). I know that Apple run one of the largest companies on the planet exclusively on Macs. These people choose (generally a macbook/pro) because its simply the best option, even if more expensive than the cheaper chromebook/PC.

A Mac Studio runs flat out all day, as it must given its graphics-intensive uses in film production, 3D animation, and game development not to mention local AI.
Since the late '80 the graphics people have been 'worked to love the Macs'. And that is fine. I'm not saying that it is not working, I'm saying it is not an objective choice. Price / performance is plain bad. There are real use examples with QWEN3.8B27 on the Mac studio M5 Ultra (30 to 45 tk/s depending on context size) compared to a RTX5090 (98 to 156 tk/ same workload as on the Mac). That is 3x slower for 3x as much money. This rate cannot be justified.

See above. Price is indeed not wonderful. Performance is though. About the only place where I'd voluntarily choose a PC these days is for games, and honestly I don't see that lasting much longer. The Apple GPUs will start to become "easily good enough" over the next few years - the one in my M4 Max is currently "easily good enough" and that's only about 2x that of the base M6 mini.

As for that RTX5090 (which is £5600 on its own at Overclockers UK and need a reasonably beefy machine to make it useful, so you're into Studio territory), it's pretty useless for larger models, no ? I tend to use Claude so I don't care these days, but I didn't find QWEN3.8B27 to be much use personally. What actually might be useful for the sort of coding I want LLM's for is Llama 4 Maverick (400B) or DeepSeek-R1 (671B), using 4-bit or 3-bit quantised weights. I could fit those into a 512GB Studio, on my desk (probably hiding behind the 4K monitor) and maybe even get useful results out of it. I'm not sure it's worth an £18k gamble though, so I'll stick with Claude until RAM prices recover in a half-decade or so, then maybe look at an M8 studio with 1TB of RAM.
 

Offline SpacedCowboy

  • Frequent Contributor
  • **
  • Posts: 435
  • Country: gb
  • Aging physicist
Re: Local AI and hardware to run it
« Reply #68 on: September 30, 2026, 11:45:57 am »
On that you'll want to run Qwen 3.8-Flash-Next in full BF16 precision, which needs around 355 GB of RAM.

The interesting question no one seems to be willing to answer, though, will it run 5 parallel agent, each averaging 300ktokens of context, and still fit the RAM at all, and if yes, what's the speed of those agents? This would be considered "light" use pattern by today's terms. A lot of RAM is useful, but performance, not just amount of RAM, is important too.

The answer to that is no. No it won't :) The benefit the Macs give you is that you can fit a large-ish model into the unified RAM. You'd need enough space for model + 3 contexts to have 3 concurrent sessions, and you'd get 1/3 of the throughput (at best, though I think it'd be close). The sort of models I want to put in there would be on the order of the size of the RAM, so I'd personally be squeezed on context.

If you want more concurrent large-models, the answer is in either just multiple machines (price = ouch) or multiple machines with RDMA (still price = ouch). I'm actually wondering if there's a project inside Platform Architecture in Cupertino where there's a dedicated box to add to a TB5 mac which gives you just memory-over-tb5. RDMA bypasses the OS anyway, so it could be done just with a hardware ASIC. The latency is of course much higher than built-in RAM but you can hide a lot of that via layer-pipelining, micro-tiling and ring-based all-reduce. Couple that with RDMA reducing the native TB5 latency from ~300ms to ~3ms and a plug-in memory block might be a feasible thing for AI. You'd still see the performance hit if running multiple contexts though - there's only so many compute modules, that's when you "just" buy another Studio.
[/quote]
« Last Edit: September 30, 2026, 11:51:04 am by SpacedCowboy »
 

Offline Randy222

  • Super Contributor
  • ***
  • Posts: 1803
  • Country: ca
Re: Local AI and hardware to run it
« Reply #69 on: September 30, 2026, 04:44:16 pm »
WSE3 is 900,000 cores w/ 44GB of on-chip SRAM @ 21PB/s

44GB is a bit tight though. Too much compute, not enough memory.

......... @21PB/s

If memory can move faster, do we technically need more of it? If one task has to wait for RAM but the wait time is 10x less than before, is that an issue?

Not all AI computing is the same, so I guess the debate needs to be put into various contexts. ML and Inference are very different and require different types of compute to obtain maximum performance.

WSE3 is best at ML, nvidia plays in the Inference world.

Since AI is fairly new, as of today, there's an infinite amount of ML models to train, which then means an infinite amount of Inference engines.
However with a caveat, wondering how "infinite" it is if (if) a model is created that itself can generate and build new models and then do all the ML things in automated fashion. So a bit like a nuclear fission reaction, etc.

I am however skeptical in terms of how much "I" there is in "AI". Has any new physics been developed by anything "AI"? Any current theories squashed or found to be proven true?
 

Offline KE5FX

  • Super Contributor
  • ***
  • Posts: 2638
  • Country: us
    • KE5FX.COM
Re: Local AI and hardware to run it
« Reply #70 on: September 30, 2026, 05:09:27 pm »
I do like the notion of plugging in an arbitrarily-long chain of external memory blocks. 

I have a feeling that if we were Doing It Right, whatever that may turn out to mean, the bus throughput wouldn't be especially critical.  There is such a monstrous disparity between the rate at which the layers communicate with each other and the rate at which tokens are ultimately returned to the user.  It's just crazy how much work an autoregressive model has to do to add one more token to the pile.  :(  Historically that kind of disparity suggests optimization opportunities, but it's all way above my pay grade.  MoE is a step in the right direction but it just feels too much like a hack to me.

Has any new physics been developed by anything "AI"? Any current theories squashed or found to be proven true?

I don't think mathematicians would be as butthurt about this stuff as they are if they didn't feel deeply threatened by it. 

Some of Terence Tao's videos might be of interest in that regard.  I have a lot of them on my playlist that I haven't gotten to yet, myself.  But if anyone can speak authoritatively about what is and is not realistic to expect from LLMs, it would be him.
 

Online Siwastaja

  • Super Contributor
  • ***
  • Posts: 11222
  • Country: fi
Re: Local AI and hardware to run it
« Reply #71 on: September 30, 2026, 06:04:50 pm »
MoE is a step in the right direction but it just feels too much like a hack to me.

Don't think that as a hack, rather, a necessity; like a very effective lossy compression. "Helpful" human slop writers have done a lot of damage by trying to explain MoE with incorrect analogies like "one part of the network is a math expert, another a programming expert" and so on. A better analogy is any familiar lossy compression (say, JPEG, MP3, etc.) - think about DTC or Fourier transform which leaves a lot of near-zero coefficients which can be then truncated to zero for significant space savings with little visible loss. MoE basically does the same - removes links that contribute almost nothing. Significant memory bandwidth saving, which you can then use to make the model itself larger, thus better.

And if you can choose between an uncompressed 640x480 pixel video, versus slightly compressed (nearly no compression artifacts) 4K stream at similar bitrate, it is obvious which one has better fidelity for most purposes. For the same reason, every modern-day, performant, large model is MoE. Very small models don't benefit similarly, and dense models can be afforded anyway.

Clearly the AI industry now focuses on optimization. Opus 5.5 feels like a really good release again; it really seems to be on par with the Fable with quality/capabilities, but it's blazing fast, so I fully believe that the significant price drop is not just marketing, it's true drop in compute requirements. Same seems to be true on the open-weight field, at least people say GLM-5.3 is significantly more capable than GLM-5.2, with exact same structure and as such, same compute requirements.

But there will be limits how far this can go. I don't believe a 100-200B parameter model can ever be Fable-level. Most tasks really need more knowledge than is obvious, and encoding that knowledge, and all the complexities of autonomous, correct behavior, simply requires certain number of gigabits.
 

Offline ejeffrey

  • Super Contributor
  • ***
  • Posts: 4841
  • Country: us
Re: Local AI and hardware to run it
« Reply #72 on: September 30, 2026, 06:49:43 pm »
Don't think that as a hack, rather, a necessity; like a very effective lossy compression. "Helpful" human slop writers have done a lot of damage by trying to explain MoE with incorrect analogies like "one part of the network is a math expert, another a programming expert" and so on. A better analogy is any familiar lossy compression (say, JPEG, MP3, etc.) - think about DTC or Fourier transform which leaves a lot of near-zero coefficients which can be then truncated to zero for significant space savings with little visible loss. MoE basically does the same - removes links that contribute almost nothing. Significant memory bandwidth saving, which you can then use to make the model itself larger, thus better.

This +1000.  MoE is a bad name that has unfortunately stuck, it's much better to call them sparse: it has huge blocks of zeros that don't have to be computed.

To understand why, you just have to look at the structure.  The "experts" are routed per-layer, per-token.  So a model with 64 layers and 256 experts/layer has 16k total experts.  So a typical query will use most or all of them even though a single token only uses 5-10%.    Any specialization is much more fine grained, than something like "math" vs "programming"

Quote
Clearly the AI industry now focuses on optimization. Opus 5.5 feels like a really good release again; it really seems to be on par with the Fable with quality/capabilities, but it's blazing fast, so I fully believe that the significant price drop is not just marketing, it's true drop in compute requirements. Same seems to be true on the open-weight field, at least people say GLM-5.3 is significantly more capable than GLM-5.2, with exact same structure and as such, same compute requirements.

Yeah, overall the biggest change I have seen over the last 6 months is that mid-range models have significantly closed the gap with frontier models.  The open weight "flash" models have made tremendous improvements.  Google hasn't release a pro model in 6 months, but they have released 4 flash models and their latest one is on par or exceeding the flagship models from 6 months ago, and Opus 5.5 is essentially moving in the direction of a flash model while maintaining the quality.  There are still applications for the flagship models, and I don't see that going away, but more and more things can be handled by cheaper and lighter models.

Quote
But there will be limits how far this can go. I don't believe a 100-200B parameter model can ever be Fable-level. Most tasks really need more knowledge than is obvious, and encoding that knowledge, and all the complexities of autonomous, correct behavior, simply requires certain number of gigabits.

Probably not.  But there might be room to increase sparsity further.  In the open weight world, MoE sparsity is about 3-10%, at least for the big models.  Could that be 1%?  Or could the memory locality be improved, say by making expert selection sticky between layers and/or tokens?   Or having multiple layers of sparsity -- with sparse experts?  If you could improve locality and have hierarchical sparsity maybe you could store most of the LLM on an SSD.  On the other hand, how much can be pushed into the harness?  A huge fraction of the improvements of the past year are not just in the models, they came from using better harnesses that can implement memory, tools, sub-agents, MCP servers, and feedback.  That can compensate for lack of knowledge directly encoded into the weights.  Already it looks like companies are focusing training on task completion over knowledge -- a lot of people claim this is just bench-maxing, but if you can fill the knowledge from a separate database (maybe itself a very, very spare LLM that encodes a lot of knowledge but has very short context windows?) then it's just efficient design.

I'm not an expert in LLM design at all, I assume all those ideas are already being investigated and some of them are just dumb.   But it doesn't seem like we are done yet with architecture.
 

Offline booscrawl

  • Regular Contributor
  • *
  • Posts: 189
  • Country: us
Re: Local AI and hardware to run it
« Reply #73 on: Today at 12:21:04 am »
Largest memory of any desktop computer on the planet ? check.
A Dell Precision 7960 can hold 4TB of DDR5 ECC RAM. Other brands have similar offerings. Is there an Apple computer that comes close? Upgradable?

That's not a desktop, dude, that's a floor-standing full-tower. The studio *is* a desktop, *is* quiet even under load so you don't need headphones on to use it (50dB!), and *is* about 1/15 the size of that behemoth.

I wonder if Swake has experience spending $70-80,000 for 4 TB of DDR5 and is happy with getting only 307 GB/s of bandwidth from it. Oh, by the way, that CPU-mated DDR5 isn't unified with GPU memory, so prepare for that "307 GB/s" to be choked down to barely 64 GB/s if you use PCIe. Did someone say 'useless'?

The lowest end AS Ultra chip does 800 GB/s, the highest end: 1200 GB/s. In an almost entirely silent shoebox. On less than 300W. For less than $10k.
 


Share me

Digg  Facebook  SlashDot  Delicious  Technorati  Twitter  Google  Yahoo
Smf

 

-->