768GB RAM pipe dreams make no sense to Apple. By discontinuing 256GB / 512GB M3 Ultra and raising prices $5000 -> $7000 on Macbook pro with 128GB they basically confirmed how badly RAM shortage affecting them.
768GB is 64-times of 12GB which is rumored to be amount of RAM in new iPhones. Imagine what profit margin 768GB Mac Studio gonna need in order to justify making one instead of 64 iPhones.
Apple is the company that is okay about selling microfiber cloth for $100 and wheels for $700. Imagine how bad price hike for M3 Ultra 256GB / 512GB had to be in order for them to just discontinue them instead of getting free money out of desperate local AI folks.
They could treat the extreme spec machines separately from the prosumer ones, like they did with the Xserve. Let business customers spec up to 768GB (say) who are prepared for a $20-25k price tag, while keeping them away from the stores and usual consumer supply chains (Amazon et al). It may not be a big enough market segment for them to care about anymore, though.
They can do it, but that gonna need to be different SKU not Mac Studio. Otherwise news will be full of discussions about Apple price hike from $8000 to $24,000 or who the hell knows $48,000.
So yeah the only way I see them selling it is usual "call us" enterprise price tag.
But since its not what Apple usually do its easier to sell 4x Mac Studio 256GB RAM boxes with interconnect for lets say $12,000 - $15,000 each.
I don't think it's needs 'call us' just a separate SKU, I mean if they called it Xserve Ultra and it was just a studio in a 1u format with dual PSUs and extra RAM, it would fly off the shelves.
Isn't that the same thing you can get from ordinary 2S Epyc/Xeon servers at a similar price that have 24 memory channels (when the M3 Ultra has the equivalent of 16)?
And the reason people rarely use that for AI is that the enterprise GPUs from AMD and Nvidia are only moderately more expensive but are significantly faster because they use HBM instead of DDR5.
Yeah kind of, I think a 24 channels DDR5 works out approx 1TB/s, but the cost is astronomical, a M5 studio would probably beat that performance for around half the cost. You also get to use the GPU/NPU cores of the mac vs CPU only on the servers.
M5 ultra studio with 128GB RAM could probably beat out a sever with a RTX 6000 pro at half the price.
24 channels of DDR5-6400 is 1.2TB/s, M3 ultra is 0.8TB/s. They both use DDR5-6400.
> a M5 studio would probably beat that performance for around half the cost.
A barebones 2S system with no CPUs or memory is ~$2000, a pair of 16 core CPUs another ~$1000 each, and then however much memory you want. The price seems pretty comparable. The "problem" with doing this is actually that 128GB is too little memory, because you want to populate all the channels, but even using 16GB sticks, 24x16GB is already 384GB.
> You also get to use the GPU/NPU cores of the mac vs CPU only on the servers.
You only need enough cores to make sure the bottleneck is memory bandwidth.
M3 ultra is obviously 1-2 generations behind and new the studio is expected 'any day now. Even if this was M4 Ultra it would still be ~comparable to any EPYC system in bandwidth, but get to use the GPU for compute so potentially faster than the EPYC. Total Cost of Ownership in the Epyc is going to be WAY higher because of electricity costs, the EYPC is going to be consuming probably 5X the electricity and is probably not going to sit quietly on your desk. More RAM though, but again it's more about the ratio of RAM (size) to RAM (Memory Bandwith) to Compute and you may find a model bigger than e.g 70b suddenly is bottlenecked by the CPUs or memory bandwidth and therefore the extra RAM (size) is wasted. But maybe not, different use cases will yeild different results I guess.
> A barebones 2S system with no CPUs or memory is ~$2000, a pair of 16 core CPUs another ~$1000 each, and then however much memory you want.
As you say, the thing is it's not 'however much memory you want' it's 24 sticks which at $300 a stick for 16GB is $7200, then you also need at least one NVME disk so you're looking at what $13,000?
> M3 ultra is obviously 1-2 generations behind and new the studio is expected 'any day now.
M3 Ultra uses a 1024-bit memory bus, which is a major inconvenience to Apple because they're soldering everything. In ordinary systems if you have 16 memory slots and any one of the memory chips is bad, you replace that stick. If the processor is bad, you replace the processor. If the system board is bad, you transfer the processors and memory to another one.
If any of those has a defect after you solder thousands of dollars worth of memory onto the same board as a >$1000 CPU, you're not doing well. Worse, the more memory chips you have and the more pins the CPU needs for its memory bus, the higher the chances of one of them having a defect.
In addition to that, when you get to that number of pins it starts getting harder to run the memory at the highest speeds. The M3 uses DDR5-6400 but some of the M5 line is using DDR5-9600. It may or may not be possible to do that speed when using a 1024-bit bus -- not every existing M5 even does it. If it isn't then the Ultra wouldn't be much faster than the Max since it would have to use a lower memory speed. If it is then the tolerances would have to be even tighter and increase the defect rate even more.
Which is to say, I can see why they haven't released an Ultra since the M3.
> but get to use the GPU for compute so potentially faster than the EPYC.
Only if the bottleneck is compute rather than memory bandwidth, and for LLMs it's generally memory bandwidth. And if something significant was compute bound, there are also higher core count CPUs.
> Total Cost of Ownership in the Epyc is going to be WAY higher because of electricity costs, the EYPC is going to be consuming probably 5X the electricity and is probably not going to sit quietly on your desk
The M3 Ultra has a 480W TDP. There are relevant EPYC SKUs on SP5 at 125-200W/socket.
> you may find a model bigger than e.g 70b suddenly is bottlenecked by the CPUs or memory bandwidth
The extra RAM allows you to fit the larger model in memory to begin with, without which it's pretty hopeless. The bottleneck typically is memory bandwidth for LLMs, but that's true pretty much regardless of the model size. Moreover, mixture of experts models require significantly more RAM for the same amount of compute/bandwidth.
> you also need at least one NVME disk
That's ~$100.
> As you say, the thing is it's not 'however much memory you want' it's 24 sticks which at $300 a stick for 16GB is $7200
The chips don't cost a materially different amount based on whether you solder them. Apple presumably discontinued the 256GB and 512GB versions of the M3 Ultra because they'd have had to add a similar number to the price.
Using a server system designed for having hundreds of cores and TBs of RAM for LLMs because it has a medium-high amount of memory bandwidth is a hack for enthusiasts who want to take the price/performance trade off to run big models without paying for enterprise GPUs. The market for 8GB RDIMMs would be specifically that, since the more typical server workloads that need that amount of memory bandwidth also want the larger memory sticks (e.g. database servers, systems running hundreds of VMs), or aren't bounded by memory bandwidth to begin with and then don't need to populate all the channels.
And if you wanted that market segment then what you'd really do is produce a consumer GPU with 128GB of RAM.
> That would be the kind of thing you could sell as many as you could make.
People said that about the M1 Ultra Mac Pro, a few years before it was discontinued. I don't think there are many HPC customers looking at Apple hardware.
> Let business customers spec up to 768GB who are prepared for a $20-25k price tag
There is a clear difference between $25k and $100k.
64 iPhones at retail price is already around $64k. For something at 768GB to be profitable at Apple's terms, this has to retail at $100k for it to be profitable. That was the OP's point.
Honestly you're still looking at (from my understanding) ~3 minutes prefill (TTFT) even with architectural improvements and so on with a 32k context window (against a large model). How is this going to be competitive with Nvidia and all of the tricks massive scale get's you to parallelise context across many machines?
Is it supposed to be? I think the point with some of these Macs is you get the capability in something the size of a heatsink from Intel's Netburst architecture era, or a Macbook light enough to stick in a backpack and take with you to lunch.
If you're talking about chaining together multiple GPUs you're talking about a different game -- I suspect, anyway. Seems like a high-spec Mac would be good for development and testing. Arrays of GPUs, better aimed at production use.
The message I’m getting is that Apple will never compromise on its healthy margins. If something becomes basically unaffordable for their target market, they’d cut the production and even discontinue the product, than take a hit on margins. Their business model is refreshingly simple.
Yes - they create new/cheaper products for a different consumer but not at the cost of their margins. Vision Pro may have been the only device in recent history that likely had slim margins (if any).
Apple is not gonna risk their iphones, as they are their flagship (aside even from giving them higher margins). My opinion is that, as we are talking about ram SHORTAGE (not just for ram price hike) they have to cut the more ram hungry models to be able to keep up with their projected production/demand (at reasonable ram prices). Getting iphones "sold out" is not a great thing for apple.
Once/if the ram shortage ends, they will continue increasing the ram caps as they were already doing, because then selling ram-heavy macs will not interfere with the rest of their products.
It's just a planned economy failing the way planned economies often do: the central planner failed to predict the demand correctly. Instead of trying to secure additional stock from the market at spot prices, they are simply waiting for the next batches they had planned for.
I don’t think that represents the scenario at all, not to mention the fact that it’s literally not a planned economy (but also not very analogous to one, either).
What’s really happening is that the effort of securing additional stock isn’t worth it because the price is so high that there aren’t enough buyers.
If ground beef were to suddenly cost $50/pound, McDonald’s doesn’t raise the price of the Big Mac to $25 and hope people buy it, because it makes zero sense for their business model. Sure, some fancy restaurant will still be selling hamburgers, but not your chain of thousands of working class fast food restaurants. McDonald’s would find some other alternative item to sell.
The truth is that nobody’s going to be buying Mac Studios that cost $25,000. Not even enterprises.
Businesses are usually planned economies, and supply chain management is literal central planning.
Apple failed to predict the demand for Mac Studios. Many other companies in its supply chain likely failed to predict that Apple would come back asking for more. There is no excess stock for some key components or the spare capacity to make them on demand. Apple would have to scour them from the market, likely paying much higher prices than it will pay for scheduled deliveries.
Hopefully I’m not being too pedantic, but I am saying that we can’t just use the words “planned economy” or “central planning” to describe a single company’s actions.
It’s just not what those words mean.
Apple is a large company but they are still just one company.
A single business can’t be doing “central planning,” it is by definition not in charge of the whole market.
Central planning would be if the government mandated that Micron to make ## of memory chips and distribute ## of them to Company A and ## of them to Company B.
If they cared about local AI market they would price hike M3 Ultra instead of discontinuing them. After all they conviniently introduced RDMA just few months before that.
Initially when it happened everyone expected they did it because they planned to announce M5 Ultra shortly, but its not looks like this is happening.
Now IMHO its indicates they simply run out of RAM supply.
If you’re truly serious about local AI you’re not using a Mac. A Mac is for people who want the best of both worlds, GPUs are much faster if memory is equal.
It's even worse than that. Demand is so high so every next GB will be sligly more expensive. There is a lot of smaller hardware manufacturers that unable to secure DRAM chips at any price, there just no free capacity on market.
Of course Apple is massive, but if they announce inference boxes that everyone wants they need to make 10,000s or even 100,000s of them.
And it's very much likely they'd rather sell 600,000 - 10,000,000 more iPhones or Macbooks Neo. End users bring Apple money with every OpenAI / Claude subscriprion sold through their platform.
And inference boxes is just one-off sale of hardware that will bring no further income.
They are assuming that they are able to get ram in the future, once the AI bubble either dissipates or pops. Its far easier to build something you planned for 3 years ago, than crash build it in 3 months.
Of course if RAM prices crash Apple might of bring high RAM options back. We just cant bet on it as consumers.
Right now RAM shortages are bad to the point where likely even Apple have to decide what products they make and what they discontinue.
There been short time where M3 Ultra with 256GB / 512GB been best offet on market because Apple lagged with price increase. Now HN crowd expect Apple of all companies jump into price war with Nvidia and to subsidize their inference hardware.
768GB is 64-times of 12GB which is rumored to be amount of RAM in new iPhones. Imagine what profit margin 768GB Mac Studio gonna need in order to justify making one instead of 64 iPhones.
Apple is the company that is okay about selling microfiber cloth for $100 and wheels for $700. Imagine how bad price hike for M3 Ultra 256GB / 512GB had to be in order for them to just discontinue them instead of getting free money out of desperate local AI folks.