Two stories today are about AI infrastructure getting more expensive, and one is about open models getting more capable — all three cut toward the same question.
In Plain English: Two stories today are about AI infrastructure getting more expensive, and one is about open models getting more capable — all three cut toward the same question. Nvidia has told its biggest customers that next-generation AI servers will cost more than 15% more starting in early 2027, because the memory chips that go into them are in short supply. Separately, Alibaba raised $10.2 billion to keep funding its AI buildout, but investors sent the stock down more than 10% on dilution worries, even as the company's AI profits keep shrinking. On the other side of the ledger, DeepSeek released a new open-weight model that goes toe-to-toe with Claude Opus 4.8 on image-understanding tasks while cutting the computing cost of long prompts by nearly three-quarters. The standing Apple argument below is worth a fresh look given today's cost pressures.
Nvidia has told its biggest customers to expect prices on next-generation Grace Blackwell and Vera Rubin AI servers to rise more than 15% starting in early 2027, according to Bloomberg. The culprit isn't Nvidia's own chip pricing but a severe DRAM and high-bandwidth-memory shortage — sometimes called "RAMageddon" — that has pushed memory contract prices up 58–95% quarter over quarter, with memory now one of the largest line items in an AI server's bill of materials. It's a direct, concrete example of AI infrastructure costs rising even as the industry debates whether current spending will ever be profitable.
Tom's Hardware (Bloomberg) · Infrastructure & Economics · Aug 23, 2026
Alibaba priced an HK$80 billion (roughly $10.2 billion) Hong Kong share placement to fund what it calls "full-stack AI" — infrastructure, chips, cloud, and its Qwen model line. Demand was strong (about $28 billion in orders from sovereign wealth funds and long-only investors), but Alibaba's Hong Kong shares still fell as much as 10.5% on the announcement, as investors weighed dilution from 710 million new shares against results showing net profit down 75% year-over-year even as AI cloud revenue grew 45%. CEO Eddie Wu says the AI investment could break even within two to three years — a bet playing out in real time on exactly the profitability timeline Lucien's thesis questions.
South China Morning Post · Infrastructure & Economics · Aug 23, 2026
DeepSeek released V4-Flash-Vision-Exp, a 284-billion-parameter experimental model (13B active parameters via mixture-of-experts) that scores within striking distance of, and in some cases ahead of, Anthropic's Claude Opus 4.8 on image-understanding benchmarks like ZeroBench. More notable for the thesis: new compression techniques reportedly cut the computing power needed for long, million-token prompts by about 73%. It's currently gated behind DeepSeek's paid developer platform, but a free release is expected — another concrete step in open-weight models closing the gap with closed frontier systems while getting cheaper to run.
SiliconANGLE · Open-Weight Models · Aug 21, 2026
Today's cost headlines make a good backdrop for a case some commentators keep making: the company best positioned for a shift away from renting AI by the token may not be Nvidia, but Apple. The mechanism is unified memory — a single Mac Studio, priced around $9,500, can be configured with up to 512GB, enough to hold a trillion-parameter open-weight model on one desktop machine. Matching that capacity with Nvidia's RTX Pro 6000 workstation cards would take five or six of them, on the order of $60,000–$75,000 combined, while drawing roughly ten times the power — and that's before accounting for the 15%+ price hikes Nvidia just warned about, or the dilution Alibaba is absorbing to keep building at data-center scale.
The argument compounds with every open-weight release, today's DeepSeek vision model included: as frontier-caliber models from labs like DeepSeek, Alibaba, GLM, and Kimi keep shipping every few weeks, the bottleneck shifts from "who has the smartest model" to "who can afford the box to run it locally." Multi-Mac clusters using MLX and Thunderbolt 5's RDMA networking can already pool memory across several Studios for even larger models, and Nvidia's own DGX Spark quietly concedes the point by adopting a unified-memory design of its own — notably after Nvidia stripped NVLink pooling out of its consumer and workstation cards. The fair caveat still stands: a Mac Studio serves one user running one model locally, while a data-center GPU cluster serves many people at once — so this looks like Apple leading a specific, valuable segment, not the whole of AI compute.
Limited Edition Jonathan (Substack) · Perspective
Today's three items sit on both sides of Lucien's thesis at once. Nvidia's 15%+ server price hike and Alibaba's dilutive $10.2 billion raise — against a backdrop of AI profit that fell 75% even as AI cloud revenue grew 45% — are exactly the kind of rising-cost, uncertain-payback signals that raise doubts about whether current AI infrastructure investment ever turns profitable as memory, power, and capital keep getting more expensive. Meanwhile, DeepSeek's new open-weight vision model closing in on Claude Opus 4.8 while cutting compute costs by nearly three-quarters shows the capability gap continuing to narrow from the open-source side, even as it gets cheaper to run. Set against the standing Apple argument, the throughline holds: the pricier and more capital-intensive commercial AI infrastructure becomes, the more a one-time, unified-memory machine running an increasingly capable free model looks like the better bet.
The Daily: Open Source AI — a recurring research brief for Lucien Engelen
Keynotes, masterclasses, panels and board-room sessions.