One story today, but it's a meaningful one for the "will this actually be profitable" half of the thesis.
In Plain English: One story today, but it's a meaningful one for the "will this actually be profitable" half of the thesis. The going price for AI companies to process a million words of text has fallen below $1 for the first time ever, according to an industry tracker — the latest drop in a price war that keeps squeezing everyone's margins. One Chinese lab, DeepSeek, went the opposite way and raised its prices sharply last month, hinting the market may be splitting into a cheap commodity tier and a pricier premium one. Both OpenAI and Anthropic have also quietly filed paperwork to go public, which raises the pressure to show real, durable profit rather than growth alone. It was a quiet news day overall (a US holiday), so today's Apple Angle does some extra work explaining why cheaper tokens don't necessarily settle the argument.
A widely-cited industry price tracker, the Silicon Data LLM Token Expenditure Index, dropped below $1 per million tokens for the first time — down to $0.97, an 8.6% decline in a single week — after Anthropic cut cache-read pricing on its new Fable 5.1 model by 75% (from $1.00 to $0.25 per million tokens) on September 1, following Google's and Meta's price moves two days later. Analysts describe a market splitting into two tracks: commodity-priced standard models racing toward zero, and a separate, gated premium tier (Google's new cybersecurity-focused model variant, for instance) still commanding real prices. Notably, DeepSeek broke from the pack and raised its own prices by as much as 14x in mid-August, and both OpenAI and Anthropic have filed confidential paperwork for IPOs — a step that typically increases pressure to show sustainable, not just subsidized, profit. For Lucien's thesis, this is double-edged: falling headline token prices look like exactly the margin compression the "infrastructure investment won't pay off" argument predicts, but a price war this aggressive is also, by definition, a bet that current prices are not what will stick — someone has to eventually charge enough to cover the compute that produced the tokens.
Forkast News · Infrastructure & Economics · Sep 5, 2026
Today's news is a fair moment to stress-test the standing argument rather than just restate it: if renting tokens is getting cheaper by the week, why would anyone bother buying $9,500 worth of local hardware? The honest answer is that the two numbers aren't really comparable. A price under $1 per million tokens describes light, chat-style usage of a standard-tier model — it says nothing about the agentic workloads that chew through tokens by the order of magnitude, or about the "premium, gated" tier the pricing article itself flags as exempt from the race to zero. Meanwhile, the hardware math hasn't moved: a Mac Studio configured with up to 512GB of unified memory, shared between CPU and GPU, still costs around $9,500 and still holds a trillion-parameter open-weight model entirely in memory, with zero further per-token billing regardless of how heavily it's used. Reaching an equivalent memory footprint with Nvidia's RTX Pro 6000 workstation cards takes five or six of them, somewhere around $60,000–$75,000 combined, at roughly ten times the power draw of one Mac Studio.
There's also a subtler point in today's story: a price war this sharp, arriving right as both OpenAI and Anthropic have filed confidentially to go public, is not obviously a sign of a healthy, durable business — it's what happens when providers compete for usage before profitability is proven, and DeepSeek's decision to move against the trend and raise prices 14x suggests at least one major lab thinks the going rate is unsustainable. As open-weight models keep shipping every few weeks at frontier-adjacent capability, and Apple's MLX framework plus Thunderbolt 5's RDMA networking let multiple Mac Studios pool memory for even bigger jobs, the calculation for a heavy or steady user increasingly isn't "what's the current promotional token price" but "what happens to my bill once the promotions end" — a question a fixed-cost local machine simply doesn't have to answer twice. The fair caveat stands as always: a Mac Studio serves one person running one model at a time, while the data centers behind today's price war are built to serve many thousands of concurrent users — Apple looks positioned to dominate a specific, high-value slice of AI compute, not to replace the data center outright.
Limited Edition Jonathan (Substack) · Perspective
Today's single item cuts right to the center of the profitability half of Lucien's thesis: the first-ever sub-$1-per-million-token price point is real evidence of margin compression, but the way it arrived — a 75% single-day price cut from Anthropic, both OpenAI and Anthropic quietly filing to go public, and DeepSeek pointedly moving the opposite direction — reads less like a market that has found its stable price and more like one still figuring out what these tokens actually cost to produce. Nothing here proves the on-prem transition has arrived on its own, but it sharpens the same question the thesis keeps returning to: when the promotional pricing of the current API price war eventually meets the real cost of the compute behind it, does the math still favor renting, or does it favor owning — which is exactly the bet the standing Apple Angle argues Apple is already quietly positioned to win.
The Daily: Open Source AI — a recurring research brief for Lucien Engelen
Keynotes, masterclasses, panels and board-room sessions.