Per Million Tokens Fall Below $1! Goldman Sachs: "Volume Up, Price Crash" Shakes the Logic of AI Capital Expenditures
Goldman Sachs' latest report shows that the SDLLMTK index, which measures AI inference pricing, plunged 29% in August alone, falling below the $1 per million token mark for the first time and having been cut in half from its peak. Factors such as price cuts, user migration to low-cost open-source models, intensifying competition, and local inference are structurally eroding the per-token pricing model. The "volume up, price down" divergence is undermining the core narrative that "more tokens equal more revenue."
The core narrative of AI infrastructure investment is facing severe challenges at the unit economics level. The collapse in Token prices has outpaced previous market expectations, fundamentally shaking the revenue logic that has supported AI sector valuations over the past two years.
According to the latest report from Rich Privorotsky, head of Goldman Sachs One-Delta trading desk, the Silicon Data LLM Token Expenditure Index (code: SDLLMTK), which measures the market’s weighted average price per million Tokens used, fell by 29% in August alone, ending the month at around $0.97—a record low. The cumulative drop from the May peak of about $2.05 has now exceeded 50%. This is the first time the index has fallen below the $1 mark per million Tokens.

Meanwhile, according to a JPMorgan data center report, Token usage on the OpenRouter platform surged about 47% month-on-month in August, but dollar expenditures grew by only about 7%. This stark divergence between rising volumes and falling prices reveals the core contradiction of the current AI investment narrative: returns in the stock market are calculated in dollars, not in Token quantities, and capital spending on hyperscale data centers is accounted for based on expected dollar revenues.
Privorotsky emphasized in the report, "More demand plus lower prices—if prices are falling faster than consumption is growing, it is not automatically a positive."
Token Price Collapse: It's Not About Vanishing Demand, But a Breakdown in Unit Pricing
The Silicon Data LLM Token Expenditure Index does not measure Token demand, but is a usage-weighted price index. The company itself has warned the market that the index’s decline may be due to listed price reductions, user migration to lower-cost open-source models, or both; it does not mean that AI usage is contracting.
However, this is precisely the issue. OpenRouter routing volumes surged, while H100 GPU rental prices have declined—contradicting the "Token demand surge" narrative. Privorotsky noted that August’s data showed a clear pattern: volumes were up significantly, prices fell even more, and dollar spending barely moved.
Goldman Sachs concludes: The equity pricing of the past two years was based on the assumption that "more Tokens equals more revenue," and August data clearly broke this assumption for the first time.
Intensifying Competition and Local Inference: Structural Erosion of the Per-Token Pricing Model
Privorotsky bluntly stated in the report that he questions whether "charging by the Token" is a sustainable business model. The cloud-based inference per-million-Token pricing logic relies on three assumptions being true at the same time: models are too large to run locally, users cannot replace them with open-source alternatives, and workloads bursty enough that building in-house compute is uneconomical. Today, all three assumptions are failing simultaneously.
On the hardware end, RTX Spark-level laptops, DGX Spark workstations, Mac Studios running 70-billion parameter models, and NPUs with compute in the 40 to 75+ TOPS range have pushed the marginal cost of many Tokens down to near the cost of electricity. Once a company’s monthly API bill exceeds the depreciation of a $5,000–$15,000 device, that company shifts from being a Token consumer to a one-time hardware purchaser, rather than a subscriber contributing a 40% software margin on an ongoing basis.
On the model competition front, Meta’s Muse Spark 1.3, which began rolling out on September 2, is now in the same league as GPT-5.6 Sol and Claude Opus 5 in independent benchmarks, and is particularly strong in agent tasks and code generation. Privorotsky notes, "Frontier advantages are measured in weeks—they are not moats, just product cycles." Whenever a suboptimal model becomes sufficiently close in capability and cheaper, enterprises will bypass the premium tier, and the Token price index is precisely the aggregate reflection of this routing logic.
Capital Spending Logic Under Pressure: Credit Markets Are Issuing Early Warnings
Goldman Sachs reports that the current cycle of AI capital spending is not a flexible operating expense, but rather a set of locked-in, heavy asset commitments. According to the report, the five rated hyperscale cloud vendors are expected to spend a combined $737 billion in capex by 2026, about 38% of revenues. Moody’s has already warned of compressed free cash flow, a shift from asset-light to asset-heavy balance sheets, and off-balance-sheet leasing commitments that, while not reflected in bonds, nonetheless materially limit issuers.
The reaction in the credit markets has preceded the stock market. According to a survey of relevant bonds, of the 91 hyperscale cloud vendor bonds maturing in 2026, 78 had already fallen below issue price by the end of August. Privorotsky refers to this as "credit repricing" and points out that equity valuation multiples are lagging indicators.
The logic chain is clear: Token prices drop by 30%, usage rises by 20%, inference revenue falls; inference revenue falls while depreciation and interest on the 2025–2027 construction wave are still ramping up, so ROI collapses; once ROI collapses, there’s no need for dramatic "bubble bursts"—the market simply adjusts the valuation multiples to match a utility enterprise with large assets but limited pricing power.
GPT-6 Astra: The Only Remaining Narrative Reversal Variable
In the report, Privorotsky identifies OpenAI’s next-generation model, Astra, as the only likely near-term catalyst for a reversal in the current situation. On September 1, OpenAI announced Astra had reached a critical network security threshold under its Preparedness Framework, becoming the first model to be classified in this category.
However, Goldman Sachs’ trading desk remains cautious. A highly restricted, monitored model with significant access friction does not automatically add to the Token revenue pool. If high-value workloads remain in controlled testing environments and never enter the public billing system, Astra could actually reduce billable Token volumes.
The report’s conclusion: Astra is the final near-term variable with the potential to shift the revenue mix back toward high-value tiers. Until then, the August Token price data is the truest signal of the current market state—demand may grow infinitely, but if prices at that scale cannot cover the cost of debt issued to build data centers, share prices can still fall.
Disclaimer: The content of this article solely reflects the author's opinion and does not represent the platform in any capacity. This article is not intended to serve as a reference for making investment decisions.
You may also like
LINK Could Be Ready for Another Breakout as Bullish Pennant Targets $27



Rain Protocol settles its first DAO with 7.4 billion $RAIN permanently burned

