The infrastructure race has reached unprecedented scales: OpenAI plans to spend $750 billion by 2030, and Anthropic has already hit a $47 billion annual revenue run rate. However, a new problem is taking center stage—token economics and the profitability of AI agents in production.
Anthropic's Triumph and the ROI Paradox
Matt Murphy from Menlo Ventures notes that Anthropic has made an incredible leap: from $9 billion in 2025 to a $47 billion run rate (projected annual revenue) by May 2026. The investor emphasizes that the secret to this success lies not in the fundamental model itself, but in the business strategy.
Parallel to this, the orchestration segment is maturing. Writer has introduced a solution that reduces token spend by nearly 40% without sacrificing accuracy. This is a direct answer to the ROI (return on investment) problem: scaling AI by simply throwing more raw compute at production is economically unsustainable. In my view, cost-control tools, rather than new base models, will be the next billion-dollar market.
The Infrastructure Race and the Chip Bet
OpenAI has confirmed its unprecedented appetite: $750 billion in infrastructure investments through 2030. This sum is comparable to the GDP of an entire country (Sweden) and highlights the industry's shift from development to global deployment.
On the hardware side, NVIDIA is bringing the Vera Rubin system to market. Partners, including CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud, are already deploying server racks. The primary goal of the architecture is massive scale with minimal power consumption and lower token costs for partners. Meanwhile, in China, SenseTime is launching the Galaxy Project with 20 partners to scale domestic AI chips, aiming to reduce dependence on Western technology.
Lowering Agent Costs: New Google Models
LLM developers are responding to token inflation. Google has released Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber, specifically targeting reduced latency and lower operating costs for enterprise AI agents.
The absence of the expected Gemini 3.5 Pro and the focus on "workhorses" shows a clear understanding of demand. Developers have realized that the economics of autonomous agents in production require cheap, fast models, not an endless increase in parameters.
Geopolitics of Open Weights
Chinese open-weight models have come under intense scrutiny. The release of Kimi K3 from Moonshot AI has prompted corporations to rethink their approach to such solutions as Washington prepares potential sanctions.
The escalation was triggered by an incident involving the distillation (knowledge compression) of Anthropic's Fable model by Moonshot. The U.S. Treasury is threatening sanctions, forcing enterprises to seriously evaluate the regulatory risks of implementing Chinese open-source solutions.
Security and Market Restructuring
As investments grow, so does the cost of mistakes. OpenAI has taken responsibility for a breach of the Hugging Face platform, admitting the incident occurred due to a human error while configuring an isolated sandbox for model testing.
Endpoint security in the agent era is becoming a standalone market: the startup Glow has emerged from stealth with a $1.2 billion valuation. Simultaneously, a structural business overhaul is underway: Monday.com laid off 20% of its staff (approximately 630 people) to pivot to an AI-first model and launch its own AI platform.
Bottom Line
The infrastructure boom masks a critical pivot: the industry has moved from "who can build the smartest model" to "who can provide the cheapest and most secure inference." In the coming quarters, the winners won't be those training giant LLMs, but those who can lower the cost of AI agents in corporate environments and protect them from regulatory and technical risks.
Sources
- TechCrunch: Menlo Ventures’ Matt Murphy explains why Anthropic is winning
- VentureBeat: Writer's AI harness cuts token spend nearly 40%
- TechCrunch: OpenAI’s AI spending spree has ballooned to $750B
- AI News: Google’s Gemini 3.6 Flash targets enterprise agent token costs
- TechCrunch: Google releases three new Gemini models — but no 3.5 Pro
- NVIDIA Blog: NVIDIA Vera Rubin Driving Performance Per Watt
- AI News: SenseTime’s Galaxy Project targets domestic AI chip scale-up
- AI News: Chinese open-weight models are cheap. Washington is deciding what that costs
- TechCrunch: Treasury threatens sanctions after White House claims Moonshot distilled Anthropic’s Fable
- TechCrunch: How OpenAI’s human mistake led to the AI-powered hack on Hugging Face
- TechCrunch: Glow emerges from stealth at $1.2B valuation
- TechCrunch: Monday.com lays off hundreds to focus on AI