NIGHTLY INTELLIGENCE BRIEF
〔Day Digest〕OpenAI Hits $40B Run Rate, GLM-5.3 Bets on Post-Training Scaling, While a Price War Breaks Out: Google Discounts Flash 50%, 'Cost-Per-Task' Replaces Token Math
OpenAI's annualized revenue run rate topped $40 billion, doubling 2025, while Zhipu AI's GLM-5.3 used post-training scaling to jump Terminal-Bench scores from 4.6 to 28.3 and DeepSWE from 46.2 to 66.9, but the pricing war is reshuffling value: Google's Gemini 3.7 Flash runs a 50% discount through year-end, OpenAI and Anthropic are pushing 'cost-per-task' as the new metric, and Silicon Data shows enterprise spend on top US labs down nearly 25% since mid-July. SK Hynix warns of the worst memory shortage next year and plans $720B capacity expansion, while Nvidia's CPO switches enter mass production. Whether cost-per-task becomes the standard — and how fast Chinese models close the gap — will decide the next leg.
0. Weekly Arc
The AI session revolved around two poles: OpenAI’s annualized revenue run rate crossing $40 billion — double 2025 levels — and Zhipu AI’s GLM-5.3 showing that post-training scaling, not parameter count, can move benchmarks [1][2]. Yet the pricing war is deepening: Google is selling Gemini 3.7 Flash at 50% off through year-end, while OpenAI and Anthropic push a new “cost-per-task” metric to justify premiums [3][4]. The decisive question is whether token-based pricing collapses before enterprise budgets fully shift.
1. Model Race: Post-Training Scaling vs. Bigger Bases
- **[NEW] GLM-5.3 (Zhipu AI)** keeps the 753B-parameter base unchanged; gains come from longer training, richer task environments, and larger-scale reinforcement learning — a method Zhipu calls “post-training Scaling” [1]. Founder Tang Jie had teased an “epic plus” model [1].
- Benchmarks: Terminal-Bench 3.0 jumped from 4.6 to 28.3; DeepSWE v1.1 from 46.2 to 66.9; top open-source on both [1]. GDPval-AA v2: 1769 vs GLM-5.2’s 1508 [1]. The model emphasizes coding and cybersecurity [5].
- Distribution: JD Cloud has integrated GLM-5.3 on its MaaS platform [6]; Kingsoft Office’s Lingxi Professional is first to connect it, improving code-driven office workflows [7].
- Surrounding moves: DeepSeek released V4 Pro and the Harness tool, Manus resumed independent operations, and Tencent preannounced a larger Hy4 model [8]. Apple is reportedly co-developing a China-specific LLM with Alibaba — a first for a foreign company — with Apple Intelligence to land within months [9].
2. Pricing War: Token Math Under Siege
- OpenAI and Anthropic are shifting the pricing narrative from per-token cost to “cost-per-task” [3]. OpenAI CFO Sarah Friar: cheap tokens may need more attempts, time, or human review [3]. Anthropic’s Jonathan Pelosi: token cost is only a “proxy” for compute consumed; better to measure real task completion cost [3].
- Ramp data: enterprise adoption of OpenAI/Anthropic has slowed; clients are moving to cheaper open-source models; Anthropic’s most expensive Fable 5 accounts for just 11% of Claude software spend [3].
- Google counter-punches: Gemini 3.7 Flash is priced at $0.75/M input and $3.75/M output until Dec 31 — 50% below the standard $1.5/$7.5 — then reverts on Jan 1, 2027 [4]. Google DeepMind SVP Koray Kavukcuoglu calls it the latest main model for coding and agentic workflows [4].
- DeepSeek raised prices (without specifics) [4]. Silicon Data: since mid-July, actual enterprise spending on US top labs is down nearly 25% as Chinese models gain share [10].
- Real-world adoption: FAW-Volkswagen’s ID. AURA T6 opened blind booking with integration of ByteDance’s Doubao model [11].
3. Infrastructure and Capital: The Build-Out Accelerates
- Nvidia’s Spectrum-X Ethernet photonic switches are in full mass production, using co-packaged optics; Nvidia claims 5x network power efficiency, 5x AI app uptime, and 10x mean time between failures vs pluggable optics [12]. Supply chain: TSMC (silicon photonics), SPIL (packaging/test), Lumentum and TFC (lasers), Foxconn (system assembly) [12].
- AMD raised $4.75B in bonds — a company record — settling Aug 17, to fund AI, data centers, and manufacturing; revenue is expected to grow 47% in 2026 to over $51B [13].
- Goldman Sachs is courting investors for Nvidia’s $500B AI-compute financing deal [14].
- SK Group Chairman Chey Tae-won warns of the worst “memory chip shortage” next year, calling high-end memory a “war”; customers are asking for nearly 2x supply. SK Hynix plans a $720B investment to triple capacity by 2034, has 10 long-term supply pacts, and a $500B Nvidia pact covering HBM, next-gen memory, and a joint data center by 2027 [15]. Korean equities rallied: KOSPI up 11.5% on the week, Samsung +18.83%, SK Hynix +15.68% [15].
- LG and Nvidia signed an MOU for a bipedal humanoid robot with Nvidia’s Jetson Thor in Q1 2027; six LG subsidiaries will collaborate [16].
- Server makers are seeing orders convert: Lenovo ISG revenue was $8.5B, up 98%, with AI server potential orders at $54B, up >150% QoQ; Dell posted record revenue/profit and its stock popped ~40% after hours [17].
4. Monetization and Corporate Moves
- OpenAI’s annualized revenue run rate has topped $40B, doubling 2025; July run rate grew over 20% m/m, per president Greg Brockman; CFO Sarah Friar had put last year at $20B+ [2][18]. Growth is led by AI coding, subscriptions, and emerging ads [18]. IPO may slip to next year as some investors worry about cash burn; Anthropic filed confidentially in June and could list as soon as autumn [18].
- Baidu renamed its GenFlow agent “Kuku AI”, with AI-office MAU above 25M; GenFlow itself crossed 100M MAU in April; new standalone PC/web/mini-program/enterprise apps launched [19][20].
- Douyin (ByteDance) invested in household-robot maker Weilai Buyuan, lifting registered capital to 7.6914M yuan [21].
- A new robotics venture, Shanghai Xingyi Qingkun, was set up with 61.2M yuan registered capital by Shenzhen Zongqing Robotics and Shanghai Aoyi, covering AI software, robot sales, and unmanned aerial vehicle manufacturing [22].
- (A thin single-headline [23] hints a SpaceX-Nvidia alliance could extend to orbit; no further detail.)
SOURCE TRAIL
Citations
23 records
-
[1]
第一财经 · 新闻不卷万亿参数,智谱GLM-5.3侧重编程能力与网络安全 ↗
- [2]
- [3]
-
[4]
第一财经 · 新闻DeepSeek涨价,谷歌降价:新模型限时五折 ↗
-
[5]
财新(Google News 聚合)智谱发布新模型GLM-5.3 强调模型编程与网络安全能力 - 财新 ↗
- [6]
-
[7]
36氪 · 快讯灵犀专业版首发接入GLM-5.3 ↗
- [8]
-
[9]
格隆汇 · 财经动态据报苹果联合阿里自研中国专属大模型 ↗
-
[10]
财联社 · 电报高性价比中国大模型兵临城下 美头部AI实验室大幅降价应对 ↗
-
[11]
同花顺 · 7×24 直播一汽-大众ID. AURA T6接入豆包大模型 ↗
- [12]
-
[13]
36氪 · 快讯AMD发行47.5亿美元债券,创最高规模发债纪录 ↗
- [14]
-
[15]
格隆汇 · 财经动态巨头挤爆韩国“抢芯”,SK掌门人:明年“存储荒”最严重! ↗
- [16]
-
[17]
第一财经 · 新闻AI服务器订单火爆:联想、戴尔、工业富联吃到AI基建红利 ↗
-
[18]
格隆汇 · 财经动态年化收入破400亿美元!OpenAI备战IPO底气大增 ↗
- [19]
- [20]
-
[21]
同花顺 · 7×24 直播抖音入股家庭通用机器人公司未来不远 ↗
-
[22]
36氪 · 快讯众擎机器人等成立科技公司,智能无人飞行器制造业务 ↗
- [23]