Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

TSMC accumulated 1 billion chips in backlog due to capacity crunch; DeepSeek raised inference API pricing amid supply shortage.

Acute manufacturing bottleneck cascading to end customers; pricing pressure on AI services signals critical foundry capacity exhaustion
Trade pressSlicast · August 7, 2026 · China · Source: Google News
importance 94

On August 6, Huawei Fellow and semiconductor chief scientist Liao Heng stated in an interview: "Currently, chips are almost everywhere. One could say they have become as important as 'water' and 'air.'" The critical importance of the domestic chip industry chain targeted by the Science and Technology Innovation Chip ETF Huian Addax (588750) is evident.

Especially in the context of the AI wave, chips face unprecedented massive demand. Compute shortages, chip scarcity, and chip inflation have become industry norms. DeepSeek's rare price announcement and Anthropic's announcement of self-developed chips further confirm this trend.

Looking forward, as domestic large models accelerate iteration and the "domestic models with domestic chips" trend progresses, innovation breakthroughs in domestic chips—including 3D stacking and advanced nodes—position domestic chips to leapfrog forward, entering a golden development period of volume and price expansion.

According to industry sources, TSMC, ramping high-end capacity to meet Apple's A20 Pro processor demand, has accumulated Apple processors worth $1 billion awaiting final packaging due to DRAM shortages. It must wait for DRAM delivery to complete production.

On August 6, DeepSeek released a notice stating it plans to implement an overall price increase for DeepSeek API services in the near term, with significant increases expected. Multiple factors drive this move:

First, peak-hour call volumes and concurrent access surge continuously; compute is severely constrained. OpenRouter's weekly rankings show V4-Flash calls at 7.22 trillion tokens leading the list. The open-source project OpenCode disclosed that on August 1 alone, daily usage reached 8 trillion tokens. On August 4, "unprecedented traffic" caused capacity insufficiency.

Second, inference cost structures (GPU/HBM/electricity/operations) rise during demand explosions. Previously, "extremely low prices plus high volume" eroded margins, necessitating a return to "sustainable cash flow plus value-based pricing."

Third, industry-wide price floor increases—domestic cloud vendors and multiple model providers have raised prices multiple times this year. DeepSeek's increase shifts the "low-price anchor" upward, transitioning the industry from "competing on low price" to "paying for performance/stability/ROI."

Short-term, this aids load-balancing and stabilizes core business. Medium-term, the upward price shift repairs model-side cash flow, feeding back to upstream compute and infrastructure, driving a positive flywheel of "compute rental—domestic compute—applications."

AI large models have driven training and inference compute demand to explosive levels. To meet massive chip demand and ensure supply stability, many AI leaders have entered the "self-developed chips" race.

According to media reports, AI startup Anthropic for the first time publicly confirmed it will establish a dedicated internal team to customize AI chips for its Claude large model, formally implementing its self-developed chip strategy. This leverages custom hardware to address compute shortages, fully adapting to Claude's continuously escalating large-scale training and inference demands, solidifying the hardware foundation for scaled deployment. Similarly, OpenAI previously released its first self-developed AI inference chip, Jalapeño.

In the AI wave, how massive is chip demand? Goldman Sachs projects global AI accelerator chip demand for 2026-2028 at 19 million, 27 million, and 32 million units respectively—14%-22% growth over prior forecasts.

Policy-wise, computing power network construction is accelerating, enabling efficient cross-region and cross-industry resource dispatch. As one of the "six networks" (water, power grid, computing power, next-generation communications, urban underground infrastructure, and logistics), the computing power network aims to unify nationwide computing resource scheduling, reducing per-token costs and functioning as foundational infrastructure similar to the "national grid." It directly drives investment while spurring industry-wide upgrades in AI, big data, and industrial internet through multiplier effects. In the first half of 2026, China's AI independent innovation pace significantly accelerated: the first fully domestic 100,000-card cluster was deployed; intelligent computing capacity reached 2.8 times year-over-year levels from the prior year; domestic large model downloads exceeded 10 billion; and rapid adaptation with domestic computing chips accelerated. (Source: Galaxy Securities 20260805《Industry Weekly Report: Computing Power Network Construction Accelerates, Domestic AI Full Chain Speeds Up》)

From the AI model angle, domestic large models concentrate iteration efforts, and domestic computing power logic continues strengthening. Zhongyin Securities notes that recently DeepSeek, ByteDance Seedance, MiniMax, and Moonshot undertook dense upgrades: DeepSeek-V4-Flash officially launched API open beta with Codex adaptation; ByteDance Seedance 2.5, Xiyú MiniMax H3, and Kimi K3 achieved simultaneous breakthroughs in video generation, multimodal capabilities, and parameter scale. Model development has shifted from "parameter stacking and Q&A" toward Agent capabilities, lower inference costs, and multi-model coordination, potentially pushing AI calls from low-frequency Q&A toward high-frequency, complex, multi-step tasks. Price reduction does not necessarily suppress total compute demand—if call volume growth exceeds unit token cost reduction, combined with Agent multi-round planning and tool invocation driving higher token consumption, inference compute demand still has upside potential.

As domestic model competitiveness improves and applications penetrate enterprise production workflows, beneficiaries expand from GPUs to CPUs, AI servers, storage, switches, optical modules, PCBs, liquid cooling, data centers, and power infrastructure. The favorable logic of the full domestic computing power chain continues strengthening. (Source: Zhongyin International Securities 2026-08-05《Computer Industry Weekly Decoding: Domestic Models + Domestic Computing Power Converge, AI Applications + Domestic Computing Power Hardware Continue to Benefit》)

Technology-wise, super nodes drive computing power upgrades. According to IT Home, Tencent Cloud plans large-scale domestic computing power deployment and super-node deployment in Q4 2026. Tencent Cloud views optical interconnect as critical for overcoming bandwidth and power bottlenecks as super-node scale expands. NPO represents a pragmatic path for domestic GPUs building super nodes. Compared to CPO, NPO places optoelectronic conversion chips on the same board as computing chips, shortening distances to main chips and eliminating DSP chips from optical modules.

WAIC 2026 shows domestic GPU competition has shifted from single-card performance to systemic competition across super nodes, clusters, software ecosystems, and application deployment. Birentech, Mthreads, Enflame, Muxi, and Lightningstar concentrated on demonstrating large-scale interconnection, backplaneless designs, general-purpose GPUs, open-source ecosystems, and commercialization achievements. Domestic computing power has progressed from "do we have it" to "can we use it well," with industry focus shifting from "making chips" to "using chips well." (Source: CITIC Jiantou 20260725《Weekly Perspective: Domestic Computing Power Supply-Demand Resonance》)

The AI infrastructure super-cycle remains promising, with storage, CPU, wafers, and advanced packaging supply-demand gaps all exceeding expectations. The Science and Technology Innovation Chip ETF Huian Addax (588750) index contains 86% computing power (CPU/GPU/ASIC) plus storage plus advanced manufacturing—highly focused on pure-play chip opportunities. 20CM price swings easily capture high-momentum track growth. Off-exchange investors can track linked funds (A: 020628; C: 020629), available for 7*24 trading.

For more precise domestic computing power and storage positioning, consider the Science and Technology Innovation Chip Design ETF Huian Addax (589290), with 72% computing power (CPU/GPU/ASIC) plus storage content. Under triple resonance—AI compute demand explosion, storage chip price cycle, and domestic substitution—it should fully benefit from historic opportunities in domestic computing power and storage chains amid the AI wave and self-reliance imperative. 20CM swings help easily capture this super-cycle.

Read the original
TSMC accumulated 1 billion chips in backlog… · Slicast