Etched, a San Jose-based builder of transformer-only AI inference hardware, has raised $800M across multiple rounds including a $500M Series B at $5B valuation led by Stripes. Its Sohu ASIC hardcodes attention mechanisms into silicon for 10-20x better throughput and efficiency than NVIDIA GPUs on dense transformer inference. The capital supports scaling manufacturing and first customer shipments.
Inference Specialization Gains Momentum
Etched's emergence coincides with a broader shift toward specialized inference silicon. Groq was acquired by NVIDIA for $20B in December 2025, while Cerebras completed an IPO at a $56B valuation in May 2026. Etched's fixed-function approach targets only dense transformers, carving out performance gains unavailable to programmable alternatives.
Transformer Inference Demands Efficiency
AI inference workloads now consume more compute than training at frontier labs. General-purpose GPUs leave substantial efficiency on the table for attention-heavy models. Etched's architecture eliminates programmability overhead to sustain higher FLOPs density and lower power draw on Llama-scale workloads.
Fixed-Function ASIC Targets Dense Models
Sohu runs exclusively on TSMC's N4P process with 144GB HBM3E per chip and proprietary Low-Voltage Inference and Cluster-Scale Memory techniques. An 8-chip server claims 500,000 tokens per second on Llama 70B at batch size 1. Unlike reconfigurable designs from SambaNova or Tenstorrent, the chip cannot execute MoE, diffusion, or training workloads.
As co-founder Robert Wachen noted:
"Production is the product."
Stripes Leads $500M Growth Round
Stripes led the December 2025 round alongside Positive Sum, Ribbit Capital, and Radical Ventures. Strategic backing from VentureTech Alliance (TSMC-linked) and trading firms including Jane Street validates both manufacturing readiness and demand from quantitative finance users.
AI Chip Market Expands Rapidly
The AI inference market stands at $106.15B in 2025 and is projected to reach $254.98B by 2030 at 19.2% CAGR. ASIC share of inference accelerators is expected to grow from 15% to 40% by 2026. Etched joins Cerebras, SambaNova, and Tenstorrent in challenging NVIDIA's dominance while accepting narrower architectural scope.
Harvard Dropouts Build Elite Team
Founders Gavin Uberti, Robert Wachen, and Chris Zhu are Harvard dropouts and Thiel Fellows. The 400-person team includes 22-year NVIDIA veteran Brian Loiler and former Google TPU software lead David Munday. Advisors include Geoffrey Hinton, Andrej Karpathy, and Fei-Fei Li.
First Racks Ship Summer 2026
Etched has already validated A0 silicon and signed over $1B in customer contracts. Initial rack shipments begin this summer with a path to gigawatt-scale deployments in 2027.
