Inferact: Building the Infrastructure That Runs Modern AI

a16z Deep Dives · 2026-01-22 · 44 min
https://www.youtube.com/watch?v=GsRnarLIC9gVideo summary
Inferact’s vLLM founders aim to build an open inference layer for models running across 400,000–500,000 GPUs.
Inferact co-founders Simon Mo and Wau Kuan traced vLLM to a 2022 UC Berkeley effort to speed up Meta’s 175-billion-parameter OPT model. Its distinctive challenge was scheduling constantly changing language-model requests and managing KV-cache memory, leading to the PagedAttention work. The open-source project has grown to more than 2,000 contributors, with 50 or more regular full-time contributors and a CI bill exceeding $100,000 monthly. vLLM now runs on an estimated 400,000–500,000 GPUs around the clock, across varied chips and models. The founders say inference is getting harder as models scale beyond a trillion parameters, architectures and hardware diversify, and agents add long-running tool interactions that complicate cache management. They founded Inferact to support vLLM and build a universal inference runtime for new models, hardware, and applications; open source, they say, is the company’s top priority.
Chapters
- 0:00Introduction: Inferact’s Universal Inference Layer and vLLM’s OPT Origins
- 6:01Introduction: Dynamic LLM Scheduling and the First vLLM Meetup
- 11:41Community and Collaboration in vLLM: 2,000+ Contributors and $100K-Plus Monthly CI
- 19:19Understanding Inference Engines: Components and the Multi-Trillion-Parameter Trend
- 24:27Cluster Scale and GPU Deployment: Up to 500,000 GPUs and Agent-Driven Cache Challenges
- 31:19Belief in Open Source AI: Amazon Rufus and Character.AI’s vLLM Deployments
- 35:45Founding of Inferact: vLLM’s Universal Inference Layer and Yang as Co-founder
- 40:00Future of Inference at Scale: GB200/GB300 NVL72 and a Universal Runtime
This is a Tier 1 public summary
Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.
More from this channel
To Regulate AI Effectively, Focus on How It’s Useda16z Deep DivesMartin Casado argues AI laws should target harmful uses, while regulatory uncertainty pushes US startups toward Chinese open-source models.
Mintlify and the Transition From Human Docs to Agent Infrastructurea16z Deep DivesMintlify grew from eight pivots and a two-day prototype into documentation infrastructure serving AI agents and 20 million monthly visitors.
How Palantir Scaled: Why the Best Software Is Built Backwardsa16z Deep DivesPalantir architect Akshay Krishnaswamy explains how field engineers turn customer pain into products and scale software backwards.
Temporal CEO on AI Agents & The Future of Software | Deep Dives with a16za16z Deep DivesTemporal CEO Samar Abbas says durable execution will underpin long-running AI agents, with cloud handling 150,000 actions per second.
Braintrust CEO on Where Engineering Actually Matters in AIa16z Deep DivesBraintrust CEO Ankur Goyal says AI engineering should focus on evals and harnesses, while SQL beat Bash in his agent benchmark.
Related analyses
New Grad求職有多困難?矽谷大裁員、AI取代工程師恐慌下靠什麼進入OpenAI、Google、LinkedIn、Salesforce哈佛姐夢遊矽谷 AliceInSiliconWonderlandBerkeley New Grad 求職率約三成,四位來賓分享靠 GitHub 職缺、校友推薦與真人模擬面試進入 OpenAI、Google 等公司。
Yikes.ColdFusionColdFusion traces Claude’s role in military targeting, Anthropic’s refusal of surveillance terms, and OpenAI’s controversial Pentagon deal.
How The Internet’s Favourite AI Employee Went RogueColdFusionOpenClaw promised a capable AI assistant but exposed private data, compromised 4,000 developer machines, and helped drive a $1 billion mortgage fraud investigation.
EP313. 谷歌痛失 AI 大將、Fable 5 配方解密、Meta 士氣低落中、FSD 即將入台 | M觀點M觀點Google 痛失 Gemini 共同負責人 Noam Shazeer 與諾貝爾獎得主 John Jumper,Fable 5 靠 Agent 流程變強,Meta 裁員後士氣低落。
EP315. GPT-5.6 限定推出、蘋果漲價被罵翻、Meta 新智能眼鏡 | M觀點M觀點GPT-5.6 在代理工作與資安測試勝過 Mythos;蘋果因記憶體漲價調高多款產品售價,Meta 推平價智慧眼鏡。