Inferact: Building the Infrastructure That Runs Modern AI

a16z Deep Dives · 2026-01-22 · 44 min
https://www.youtube.com/watch?v=GsRnarLIC9gVideo summary
Inferact founders explain vLLM’s rise, 500,000-GPU scale, and a universal open-source inference layer
Inferact’s founders trace vLLM from a slow 2022 Meta OPT demo at UC Berkeley into a major open-source inference engine. Its core breakthroughs addressed dynamic LLM workloads through scheduling, memory management, and PagedAttention, unlike static image-model batching. The project now has more than 2,000 GitHub contributors, over 50 regular full-time contributors, and runs on an estimated 400,000–500,000 GPUs continuously. Amazon uses vLLM for its Rufus assistant, while LinkedIn and Character AI have adopted cutting-edge features early. The founders say inference is becoming harder as models exceed one trillion parameters, hardware and architectures diversify, and agents create long-lived, tool-using sessions. Their new company, Inferact, aims to steward vLLM and build a universal runtime spanning models, chips, applications, and deployment environments.
Chapters
- 0:00Introduction: Inferact’s Universal Open-Source Inference Layer
- 6:01Introduction: From OPT’s Slow Demo to vLLM’s Scheduling Breakthrough
- 11:41Community and Collaboration in vLLM: 2,000 Contributors and $100K Monthly CI
- 19:19Understanding Inference Engines: Tokenized Output, Scheduling, and KV Cache
- 24:27Cluster Scale and GPU Deployment: 400K–500K GPUs and Agentic Cache Challenges
- 31:19Belief in Open Source AI: Diversity Across Models, Chips, and Use Cases
- 35:45Founding of Inferact: A Universal Inference Layer Built Around vLLM
- 40:00Future of Inference at Scale: Universal Runtime for GB200 and GB300 Systems
This is a Tier 1 public summary
Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.
More from this channel
Mintlify and the Transition From Human Docs to Agent Infrastructurea16z Deep DivesMintlify’s Han Wang and Hahnbee Lee explain eight pivots, a two-day prototype, and docs becoming AI infrastructure.
To Regulate AI Effectively, Focus on How It’s Useda16z Deep DivesMartin Casado argues AI laws should target illegal use, while US regulatory uncertainty pushes startups toward Chinese open-source models.
How Palantir Scaled: Why the Best Software Is Built Backwardsa16z Deep DivesPalantir’s Akshay Krishnaswamy explains FDEs, backward product building, Foundry, and avoiding the consultancy trap.
AI Copilots Are a Dead End. Here's What Actually Works | Kavak CEOa16z Deep DivesKavak uses AI agents for 90–95% of customer interactions after a flat 2023, then grows four times.
Temporal CEO on AI Agents & The Future of Software | Deep Dives with a16za16z Deep DivesTemporal powers OpenAI Codex and Snap at scale, bringing durable state, recovery, and RPC to long-running AI agents.
Related analyses
2026/8/27(四)輝達財報再超標!追高意願低?美股為何原地踏步?【早晨財經速解讀】游庭皓的財經皓角英偉達營收年增106%、資料中心占92%,AI需求強勁但表外承諾與資料中心延誤成最大風險
Rebuilding Git for AI Agents and The Future of Developer Tools | Deep Dives with a16za16z Deep DivesGitButler rethinks Git for AI agents with parallel branches, agent-ready CLIs, and patch-based code review
EP313. 谷歌痛失 AI 大將、Fable 5 配方解密、Meta 士氣低落中、FSD 即將入台 | M觀點M觀點Google痛失兩名AI大將,Anthropic揭露Fable 5靠Agent Harness變強,Tesla FSD送件台灣
EP315. GPT-5.6 限定推出、蘋果漲價被罵翻、Meta 新智能眼鏡 | M觀點M觀點GPT-5.6限量預覽擊敗Claude Mythos,蘋果MacBook Neo漲價3000元引發爭議
EP316. Claude Sonnet 5、Meta 也要賣算力、PLTR 合作 NVDA | M觀點M觀點Anthropic 推出 Claude Sonnet 5,Meta 評估出租 GPU,Palantir 與 NVIDIA 搶攻企業私有 AI