Share:

Video summary

Inferact’s vLLM founders aim to build an open inference layer for models running across 400,000–500,000 GPUs.

Inferact co-founders Simon Mo and Wau Kuan traced vLLM to a 2022 UC Berkeley effort to speed up Meta’s 175-billion-parameter OPT model. Its distinctive challenge was scheduling constantly changing language-model requests and managing KV-cache memory, leading to the PagedAttention work. The open-source project has grown to more than 2,000 contributors, with 50 or more regular full-time contributors and a CI bill exceeding $100,000 monthly. vLLM now runs on an estimated 400,000–500,000 GPUs around the clock, across varied chips and models. The founders say inference is getting harder as models scale beyond a trillion parameters, architectures and hardware diversify, and agents add long-running tool interactions that complicate cache management. They founded Inferact to support vLLM and build a universal inference runtime for new models, hardware, and applications; open source, they say, is the company’s top priority.

Chapters

  1. 0:00Introduction: Inferact’s Universal Inference Layer and vLLM’s OPT Origins
  2. 6:01Introduction: Dynamic LLM Scheduling and the First vLLM Meetup
  3. 11:41Community and Collaboration in vLLM: 2,000+ Contributors and $100K-Plus Monthly CI
  4. 19:19Understanding Inference Engines: Components and the Multi-Trillion-Parameter Trend
  5. 24:27Cluster Scale and GPU Deployment: Up to 500,000 GPUs and Agent-Driven Cache Challenges
  6. 31:19Belief in Open Source AI: Amazon Rufus and Character.AI’s vLLM Deployments
  7. 35:45Founding of Inferact: vLLM’s Universal Inference Layer and Yang as Co-founder
  8. 40:00Future of Inference at Scale: GB200/GB300 NVL72 and a Universal Runtime

This is a Tier 1 public summary

Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.

More from this channel

Related analyses