Why The Harness Matters More Than The Model | YC Paper Club

Y Combinator · 2026-09-07 · 60 min
https://www.youtube.com/watch?v=n9xKblqyQ28Video summary
YC Harness Night shows Prime Agent reaching 95% on ARC-AGI, OpenJarvis cutting costs 800x, and QM scaling agents across YC.
At YC Harness Night, speakers argued that the harness—not merely model weights—determines what AI agents can accomplish. The same Claude Opus model reportedly rose from 30% to 95% on ARC-AGI with a stronger harness, while Prime Agent demonstrated persistent sub-agents, programmable memory, context management, and long-horizon work, including 633 agents producing 23 million tokens over a seven-day factorial run. OpenJarvis presented a privacy-focused local AI stack that uses cloud models to optimize on-device configurations, claiming up to 800x lower inference cost. YC’s QM showed how one open-source harness gives every employee personalized Slack and web assistants with sandboxes, scheduled jobs, internal data access, and shared resources—while highlighting unresolved challenges around persistence, permissions, social context, and agents giving up too early.
Chapters
- 0:00Why harnesses matter: ARC-AGI rises from 30% to 95%
- 4:27Building an auto-researcher by accident: eight H100 nodes producing papers
- 7:13A five-minute history of harnesses: from GPT-2 loops to recursive RLMs
- 13:56Self-improving harnesses: DSPy prompts to Darwin Machines editing code
- 17:22Tonight's speakers: Seth Karten, OpenJarvis, and QM
- 18:35Seth Karten: Prime Agent as a self-improving RLM harness
- 21:44Context as an L1, L2, L3 cache: compaction, RAM, and refinement
- 24:51From Turing machine to von Neumann computer: expressive harnesses and persistent sub-agents
- 28:24Messaging between agents: coordination for long-horizon performance
- 30:04ARC-AGI results: Prime Agent reaches 95.5% with Opus
- 33:09Emulator Bench and GPU kernels: out-of-loop experiments scale to 633 agents
- 39:21The five primitives of a personal AI stack: OpenJarvis runs locally
- 42:47Letting cloud models optimize your local stack: OpenJarvis beats out-of-box deployments
- 43:53800x cheaper than the cloud: local inference benefits from cloud optimization
- 45:58Josh France and Regan Bell: QM, YC's agent harness for work
- 47:29A history of YC's internal agents: from the January 2025 general agent to coding bots
- 49:24OpenClaw and a fleet of 50 agents: YC scales personal assistants
- 51:04Pulling the brain out of the sandbox: QM centralizes context in Postgres
- 54:43Letting the agent choose its own sandbox and model: QM keeps the harness thin
- 57:16The grind tool: budgets on goals prevent agents from quitting early
- 58:50Agents don't understand social context: permission systems limit shared knowledge
This is a Tier 1 public summary
Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.
More from this channel
Jeff Dean: The 1% Rule for Building in AIY CombinatorJeff Dean links Google Search, TPUs, Gemini agents, and founder strategy to 1,000x energy gaps and 0%-success niches.
Patrick Collison: Is AI Breaking the Lean Startup Playbook?Y CombinatorPatrick Collison says Stripe took two years to launch, while Stripe data shows AI-era startups nearly doubled year over year.
Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The WorkY CombinatorWaymo’s demo took 18 months, its product 15 years, now reaching 17× human safety across 220 million autonomous miles.
Garry Tan: Own Your IntelligenceY CombinatorGarry Tan argues personal AGI, G Brain, and Markdown skills let one founder achieve 400x leverage while owning their intelligence.
Max Hodak: Average Is Not Good EnoughY CombinatorMax Hodak explains how Science’s Helix infrastructure, hiring system, and retinal implant work turn speed into startup advantage.