Open Models Change The Economics of AI

Y Combinator · 2026-09-04 · 57 min
https://www.youtube.com/watch?v=rY0wnfFHYbsVideo summary
Ollama CEO Jeffrey Morgan says open models could handle 80–90% of enterprise tokens as cloud usage surged 150x.
Ollama co-founder and CEO Jeffrey Morgan says enterprise AI use is shifting toward open models, driven first by coding agents and then by OpenClaw and Hermes, which brought long-running agent tasks to non-developers. Ollama Cloud token usage has risen 150x since the start of the year, while AT&T has shifted 40% of its token consumption to open models. Morgan expects open models to process 80–90% of business tokens, though frontier models will remain important for the hardest tasks. He describes Ollama as an operating-system-like layer that connects models, hardware, inference providers, and agent harnesses, and predicts a hybrid future: run simpler work locally and send demanding coding tasks to the cloud. The conversation also traces Ollama’s path from Docker and years of searching for a product-market fit to its 2023 launch, when making local models easy to run quickly attracted users and enterprise adoption.
Chapters
- 0:00Intro: Ollama serves 9 million developers as enterprise use shifts to open models
- 2:13Is It All About Cost?: AT&T moves 40% of token use to open models
- 3:31The Token Usage Explosion: Ollama Cloud token usage grows 150x
- 5:31Fine Tuning: Faster DeepSeek Releases and Open Models for Security Testing
- 8:26Launching Models at Scale: DeepSeek’s multimodal launch and Ollama’s day-zero playbook
- 11:31Ollama as the OS Layer: Connecting inference, hardware, and developer runtimes
- 13:58Hidden Layers Between Model and App: Knowledge, orchestration, and persistent memory
- 17:28Open vs. Closed: Open models may handle 80–90% of enterprise tokens
- 20:41Local vs. Cloud Models: Quinn 3.8 38B benchmarks match Opus 4.6 for coding
- 24:00NVIDIA's Open Source Play: DGX Spark runs 20–120B models with 128 GB memory
- 26:03The Return to Local: GB300 desktops could bring coding agents back from the cloud
- 27:15Getting GPUs Is Hard: B200 and B300 Access Requires Provider Partnerships
- 29:16The Flash Model Revolution: DeepSeek Flash Makes High-Volume AI Cheaper
- 31:35God Model vs. Orchestration: Kimi’s Web-Development Results Intensify Competition
- 33:37The Geopolitics Question: Model Origin, US Hosting, and Llama at a Finnish Power Plant
- 37:12Applied to YC With the Wrong Idea: From Kubernetes Desktop to Ollama in 2023
- 40:36Lost in the Wilderness for Two Years: A Team of 10+ Found Its Direction by Running Llama
- 42:15The Pivot That Changed Everything: Ollama’s First Version Shipped in Two Weeks
- 44:49100K GitHub Stars, No Revenue: Ollama Cloud Monetization Followed Enterprise Adoption
- 47:40How Do You Monetize Open Source? Coding Agents Made Open Models Ready for Harder Problems
- 49:43Why Do YC as a Second-Time Founder? Founder Peers and Lessons from Docker
- 51:49What Seeing “Good” Actually Does for You: Product-Market Fit and Two-Year Software Quality
- 53:38Old DevOps Rules That Don't Apply Anymore: Ollama's Curation of Fragmented Models
This is a Tier 1 public summary
Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.
More from this channel
Patrick Collison: Is AI Breaking the Lean Startup Playbook?Y CombinatorPatrick Collison says Stripe’s new-business starts are nearly 2x year over year, making this a better time than ever to found a company.
The Case For Data Centers In SpaceY CombinatorStarcloud CEO Philip Johnston explains how an H100 reached orbit and why 88,000 satellites could host space-based AI data centers.
Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The WorkY CombinatorWaymo co-CEO Dmitri Dolgov says the demo took 18 months, but building a safe robotaxi product took 15 years.
Garry Tan: Own Your IntelligenceY CombinatorGarry Tan says personal AGI can make one founder 400x more productive—and argues your AI skills should stay in your own repo.
Max Hodak: Average Is Not Good EnoughY CombinatorScience CEO Max Hodak says Helix and faster iteration drive startup speed, while one retinal implant patient read a 300-page novel.
Related analyses
Peter Steinberger: "Fun Is Velocity"Y CombinatorPeter Steinberger says OpenClaw hit 4.7 million weekly downloads, but 9,500 configuration options and security work nearly overwhelmed him.
Braintrust CEO on Where Engineering Actually Matters in AIa16z Deep DivesBraintrust CEO Ankur Goyal says AI engineering should focus on evals and harnesses, while SQL beat Bash in his agent benchmark.
How The Internet’s Favourite AI Employee Went RogueColdFusionOpenClaw promised a capable AI assistant but exposed private data, compromised 4,000 developer machines, and helped drive a $1 billion mortgage fraud investigation.
Pyramid of Work and The Future of Enterprise Automation | The a16z Showa16z Deep DivesHappy Robot’s agents serve nine of the top 10 U.S. freight brokers and help enterprises coordinate work across teams and countries.
EP316. Claude Sonnet 5、Meta 也要賣算力、PLTR 合作 NVDA | M觀點M觀點Claude Sonnet 5 以每百萬輸入 Token 3 美元逼近 Opus 4.8,Meta 也考慮出租閒置 GPU 算力。