Share:

Video summary

Ollama CEO Jeffrey Morgan says open models could handle 80–90% of enterprise tokens as cloud usage surged 150x.

Ollama co-founder and CEO Jeffrey Morgan says enterprise AI use is shifting toward open models, driven first by coding agents and then by OpenClaw and Hermes, which brought long-running agent tasks to non-developers. Ollama Cloud token usage has risen 150x since the start of the year, while AT&T has shifted 40% of its token consumption to open models. Morgan expects open models to process 80–90% of business tokens, though frontier models will remain important for the hardest tasks. He describes Ollama as an operating-system-like layer that connects models, hardware, inference providers, and agent harnesses, and predicts a hybrid future: run simpler work locally and send demanding coding tasks to the cloud. The conversation also traces Ollama’s path from Docker and years of searching for a product-market fit to its 2023 launch, when making local models easy to run quickly attracted users and enterprise adoption.

Chapters

  1. 0:00Intro: Ollama serves 9 million developers as enterprise use shifts to open models
  2. 2:13Is It All About Cost?: AT&T moves 40% of token use to open models
  3. 3:31The Token Usage Explosion: Ollama Cloud token usage grows 150x
  4. 5:31Fine Tuning: Faster DeepSeek Releases and Open Models for Security Testing
  5. 8:26Launching Models at Scale: DeepSeek’s multimodal launch and Ollama’s day-zero playbook
  6. 11:31Ollama as the OS Layer: Connecting inference, hardware, and developer runtimes
  7. 13:58Hidden Layers Between Model and App: Knowledge, orchestration, and persistent memory
  8. 17:28Open vs. Closed: Open models may handle 80–90% of enterprise tokens
  9. 20:41Local vs. Cloud Models: Quinn 3.8 38B benchmarks match Opus 4.6 for coding
  10. 24:00NVIDIA's Open Source Play: DGX Spark runs 20–120B models with 128 GB memory
  11. 26:03The Return to Local: GB300 desktops could bring coding agents back from the cloud
  12. 27:15Getting GPUs Is Hard: B200 and B300 Access Requires Provider Partnerships
  13. 29:16The Flash Model Revolution: DeepSeek Flash Makes High-Volume AI Cheaper
  14. 31:35God Model vs. Orchestration: Kimi’s Web-Development Results Intensify Competition
  15. 33:37The Geopolitics Question: Model Origin, US Hosting, and Llama at a Finnish Power Plant
  16. 37:12Applied to YC With the Wrong Idea: From Kubernetes Desktop to Ollama in 2023
  17. 40:36Lost in the Wilderness for Two Years: A Team of 10+ Found Its Direction by Running Llama
  18. 42:15The Pivot That Changed Everything: Ollama’s First Version Shipped in Two Weeks
  19. 44:49100K GitHub Stars, No Revenue: Ollama Cloud Monetization Followed Enterprise Adoption
  20. 47:40How Do You Monetize Open Source? Coding Agents Made Open Models Ready for Harder Problems
  21. 49:43Why Do YC as a Second-Time Founder? Founder Peers and Lessons from Docker
  22. 51:49What Seeing “Good” Actually Does for You: Product-Market Fit and Two-Year Software Quality
  23. 53:38Old DevOps Rules That Don't Apply Anymore: Ollama's Curation of Fragmented Models

This is a Tier 1 public summary

Whether the chapter key points, section summaries and mind map are public is up to the person who shared it. Want the full analysis?Submit one yourself.

More from this channel

Related analyses