In this episode of the Lightcone, Ollama Co-founder Jeffrey Morgan sits down with the hosts to talk about the massive shift toward open models in enterprise, the 150X growth in tokens since the start of the year driven by coding agents, and how Chinese models now dominate cloud token consumption.
Apply to Y Combinator: https://www.ycombinator.com/apply
Work at a startup: https://www.ycombinator.com/jobs
00:00 — Intro
01:16 — Who’s Actually Using Open Models?
02:13 — Is It All About Cost?
03:31 — The Token Usage Explosion
05:31 — Fine Tuning: Hype Cycle or Here to Stay?
07:27 — Where Open Models Beat Claude
08:26 — Launching Models at Scale
11:31 — Olama as the OS Layer
13:58 — Hidden Layers Between Model and App
17:28 — Open vs. Closed: The Steady State
20:41 — Local vs. Cloud Models
23:12 — Why Chinese Models Dominate Cloud
24:00 — NVIDIA’s Open Source Play
26:03 — The Return to Local
27:15 — Getting GPUs Is Hard
29:16 — The Flash Model Revolution
31:35 — God Model vs. Orchestration
33:37 — The Geopolitics Question
36:14 — From Docker to Ollama
37:12 — Applied to YC With the Wrong Idea
40:36 — Lost in the Wilderness for Two Years
42:15 — The Pivot That Changed Everything
44:49 — 100K GitHub Stars, No Revenue
47:40 — How Do You Monetize Open Source?
49:43 — Why Do YC as a Second-Time Founder?
51:49 — What Seeing “Good” Actually Does for You
53:38 — Old DevOps Rules That Don’t Apply Anymore
source








