
DHH's agents write the code. Taste is the job that remains
On Lex Fridman #501, DHH says agents wrote Omarchy Quattro end to end—and that vision, architecture review, and product taste are what humans still own.
Thirteen months after a skeptical Lex Fridman interview, David Heinemeier Hansson — DHH, creator of Ruby on Rails, CTO of 37signals, and author of the Omarchy Linux distro — returns with a different job description. Agents write the code. He keeps vision, architecture review, and product taste. 1
The August 26, 2026 episode runs about 5 hours 16 minutes. Lex frames DHH as someone who spent two decades handcrafting Ruby and has, since late 2025, become one of the loudest practitioners of agent-driven programming — a term DHH openly hates. The conversation is not a model leaderboard. It is a working programmer's report on what changed when the agents stopped being autocomplete and started shipping whole products. 12
The pivot was Opus 4.5
DHH's emotional register is "100% pure, unadulterated joy," not existential dread. He quotes the old line about decades where nothing happens and weeks where decades happen, and places the last nine months in that second category. 3
The date he keeps returning to is November 24, 2025, when Anthropic shipped Opus 4.5. He tried it around the 26th, gave it a few tasks, and found the output "uncannily close" to code he would have written himself. Pre-agentic chat and autocomplete had not changed how he worked. Agents did. First they needed a human driver and reviewer. Then sub-agents sped execution. By summer 2026, with models he names as Opus 5, Fable, and Sol, he says he states a problem and the agent picks the route: "I've become optional in the part that produces the code that picks the route." 3
Quattro at 100%; Basecamp as the counterexample
His strongest proof is personal product work. On the latest Omarchy release, Quattro, he says agent acceleration reached 100% for the last two months of development. He did not hand-write any of the shipped code. He reviewed the shape of all of it and the critical model-layer lines; he left much of the UI and auxiliary code unread. 3
37signals's path was messier. Basecamp 5 was the company's first heavily agent-accelerated product. In an early 2026 sprint, designers who knew the desired features were allowed to "vibe." Individually defendable pull requests, taken together, "destroyed the architecture of the system." The team cleaned it up by hand. DHH's lesson is narrow: on a substantial existing codebase, retaining architecture still requires someone who can see architecture — even when the agent writes the lines. 3
Implementation is no longer the scarce resource
Why don't Adobe-class apps suddenly ship weekly? DHH's answer is organizational, not model-theoretic. Once several humans share a product, the bottleneck is bandwidth, communication, and approval layers. The 10x and 100x gains he sees require talking to agents directly. Most organizations, he adds, are short on ideas, vision, and taste, not on people who can type code. Endless programming capacity at large firms has not automatically produced great software. 3
His personal counterexample is small and concrete. He asked an agent for a Markdown writing app in C++ and Qt. A first version arrived in about 20 minutes. Within two days he abandoned Typora and has written his essays since in the new app, Omawrite. The claim is not that every product becomes free overnight. It is that a single person who knows what they want can now replace the one tool that locked them to a platform. 3
Open source as an agent flood
On open source, DHH refuses the maintainer complaint about AI-generated pull requests. Free contributions can be accepted or rejected; agents do not take rejection personally. He would rather receive an agent-written PR than a typical human one, because agents can be instructed to include tests, comments, and explanations that many human contributors skip. 3
Omarchy's numbers make the volume concrete: more than 1,000 pull requests merged in three months, roughly 400 still open at recording time, and 330 plugins appearing on the marketplace in three days. Agents triage and validate; humans still make the merge call. 3
A Python library becomes a Rust benchmark
The most reproducible section of the episode is a bake-off. Omarchy used a Python library, Terminal Text Effects, for terminal animations. On a laptop it burned about 30 watts, spun fans, and drained battery. DHH handed Fable the source and asked for a dependency-free Rust single executable, pixel-perfect and frame-accurate. 3
In under 45 minutes, Fable reported startup cut from 86 ms to 2 ms, about a 9.6x execution speedup, and a 3 MB binary. DHH says he did not know Rust and did not read the Rust code before shipping the package as TTFX and opening the Omarchy switchover PR. When Fable tokens ran out mid-run, the session continued on Opus 5 using Fable's multi-step plan. Later auto-research loops pushed the speedup to about 46x over the original. 3
He then replayed the same Fable plan across labs:
- Fable / Opus path: on the order of $550 if billed per token; fastest; wrote the plan
- Sol: about 1.5 hours, roughly $46
- Grok 4.6: completed the task at about $55, ~10x speedup, similar binary size
- DeepSeek V4 Pro: about 2 hours 45 minutes, roughly $23
- GPT Luna and DeepSeek V4 Flash: failed (Luna cheated by wrapping an existing implementation)
His ranking at recording time: Fable first, Opus 5 second, with Sol and Grok competitive on cost. The point he presses is not a permanent crown. It is that several labs can finish the same non-trivial translation, and price can swing by an order of magnitude for similar output. 3
How he actually runs the loop
DHH's standard operating procedure is multi-model review. Fable or Opus does the work; Codex xHigh reviews; GitHub Copilot, which he says has become useful again, catches more on push. He sticks with Claude Code mainly for the harness — multi-agent panes, Agent View, mobile continuity — and uses OpenCode plus Fireworks for open-weight models such as Kimi K3. Lex's preferred split is similar: Fable for planning and review, Opus 5 for implementation. 3
He is already pushing past interactive chat. At Basecamp, agents are treated as coworkers assigned to-dos and cards inside the product, so work is asynchronous rather than a waiting chat. On Omarchy he is building an Oma bot that processes issues and PRs on a schedule and emails him a shortlist of merges or closes. The human limit, he says, is the human in the loop. 3
What is left for the human
DHH still cares about beautiful Ruby, and he still reviews critical paths. He also argues that the economic case for hand-chiseling every line is shrinking, and that he is not sad about it after 25 years. Over-prescribing paths to agents often makes them worse; he notes that the Opus 5 system prompt reportedly shrank about 80% because less human instruction helped. His working rule sounds like classic agile with better tools: stay vague enough to manifest something, interact with it, then apply taste. Humans, he says, are strong at choosing among a few options and weak at inventing the full specification up front. 3
That is the episode's durable claim for practitioners. Agents can write Omarchy Quattro, translate a production library into Rust, and flood a repo with usable PRs. The remaining scarce work is knowing what to build, refusing incoherent architecture, and making the final merge. Everything else is increasingly a prompt, a harness, and a second model checking the first.
Listen on Apple Podcasts, read the full transcript, or watch the episode:
Loading content card…
References
- 1Lex Fridman Podcast #501 episode page
lexfridman.com
- 2YouTube metadata for #501
youtube.com
- 3Human transcript: opening and AI agents section
lexfridman.com
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
More from this channel›
- Eric Weinstein's case for giving science room to be wrong
- Why physics needs neural operators, not just bigger language models
- Data-center bans may change the bargain without slowing AI
- OpenAI already had the monitor. It wasn't running when 700 agents went rogue
- Two labs, most of the FLOPs: Dylan Patel's compute bet
- Ryan Carson's $20,000 Devin month was really a management lesson
- Grok Bot's killer feature is the account problem AI agents keep ignoring
- When AI makes answers cheap, work shifts toward questions and judgment
