
Ryan Carson's $20,000 Devin month was really a management lesson
Ryan Carson's Devin workflow shows how founders can manage agent throughput while keeping customer insight and product judgment human-owned.
Ryan Carson's conversation with Claire Vo begins with an eye-catching number: he once spent $20,000 on Devin in one month. The useful lesson is the operating system around that spend. Carson is no longer treating a coding agent as a faster pair programmer. He is treating a group of agents as a workforce that needs priorities, review loops, and a human owner for product judgment. 1
Carson is a five-time founder and the solo founder of Untangle, a B2B software product for family-law firms. He joined Claire Vo on How I AI to compare how two founders now use cloud agents, local agents, and repeatable playbooks. 1
The job is managing the agent queue
Carson's work moved from local development into the cloud after he found Devin mature enough for production work. He says he was spending about $5,000 on Devin before reaching $20,000 in one month. He runs roughly 10 to 15 concurrent threads and organizes them into folders: bugs, P0, P1, and P2. The folders answer a management question before an agent question: which work must move the business forward today? 2
A handwritten weekly priority list sits beside the screens. The paper is a constraint on the founder, not a data source for the agents. It keeps the human's attention on a few business priorities while the agents produce a stream of pull requests, questions, and updates. Carson's method resembles ordinary delegation: define the work, rank it, and let the delegate operate without requiring constant instructions.
The arrangement changes what skill matters. Carson and Claire describe agent management as a new form of management work: setting goals, deciding what to delegate, checking results, and choosing where to intervene. The analogy to a larger organization is direct. A founder cannot track every action by dozens of workers, so the founder needs structure that makes attention selective. 2
Watchdog turns background work into a business review
Carson's Watchdog workflow shows why the agent's value extends beyond writing code. Untangle serves multiple family-law firms, and Carson needs to know what is happening across their cases and accounts. Watchdog checks recent activity, Sentry errors, and other account signals. It then surfaces the three most important problems and checks whether an open pull request already fixed each one. 2
The important design choice is the filter. Carson does not ask for a pile of raw events. He asks for a ranked view of what needs attention, followed by the status of the corresponding fix. The workflow turns a recurring anxiety—"what is happening in the business right now?"—into a repeatable review.
The same pattern applies to investor updates, customer triage, and operations. Carson built an investor-update skill, and he uses Devin for quoting, custom deal-desk work, and other tasks that require knowledge of the codebase plus the ability to create a small tool. A coding agent becomes more useful when the owner asks what the agent can do for the whole business, rather than limiting the agent to files in a repository. 1
More code still cannot choose the product
The conversation repeatedly returns to a boundary that agent enthusiasm can blur. Faster output does not create a customer or decide which feature deserves a place in the product.
Carson found product-market fit for Untangle through a change in customer and setting. He first aimed the product at consumers going through divorce. Those customers wanted lawyers and other people involved in the process, so Carson changed direction. A conversation with family-law attorney Renee Bauer revealed a different problem: firms faced a paralegal shortage and a painful discovery workflow. Bauer wanted to pay for a tool that addressed that problem. 2
The episode treats that conversation as part of the product process, not as a task an agent can cheaply simulate. Agents can prototype the workflow after a customer explains the problem. The customer still supplies the evidence that the problem matters and the judgment that a proposed solution is worth buying.
Claire makes the same point from the product side. She runs fewer agents overnight when the work involves deciding what ideas deserve development. Both founders would rather constrain output than create a larger queue of features that nobody needs. Their rule is easy to apply: automate the repeatable work, then spend human time on customers, priorities, and the decision to ship.
Match the agent to the environment
Carson's setup is not a simple replacement of local tools with cloud tools. He uses cloud agents for bugs, asynchronous work, operations, and tasks that can proceed while he is away from the computer. Claire uses Codex for low-latency, browser-connected feature work that needs close product supervision. She also uses it for verification: a browser agent can follow user stories in a preview environment and report which flows pass. 2
Both founders add review layers around the output. Claire's Merge Mommy scores pull requests for risk and automatically approves low-risk changes while sending medium- and high-risk changes to a human reviewer. Carson's Land PR playbook runs a fresh review, allows two correction loops, and produces a video walkthrough before he approves a merge. 2
Those examples make the control problem visible. An agent can generate a change quickly. A production workflow still needs a way to inspect the change, test the result, and route risk to a person who can own the consequence.
The organization changes around the agent
The hosts disagree about where agent work should live. Carson prefers agent threads because they keep context and work together. Claire likes public Slack threads because they put the agent's work where a team can see it. Their shared concern is private, contextless chat: a conversation can consume time without producing a task, decision, or artifact. 2
Carson's hiring process follows the same logic. He asks candidates to record themselves building a real feature with an agent. The recording shows how a candidate frames a problem, gives instructions, checks output, and keeps the work moving. A later stage gives the candidate access to a real system and asks for another recorded task before a short human conversation. 2
The episode's operating rule is compact. Classify work before assigning it. Give background agents repeatable playbooks. Keep a visible priority list. Add review and verification to the path into production. Use customer conversations to decide what deserves to exist. The founder's leverage comes from managing those boundaries, not from maximizing the number of generated pull requests.
Listen to the full episode on Apple Podcasts or watch it here:
Loading content card…
References
- 1I spent $20,000 on Devin in a month — Lenny's Newsletter
lennysnewsletter.com
- 2
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
More from this channel›
- Data-center bans may change the bargain without slowing AI
- OpenAI already had the monitor. It wasn't running when 700 agents went rogue
- DHH's agents write the code. Taste is the job that remains
- Two labs, most of the FLOPs: Dylan Patel's compute bet
- Grok Bot's killer feature is the account problem AI agents keep ignoring
- When AI makes answers cheap, work shifts toward questions and judgment
- Dario, data centers, open models: All-In's argument over who should control AI
- Why AI data centers became a bipartisan local revolt
