The eval is the new PRD: Anthropic's product lesson from Dianne Penn

The eval is the new PRD: Anthropic's product lesson from Dianne Penn

Dianne Penn's account of Anthropic shows why strong AI products come from turning user failures into evals, pairing frontier models with usable workflows, and keeping human judgment in the loop.

The product advantage is the loop, not the prompt

Anthropic's most useful product lesson is easy to miss in a conversation about faster models: the advantage comes from turning vague user frustration into a measurable test, then feeding that test back into research. Dianne Penn, the company's Head of Product for AI Research and Labs, describes a product organization built around that loop. She joined Anthropic in 2023 as its first technical product manager, when the product group had five engineers and one engineer handled the entire API business. 1
Her shorthand is blunt: "evals are the new PRDs." An evaluation, in this context, is a repeatable test for whether a model handles a real user need. The shift matters because a conventional product document can describe an intended experience, while an eval can tell a research team exactly where the current model fails and whether a new version improved it. 2

A small team learned by shipping odd things

Penn's story of Anthropic's early years is less polished than the usual account of a frontier lab. The team was still working out what its technology was for. It had Claude.ai, tool use, and a large amount of uncertainty about how to translate research into user and social value. One early experiment helped the group find a product identity: Golden Gate Claude.
After Anthropic published interpretability research showing that model features could be associated with themes, a cross-functional group dialed up a Golden Gate Bridge feature and made Claude obsess over the bridge. Ask for a spaghetti recipe and the response might return to the bridge's orange color. Engineering, product, design, and research built the experience in about 24 hours. Penn estimates that roughly 2,000 people used it, but the audience size was not the point. The experiment showed that Anthropic could turn a research result into a distinctive public experience at startup speed. 2
That is a useful distinction for AI product teams. A research demo is not automatically a product, but a product team can use a small, strange release to learn what kind of company it knows how to be. The team did not need a complete roadmap before it could test that proposition.

Coding became a product thesis

The more consequential shift came from watching how users worked. When Penn joined in 2023, she says people did not naturally associate Anthropic or Claude with coding. Models were used for autocomplete and many other tasks, but users were beginning to write long-form code with them. Anthropic treated that behavior as a training and product opportunity, improving the coding performance of Opus 3. Penn describes the change as relatively small from a training perspective, but large enough to give early developers a reason to choose Claude. 2
The follow-on lesson is stronger than the coding example itself. Penn argues that a frontier model needs a frontier product experience. Claude Code gave model capability a place where users could feel it, while the stronger model accelerated adoption of the product. A model release and a user workflow therefore formed a feedback loop: capability created a new behavior, the workflow exposed more failures, and those failures gave the lab better targets for the next iteration.

Evals shorten the distance from pain to action

Penn's method starts with a complaint, but it does not stop at the complaint. "Claude hallucinated" is too vague for a researcher to fix. The product team has to inspect the trajectory and determine whether the model used the wrong tool, searched the wrong document, selected the wrong facts, synthesized them badly, or behaved with unjustified confidence. Each diagnosis points to a different intervention and a different test.
She gives an early Claude example involving structured output. Users said the model was bad at following instructions. After examining the underlying cases, the team found that about 80% of the complaint referred to Claude failing to produce the required JSON. The team collected 30 to 40 examples, turned them into an eval set, and ran that set against later versions. Penn says the failure rate eventually fell to the point where the team saw roughly 99.9% success. 2
This is product management as test design. The work is not merely to rank feature requests. It is to make a user need concrete enough that a model team can act on it and a product team can tell whether the fix holds. The eval also becomes a shared language across research, engineering, safety, and product.
That does not make the traditional PRD obsolete. Penn still uses PRDs to align a large group around a model release and to explore ambiguous, early-stage work such as computer use. The difference is where the artifact is most useful. An eval is a compact instrument for a known failure; a PRD remains a coordination and exploration document when the problem itself is still being defined. 2

The jagged frontier makes measurement mandatory

The conversation also explains why this loop cannot be replaced by a one-time model score. Penn distinguishes between smooth improvement in training loss and discontinuous jumps in individual capabilities. A model may move from unreliable arithmetic to reliable arithmetic, or suddenly become able to perform a class of tasks that the product team was not expecting. The model's general trend can look predictable while its useful capabilities arrive unevenly.
That unevenness creates both product opportunity and safety risk. Teams have to discover what a model can do, identify what users will actually value, and test the new behavior before releasing it into workflows with real permissions. In Penn's terms, current models still have product and user overhang: capabilities exist that have not yet been turned into good experiences. The work is exploratory, but it still needs instrumentation. 2

Token maxing is really learning maxing

Lenny frames "token maxing" through a claim from YC president Garry Tan: someone willing to spend $100,000 a year on tokens today is living like a person in 2028. Penn accepts the underlying intuition but changes the unit of value. Token spend is input; experimentation is output.
That framing explains why Anthropic's most creative internal users spend time with every new research model. They are not simply trying to automate an existing task. They are using the model to discover what tasks, interfaces, and products should exist. Penn describes early company-wide Slack experimentation in which one person would try a use case, others would adapt it, and a promising workflow could emerge after roughly ten requests. The social part matters: experimentation is a team activity, not a private contest to write the cleverest prompt. 2
For teams outside a lab, the practical version is modest. Pick one or two important workflows, use the newest capable model deeply enough to expose its limits, save representative failures, and turn the recurring failures into tests. Broad, shallow exposure creates anecdotes. Concentrated use creates product knowledge.

Human judgment moves up the stack

Penn's argument ends where the usual automation story begins. As building becomes easier, the scarce work is deciding what deserves to be built, whether it solves the user's actual problem, and whether the result is good enough to trust. That is why she insists that managers and senior product people stay hands-on. She still owns one or two workstreams during model development so she can maintain a working sense of how the models behave and where their limits are. 2
She makes the same point about Claude's personality. A useful thinking partner cannot agree with every idea. It needs to push back when a plan is weak, and the user needs enough judgment to decide whether that disagreement is right. The product is better when the model adds friction in the right place rather than maximizing compliance.
The result is a less glamorous but more durable picture of AI product work. Progress does not come from a perfect prompt or a single benchmark. It comes from close contact with users, careful failure analysis, repeatable evals, and people willing to revise their assumptions as capabilities jump. The model may write the code. Someone still has to decide what the code should prove.
Loading content card…
Read the full episode on Lenny's Newsletter or listen through Apple Podcasts.

Related content

  • Sign in to comment.
More from this channel