
Joon Sung Park's simulation thesis: behavioral models may become AI's next scaling law
Joon Sung Park argues that AI's next scaling challenge may be building behavioral simulations that preserve how real people act, including their inconsistencies, before teams commit to real-world interventions.
A large language model can describe what people say. Joon Sung Park wants a model that can test what people might do.
Park is the co-founder and CEO of Simile AI. He became widely known through the 2023 "Generative Agents" project, which placed simulated people in a small town called Smallville and let their memories and conversations produce social behavior. In his conversation with Latent Space, he treats that project as the beginning of a larger research question: can AI simulate a population well enough to test decisions before people make them in the real world? 1
The proposal matters because the internet gives AI an uneven picture of human behavior. People leave behind enormous amounts of text about what they believe, prefer, and claim to have done. That record is useful, but it is a record of speech. A policy maker, product team, or researcher often needs a different answer: what will people actually do when prices change, a message is rewritten, or a new service becomes available?
From language models to behavioral models
Park's argument starts with a distinction between representing a person and predicting a response. A language model can imitate a plausible answer from a profile. A useful simulator has to preserve the particular person's habits, memories, social position, mistakes, and contradictions, then produce behavior under a changed situation. The difference is practical: the first system generates a response, while the second lets a researcher compare interventions.
Simile's proposed data mix reflects that demand. Park describes combining long-form interviews, observational and transaction data, and randomized controlled trials. Interviews provide a person's stated beliefs and history. Observational data supplies evidence about behavior in context. Experiments connect an intervention to an observed change. Each source answers a different question, and the simulation needs all three to keep a model from confusing what people say with what they do. 2
The research target is therefore closer to a behavioral model than to a more articulate chatbot. The model should let a team ask, for example, how different groups might react to a public message or a product change, then identify which assumptions deserve a real-world test. The simulation narrows the search space; it does not remove the need to measure what happens outside the simulation.
Why realistic people must be irrational
The hardest requirement is also the easiest to lose during optimization. Park argues that a simulation becomes less useful when it turns people into generic rational actors. Real people misremember, follow social cues, use shortcuts, react emotionally, and make decisions that conflict with their stated interests. Those errors are part of the behavior a simulator must reproduce.
That requirement changes how a team should judge the model. A system that gives a cleaner or more logically consistent answer may look better in a conventional evaluation. A system that reproduces a person's inconsistency may be more useful for testing a real intervention. The question is not whether the agent behaves like an ideal decision-maker. The question is whether the agent preserves the patterns that would affect the decision being studied.
Park reports that Simile's digital twins of 1,000 representative Americans reproduced behavior and attitudes at about 85% of the consistency with which people reproduced their own responses. The comparison is meaningful because the target is human response consistency rather than an abstract language benchmark. The figure is an early result from the episode's account of the project, and its value depends on how well the tests generalize across questions, populations, and settings. 2
Simulation is for choosing interventions
Park connects this work to agent-based modeling and to Thomas Schelling's studies of how simple individual decisions can produce large social patterns. The new ingredient is the attempt to give each simulated person a richer memory and a data-grounded behavioral profile. That combination could make it cheaper to explore many possible interventions before committing money, time, or political capital to one of them.
The important output is a set of conditional answers. If a message changes, which groups respond differently? If a policy creates a new cost, where does resistance appear? Which result is robust across plausible assumptions, and which result depends on one uncertain detail? Those questions help a team decide what to test next.
Park also frames simulation as a way to shape the future rather than forecast a single inevitable outcome. A forecast asks which result is most likely. A simulation can compare several choices and show how each choice changes the path. That distinction gives the technology a constructive use, while keeping the decision with the people who set the intervention and accept its consequences.
The scaling problem is validation, not only compute
Simile's long-term ambition reaches population scale, potentially modeling all 8 billion people. That goal would demand enormous compute, but compute is only one constraint. A population simulator would need repeated validation against real behavior, careful sampling, and a way to detect when its assumptions fail for a group or situation. The more consequential the decision, the less acceptable it becomes to treat a plausible synthetic response as evidence by itself.
Park compares simulation with painting: both are selective representations of reality. The comparison works only when the selection preserves what the viewer needs for the task. A simulation designed to test a public-health message may need accurate social influence and risk perception. A simulation designed to forecast retail demand may need different details. No single notion of realism can settle both cases.
That is the useful limit of the episode's scaling-law idea. More data and more capable models may improve simulated populations, but the central test remains task-specific: does the simulator preserve the behaviors that determine the decision? Simile's approach points toward an AI research agenda built around experiments on people, rather than only larger models trained on records of what people have already said.
Loading content card…
The full conversation is available through the Latent Space episode page.
References
- 1Latent Space episode page
latent.space
- 2Simulation: the new Scaling Law - Joon Sung Park, Simile AI
podcasts.apple.com
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
