
Typed decisions, independent verifiers, and five agent-builder signals from September 19-26
Five in-window builds show agent builders writing limits into their systems and shipping the check that enforces them: a security-audit skill whose verifier never found the bug, a model that returns typed decisions, a key service that keeps keys out of the agent's context, a runtime that publishes what it lacks, and speech built from a staged conversation.
The week in one line
Between September 19 and September 26, five agent-building signals landed, and each one writes down what may cross the boundary between a model and the rest of the system. The ones worth copying also ship the thing that checks the rule.
Cloudflare's security-audit skill sat near the top of GitHub's weekly Trending board, and builders who pointed it at their own code kept reporting the same useful half: a second wave of agents whose only job is to disprove the first wave's findings 1. TypeSafe AI's Jev returns typed numbers where a chat model returns text, and hosting it pushed Simon Willison to add a capability flag to his
llm library for models that accept no conversation history 2. That library also gained a small key service, so an API key can be installed on a remote agent machine and stay out of the agent's context 3. Cloudflare's Python Workers reached general availability with a published list of the Python features its WebAssembly runtime lacks 4. Google released speech models whose unit of work is a staged conversation between named voices 5.Five signals follow, each with the artifact to open and one test to run.
The comparison map
| Signal | Builder or project | In-window date | Inspectable artifact | First adoption test |
|---|---|---|---|---|
| The verifier never found the bug | Cloudflare, security-audit-skill | Sep 26, weekly Trending read | findings.json checked against report-schema.json, plus coverage-ledger.json 1 | Read the rejected verdicts before the confirmed ones. |
| A model whose output type is fixed before the call | TypeSafe AI, Jev | Sep 21-22 | A query of typed questions returning probabilities and confidences 2 | Rescore 100 cheap retrieval hits and compare the new order. |
| A key the agent never reads | Simon Willison, llm-keys-ui 0.1 | Sep 20 | A local web form, then llm keys get inside a shell command 3 | Route the next key on a remote box through the form. |
| A runtime that publishes what it lacks | Cloudflare, Python Workers | Sep 21 | The Workers page for the Python standard library 4 | Read that page before porting a worker that uses threads. |
| Speech whose unit is a conversation | Google, Gemini 3.8 Flash TTS | Sep 23 | A script with named speakers, per-line delivery styles, and 2,000+ voices 5 | Direct one two-speaker exchange line by line. |
Five signals to inspect
1. The verifier never found the bug
Cloudflare published
security-audit-skill, a coding-agent skill that turns a coding agent into a security auditor. The skill runs six phases: reconnaissance, coverage-led hunting, candidate validation, structured output, independent record verification, and target-neutral reporting 1. The repository carries the MIT license, holds 14 commits, and its newest commit dates from September 14, before this window 1.During the window, other people began pointing the skill at their own code. It reached second place on GitHub's weekly Trending board with about 9,500 stars gained over the week, read on September 26 6. On September 25 the team behind the Kaspa wallet Kaspire published the result of a four-hour audit 7. The audit returned mostly minor issues and one real finding: the site's post-build gate treated a payload's SHA-256 as authority for a staged Android update APK, skipping verification of the envelope's RSA signature 7.
Loading content card…
Builders keep naming the second wave as the part that pays. Every unique candidate goes to a fresh verifier whose job is to disprove it, and fresh agents re-verify the final source claims, so the agent that checks a finding never found it 1. Verdicts land in three buckets:
confirmed carries a complete source trace and a bounded observed result, needs_validation carries one exact unresolved fact while the skill withholds a severity, and rejected records a disproof 1.Two zero-dependency scripts police the contract.
validate-findings.cjs checks findings.json against report-schema.json in phase four and again after every phase-five replacement, and validate-coverage-ledger.cjs checks coverage-ledger.json once the ledger exists and after each later update 1. Repeat runs compose: the skill reads prior ledgers and findings to target gaps and revalidate changed source 1.Cloudflare's own test numbers set the expectation. In its test runs, a single pass found roughly half of the vulnerabilities that repeated passes found in total, and the findings skewed toward the simpler cases 8. That post is dated June 18 and describes the fleet-wide harness this skill seeded, so the coverage figure is Cloudflare's observation from its own tests.
One prerequisite decides whether the workflow can do anything at all. Running target-controlled builds, tests, browsers, or fuzzers requires an OS-enforced sandbox that disables external networking, uses a sanitized allowlisted environment, enforces resource limits, and permits writes only to assigned scratch paths. With those controls absent, the skill keeps a lead as
needs_validation and stops there 1.Try it. Install the skill, point it at the smallest repository you own, and read the
rejected records first. Then count the needs_validation leads that arrived with an exact unresolved fact attached, because those are the ones a second run can close.2. A model whose output type is fixed before the call
TypeSafe AI announced Jev on September 15, before this window, as the first of what it calls System One models: text in, floating-point numbers out 9. The window holds what builders did with it. Simon Willison published notes on September 21 2, and Latent Space ran an episode with TypeSafe's founder the same day 10.
A call composes a
state object, which can be a string, a list of strings, or a set of name-value pairs, and sends it with one or more questions. Yes/no questions return a confidence between 0 and 1, choice questions return a probability distribution across the options you supplied, and score questions return a float along a described range. Questions are evaluated in parallel, so a batch takes about as long as a single question 2.TypeSafe's own numbers are the first thing to test. The company lists $0.042 per million input tokens with output billed at nothing, and end-to-end latency of 70 to 500 milliseconds against 3 to 329 seconds for frontier chat models on the same class of task 9. TypeSafe also says schema matching is guaranteed, which is where its claim of no hallucination comes from 9.

Hosting the model needed a new contract in Simon Willison's
llm library. Release 0.36, published September 22, lets a model plugin declare supports_conversation = False; the library then raises llm.ConversationNotSupported when such a model receives assistant or tool history, and llm chat rejects the model before a session starts 11. The first plugin to use the flag is llm-typesafe, the plugin that wraps Jev 12. The same release wraps reasoning traces in tags inside the Markdown output of llm logs 11.The design leaves a gap a builder has to plan around. A confidence score arrives as a bare number, so a spam verdict or a ranking comes with fewer surfaces to inspect than a chat model's written justification. Simon ran one experiment that scored every city in the San Francisco Bay Area on a "Good city?" question and got Cupertino on top and East Palo Alto at the bottom 2. TypeSafe's own jaggedness guidance reports the model is weak on numbers, dates, and adversarial content 2. Bias in a single float has to be found by experiment, and the price makes hundreds of probes affordable.
Try it. Take the classification you trust least in your own pipeline, send 100 candidates to a decision model as a score question, and compare its order against the one you ship today. Then ask what would have to be true for you to act on a bare number.
3. A key the agent never reads
Simon Willison published
llm-keys-ui 0.1 on September 20 for one narrow problem 3. He had started using Codex Remote to run coding agents on several machines while controlling them from his phone, and he did not want to paste API keys into an agent session 3.The plugin serves a small web interface on the machine. Running
uvx --with llm-keys-ui llm keys-ui --all makes the agent report the server's URLs, including local network and Tailscale addresses, and a form at one of those addresses saves new keys 3. The interface never displays an existing key value, and later work reaches the key through a command such as llm keys get anthropic inside a shell command 3. The release is tagged 0.1 on GitHub 13.The advertised addresses are the part to weigh. Because the server lists the machine's local network and Tailscale addresses, anything on those networks can reach the form, so the reachability that makes remote setup convenient is the same reachability worth restricting. Codex Remote itself is documented as a way to run coding agents on a machine while driving them from the ChatGPT app 14.
Try it. Next time you configure a key on a machine that an agent controls, run the key service there and hand the key over through the form rather than pasting it into the conversation. Firewall the port to the networks you actually trust.
4. A runtime that publishes what it lacks
Cloudflare's Python Workers reached general availability on September 21, after two years in preview 4. The runtime compiles Python to WebAssembly through Pyodide and runs it inside
workerd, Cloudflare's V8-based runtime 4.Three changes make the runtime usable for ordinary Python work. Bindings now convert types inside the runtime, so
self.env.QUEUE.send({"key": "value"}) works, replacing the to_js glue that used to sit at the RPC boundary 4. A workers.asgi connector runs FastAPI applications and workers.wsgi runs Django or Flask, with the Workers platform taking the place of a web server 4. Hyperdrive reaches Postgres and MySQL because Cloudflare implemented socket system calls on top of the Workers connect API, where a WebAssembly sandbox normally stubs the POSIX networking calls out 4.The published limits are the part to read before porting anything. The Workers page for the Python standard library states that both
multiprocessing and threading are non-functional in the WebAssembly virtual machine 15. Simon Willison flagged the same two features while covering the release 16.Simon also described the local development story. Cloudflare's
pywrangler tool runs a full local simulation of the platform, which for him meant executing Python through Pyodide in WebAssembly inside V8 inside a 123 MB workerd binary that landed under node_modules 16. Cloudflare credits the release to Gyeongjae Choi, Dominik Picheta, and Hood Chatham, two of whom are Pyodide core maintainers 16.For an agent builder, this is somewhere agent-written Python can run, with a written list of what will fail there. Check that list against the libraries you intend to let an agent use.
Try it. Pick the two standard-library features your code depends on most, and read the standard library page for Python Workers before you let an agent generate a worker. Port the smallest one first and see which import breaks.
5. Speech whose unit is a conversation
Google released two speech models on September 23. Gemini 3.8 Flash TTS handles voice design and character work, and Gemini 3.8 Flash-Lite TTS targets high-volume dubbing and voice agents 5. The voice library holds more than 2,000 production voices across more than 100 languages, and a voice can be replicated from a 30-second sample with consent verification, SynthID watermarking, and C2PA credentials attached to the result 5.
The interface change matters more than the voice count. Both models accept a script in which speakers carry names and each line carries its own direction, and the API supports native two-speaker scene staging with conversational turn-taking, plus non-verbal cues and active-listening interjections for timing 5. Google's post reports first and second place on Hume AI's Overall Quality Index and first on Hume AI's Voice Design Benchmark at 71.4, alongside top positions in blind preference tests on Voice Arena; those ratings belong to the benchmark owners, and the post is where they are collected 5.

speech_metadata annotations on each dialogue line. 17Simon built that playground the same day, using the API's open CORS policy to keep it a bring-your-own-key page that talks to Google directly 17. His measurements are one person's, on one day: about 20 seconds to generate 1 minute 18 seconds of audio with Gemini 3.8 Flash TTS, for 2.74 cents, and the page that loads the voice catalogue reported 2,089 voices 17.
The interesting part for an agent builder is where the prompting sits. A voice agent that once assembled a string per turn can now declare a cast before the conversation starts, then hand each line a speaker and a delivery style. That is the same shape as a typed tool call: the model gets fewer choices about who is speaking and how.
Try it. Write a two-speaker exchange of six lines, give each speaker a voice and each line a delivery style, and generate it in the playground. Then check whether your own voice pipeline can carry a cast through to the request, or whether it flattens everything to one voice per call.
What to try this week
- Audit one repository and read the negative results. Point the security-audit skill at the smallest repository you own and start with the
rejectedverdicts. - Give one classification job to a decision model. Send 100 candidates as a score question and compare the order it returns with the one you ship today.
- Move one API key out of the agent's context. Run the key service on a remote agent machine and hand the key over through the local form.
- Read a runtime's standard-library page before generating code for it. Check the two features you rely on most against what the runtime documents as missing.
- Direct a cast, then the lines. Give one audio request a cast, a voice per speaker, and a delivery style per line.
The pattern
Five artifacts, and a limit written into each one. The question that separates them is what enforces that limit.
Three are enforced by machinery that refuses to carry on. A
findings.json file that fails report-schema.json stops the phase 1. A model plugin marked supports_conversation = False makes the llm library raise before it sends history to a model that takes none 11. An API key handed over through the local form stays on the machine and out of the conversation 3.Two rest on the reader. The Python Workers limits live on a documentation page, and code that reaches for a thread will fail when it runs 15. A replicated voice carries consent verification, a SynthID watermark, and C2PA credentials, and the voice catalogue describes what can be generated, so any particular voice is your own sample to test 5.
The decision model inverts the problem: its confidence score is calibrated by design and arrives as a bare number, so the only check you get is an experiment against your own task 2.
The security-audit skill puts the check somewhere else entirely: a second agent whose only instruction is to disprove the first. That design costs a sandbox before it will execute anything, and Cloudflare's own test observation says one pass finds about half the bugs that repeated passes find 1.
The rule to carry out of this week: when a tool states a limit, find the thing that enforces it. A schema, a raised exception, and a key held outside the transcript are checks you can build on. A confidence score, a catalogue count, and a documentation page are claims you inherit, and each one becomes visible the first time something runs.
This issue covers September 19 through September 26; the next one arrives next Saturday.
References
- 1cloudflare/security-audit-skill
github.com
- 2Jev introduces a new shape of LLM
simonwillison.net
- 3llm-keys-ui 0.1
simonwillison.net
- 4Python Workers are now generally available
blog.cloudflare.com
- 5Gemini 3.8 text-to-speech says hello
blog.google
- 6GitHub Trending, weekly
github.com
- 7
- 8Build your own vulnerability harness
blog.cloudflare.com
- 9Introducing System One Models & Jev
typesafe.ai
- 10
- 11llm 0.36
github.com
- 12simonw/llm-typesafe
github.com
- 13llm-keys-ui 0.1 release
github.com
- 14Codex Remote
learn.chatgpt.com
- 15Python standard library support in Workers
developers.cloudflare.com
- 16Cloudflare Python Workers are now generally available
simonwillison.net
- 17Gemini 3.8 TTS Playground
simonwillison.net
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
