Typed decisions, independent verifiers, and five agent-builder signals from September 19-26

Typed decisions, independent verifiers, and five agent-builder signals from September 19-26

Five in-window builds show agent builders writing limits into their systems and shipping the check that enforces them: a security-audit skill whose verifier never found the bug, a model that returns typed decisions, a key service that keeps keys out of the agent's context, a runtime that publishes what it lacks, and speech built from a staged conversation.

The week in one line

Between September 19 and September 26, five agent-building signals landed, and each one writes down what may cross the boundary between a model and the rest of the system. The ones worth copying also ship the thing that checks the rule.
Cloudflare's security-audit skill sat near the top of GitHub's weekly Trending board, and builders who pointed it at their own code kept reporting the same useful half: a second wave of agents whose only job is to disprove the first wave's findings 1. TypeSafe AI's Jev returns typed numbers where a chat model returns text, and hosting it pushed Simon Willison to add a capability flag to his llm library for models that accept no conversation history 2. That library also gained a small key service, so an API key can be installed on a remote agent machine and stay out of the agent's context 3. Cloudflare's Python Workers reached general availability with a published list of the Python features its WebAssembly runtime lacks 4. Google released speech models whose unit of work is a staged conversation between named voices 5.
Five signals follow, each with the artifact to open and one test to run.

The comparison map

SignalBuilder or projectIn-window dateInspectable artifactFirst adoption test
The verifier never found the bugCloudflare, security-audit-skillSep 26, weekly Trending readfindings.json checked against report-schema.json, plus coverage-ledger.json 1Read the rejected verdicts before the confirmed ones.
A model whose output type is fixed before the callTypeSafe AI, JevSep 21-22A query of typed questions returning probabilities and confidences 2Rescore 100 cheap retrieval hits and compare the new order.
A key the agent never readsSimon Willison, llm-keys-ui 0.1Sep 20A local web form, then llm keys get inside a shell command 3Route the next key on a remote box through the form.
A runtime that publishes what it lacksCloudflare, Python WorkersSep 21The Workers page for the Python standard library 4Read that page before porting a worker that uses threads.
Speech whose unit is a conversationGoogle, Gemini 3.8 Flash TTSSep 23A script with named speakers, per-line delivery styles, and 2,000+ voices 5Direct one two-speaker exchange line by line.

Five signals to inspect

1. The verifier never found the bug

Cloudflare published security-audit-skill, a coding-agent skill that turns a coding agent into a security auditor. The skill runs six phases: reconnaissance, coverage-led hunting, candidate validation, structured output, independent record verification, and target-neutral reporting 1. The repository carries the MIT license, holds 14 commits, and its newest commit dates from September 14, before this window 1.
During the window, other people began pointing the skill at their own code. It reached second place on GitHub's weekly Trending board with about 9,500 stars gained over the week, read on September 26 6. On September 25 the team behind the Kaspa wallet Kaspire published the result of a four-hour audit 7. The audit returned mostly minor issues and one real finding: the site's post-build gate treated a payload's SHA-256 as authority for a staged Android update APK, skipping verification of the envelope's RSA signature 7.
Loading content card…
Builders keep naming the second wave as the part that pays. Every unique candidate goes to a fresh verifier whose job is to disprove it, and fresh agents re-verify the final source claims, so the agent that checks a finding never found it 1. Verdicts land in three buckets: confirmed carries a complete source trace and a bounded observed result, needs_validation carries one exact unresolved fact while the skill withholds a severity, and rejected records a disproof 1.
Two zero-dependency scripts police the contract. validate-findings.cjs checks findings.json against report-schema.json in phase four and again after every phase-five replacement, and validate-coverage-ledger.cjs checks coverage-ledger.json once the ledger exists and after each later update 1. Repeat runs compose: the skill reads prior ledgers and findings to target gaps and revalidate changed source 1.
Cloudflare's own test numbers set the expectation. In its test runs, a single pass found roughly half of the vulnerabilities that repeated passes found in total, and the findings skewed toward the simpler cases 8. That post is dated June 18 and describes the fleet-wide harness this skill seeded, so the coverage figure is Cloudflare's observation from its own tests.
One prerequisite decides whether the workflow can do anything at all. Running target-controlled builds, tests, browsers, or fuzzers requires an OS-enforced sandbox that disables external networking, uses a sanitized allowlisted environment, enforces resource limits, and permits writes only to assigned scratch paths. With those controls absent, the skill keeps a lead as needs_validation and stops there 1.
Try it. Install the skill, point it at the smallest repository you own, and read the rejected records first. Then count the needs_validation leads that arrived with an exact unresolved fact attached, because those are the ones a second run can close.

2. A model whose output type is fixed before the call

TypeSafe AI announced Jev on September 15, before this window, as the first of what it calls System One models: text in, floating-point numbers out 9. The window holds what builders did with it. Simon Willison published notes on September 21 2, and Latent Space ran an episode with TypeSafe's founder the same day 10.
A call composes a state object, which can be a string, a list of strings, or a set of name-value pairs, and sends it with one or more questions. Yes/no questions return a confidence between 0 and 1, choice questions return a probability distribution across the options you supplied, and score questions return a float along a described range. Questions are evaluated in parallel, so a batch takes about as long as a single question 2.
TypeSafe's own numbers are the first thing to test. The company lists $0.042 per million input tokens with output billed at nothing, and end-to-end latency of 70 to 500 milliseconds against 3 to 329 seconds for frontier chat models on the same class of task 9. TypeSafe also says schema matching is guaranteed, which is where its claim of no hallucination comes from 9.
A scatter chart titled Average of 4 workflows: accuracy vs cost, plotting accuracy on the vertical axis against cost per workflow on a logarithmic horizontal axis; a pink diamond labelled Jev sits at the far left at roughly 67 percent accuracy and a cost below $0.001 per workflow, while labelled OpenAI, Anthropic and Fireworks models cluster further right at higher cost
TypeSafe AI's own workflow-eval chart, published with the Jev announcement on September 15, 2026. Each model runs the same four workflows, and the reference answer is the average of GPT-6 Astra and Fable 5.1, so the comparison baseline is two named chat models and the plotted numbers are the vendor's. 9
Hosting the model needed a new contract in Simon Willison's llm library. Release 0.36, published September 22, lets a model plugin declare supports_conversation = False; the library then raises llm.ConversationNotSupported when such a model receives assistant or tool history, and llm chat rejects the model before a session starts 11. The first plugin to use the flag is llm-typesafe, the plugin that wraps Jev 12. The same release wraps reasoning traces in tags inside the Markdown output of llm logs 11.
The design leaves a gap a builder has to plan around. A confidence score arrives as a bare number, so a spam verdict or a ranking comes with fewer surfaces to inspect than a chat model's written justification. Simon ran one experiment that scored every city in the San Francisco Bay Area on a "Good city?" question and got Cupertino on top and East Palo Alto at the bottom 2. TypeSafe's own jaggedness guidance reports the model is weak on numbers, dates, and adversarial content 2. Bias in a single float has to be found by experiment, and the price makes hundreds of probes affordable.
Try it. Take the classification you trust least in your own pipeline, send 100 candidates to a decision model as a score question, and compare its order against the one you ship today. Then ask what would have to be true for you to act on a bare number.

3. A key the agent never reads

Simon Willison published llm-keys-ui 0.1 on September 20 for one narrow problem 3. He had started using Codex Remote to run coding agents on several machines while controlling them from his phone, and he did not want to paste API keys into an agent session 3.
The plugin serves a small web interface on the machine. Running uvx --with llm-keys-ui llm keys-ui --all makes the agent report the server's URLs, including local network and Tailscale addresses, and a form at one of those addresses saves new keys 3. The interface never displays an existing key value, and later work reaches the key through a command such as llm keys get anthropic inside a shell command 3. The release is tagged 0.1 on GitHub 13.
The advertised addresses are the part to weigh. Because the server lists the machine's local network and Tailscale addresses, anything on those networks can reach the form, so the reachability that makes remote setup convenient is the same reachability worth restricting. Codex Remote itself is documented as a way to run coding agents on a machine while driving them from the ChatGPT app 14.
Try it. Next time you configure a key on a machine that an agent controls, run the key service there and hand the key over through the form rather than pasting it into the conversation. Firewall the port to the networks you actually trust.

4. A runtime that publishes what it lacks

Cloudflare's Python Workers reached general availability on September 21, after two years in preview 4. The runtime compiles Python to WebAssembly through Pyodide and runs it inside workerd, Cloudflare's V8-based runtime 4.
Three changes make the runtime usable for ordinary Python work. Bindings now convert types inside the runtime, so self.env.QUEUE.send({"key": "value"}) works, replacing the to_js glue that used to sit at the RPC boundary 4. A workers.asgi connector runs FastAPI applications and workers.wsgi runs Django or Flask, with the Workers platform taking the place of a web server 4. Hyperdrive reaches Postgres and MySQL because Cloudflare implemented socket system calls on top of the Workers connect API, where a WebAssembly sandbox normally stubs the POSIX networking calls out 4.
The published limits are the part to read before porting anything. The Workers page for the Python standard library states that both multiprocessing and threading are non-functional in the WebAssembly virtual machine 15. Simon Willison flagged the same two features while covering the release 16.
Simon also described the local development story. Cloudflare's pywrangler tool runs a full local simulation of the platform, which for him meant executing Python through Pyodide in WebAssembly inside V8 inside a 123 MB workerd binary that landed under node_modules 16. Cloudflare credits the release to Gyeongjae Choi, Dominik Picheta, and Hood Chatham, two of whom are Pyodide core maintainers 16.
For an agent builder, this is somewhere agent-written Python can run, with a written list of what will fail there. Check that list against the libraries you intend to let an agent use.
Try it. Pick the two standard-library features your code depends on most, and read the standard library page for Python Workers before you let an agent generate a worker. Port the smallest one first and see which import breaks.

5. Speech whose unit is a conversation

Google released two speech models on September 23. Gemini 3.8 Flash TTS handles voice design and character work, and Gemini 3.8 Flash-Lite TTS targets high-volume dubbing and voice agents 5. The voice library holds more than 2,000 production voices across more than 100 languages, and a voice can be replicated from a 30-second sample with consent verification, SynthID watermarking, and C2PA credentials attached to the result 5.
The interface change matters more than the voice count. Both models accept a script in which speakers carry names and each line carries its own direction, and the API supports native two-speaker scene staging with conversational turn-taking, plus non-verbal cues and active-listening interjections for timing 5. Google's post reports first and second place on Hume AI's Overall Quality Index and first on Hume AI's Voice Design Benchmark at 71.4, alongside top positions in blind preference tests on Voice Arena; those ratings belong to the benchmark owners, and the post is where they are collected 5.
Screenshot of a web page for composing multi-speaker speech, with a Compose panel containing a Conversation toggle, a Cast section pairing the speaker names Gus and Pearl with voices, and numbered dialogue lines that each carry a speaker, a delivery style and the spoken text, beside a Connection panel holding an API key field and a model selector
Simon Willison's screenshot of his Gemini 3.8 TTS playground, published September 23, 2026. The Compose panel shows the API's unit of work: a cast of named speakers with one voice each, and per-line delivery styles that serialize into speech_metadata annotations on each dialogue line. 17
Simon built that playground the same day, using the API's open CORS policy to keep it a bring-your-own-key page that talks to Google directly 17. His measurements are one person's, on one day: about 20 seconds to generate 1 minute 18 seconds of audio with Gemini 3.8 Flash TTS, for 2.74 cents, and the page that loads the voice catalogue reported 2,089 voices 17.
The interesting part for an agent builder is where the prompting sits. A voice agent that once assembled a string per turn can now declare a cast before the conversation starts, then hand each line a speaker and a delivery style. That is the same shape as a typed tool call: the model gets fewer choices about who is speaking and how.
Try it. Write a two-speaker exchange of six lines, give each speaker a voice and each line a delivery style, and generate it in the playground. Then check whether your own voice pipeline can carry a cast through to the request, or whether it flattens everything to one voice per call.

What to try this week

  1. Audit one repository and read the negative results. Point the security-audit skill at the smallest repository you own and start with the rejected verdicts.
  2. Give one classification job to a decision model. Send 100 candidates as a score question and compare the order it returns with the one you ship today.
  3. Move one API key out of the agent's context. Run the key service on a remote agent machine and hand the key over through the local form.
  4. Read a runtime's standard-library page before generating code for it. Check the two features you rely on most against what the runtime documents as missing.
  5. Direct a cast, then the lines. Give one audio request a cast, a voice per speaker, and a delivery style per line.

The pattern

Five artifacts, and a limit written into each one. The question that separates them is what enforces that limit.
Three are enforced by machinery that refuses to carry on. A findings.json file that fails report-schema.json stops the phase 1. A model plugin marked supports_conversation = False makes the llm library raise before it sends history to a model that takes none 11. An API key handed over through the local form stays on the machine and out of the conversation 3.
Two rest on the reader. The Python Workers limits live on a documentation page, and code that reaches for a thread will fail when it runs 15. A replicated voice carries consent verification, a SynthID watermark, and C2PA credentials, and the voice catalogue describes what can be generated, so any particular voice is your own sample to test 5.
The decision model inverts the problem: its confidence score is calibrated by design and arrives as a bare number, so the only check you get is an experiment against your own task 2.
The security-audit skill puts the check somewhere else entirely: a second agent whose only instruction is to disprove the first. That design costs a sandbox before it will execute anything, and Cloudflare's own test observation says one pass finds about half the bugs that repeated passes find 1.
The rule to carry out of this week: when a tool states a limit, find the thing that enforces it. A schema, a raised exception, and a key held outside the transcript are checks you can build on. A confidence score, a catalogue count, and a documentation page are claims you inherit, and each one becomes visible the first time something runs.
This issue covers September 19 through September 26; the next one arrives next Saturday.

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content