Astra for Law, Anthropic's R&D automation index, parallel Claude Code projects, and the unredacted NYT filings: four AI measurements and their workings

Astra for Law, Anthropic's R&D automation index, parallel Claude Code projects, and the unredacted NYT filings: four AI measurements and their workings

Four developments from 17 September 2026 — OpenAI's Astra for Law and the benchmark score it published, Anthropic's index of how much of its own model research Claude now leads, the redesigned Claude Code Projects, and the newly unredacted filings in The New York Times' suit against OpenAI and Microsoft — each with the measurement behind the claim and the caveat its publisher attached.

Four documents published on 17 September 2026 each put a number on work that people used to do. OpenAI began selling a version of its most capable model configured for legal research, and printed the score it earned against a research benchmark. Anthropic published an index of how much of its own model development Claude now leads, next to the rate at which its monitors stop its agents. Anthropic also rebuilt Claude Code so that a single conversation dispatches a fleet of cloud sessions. And previously redacted material in The New York Times' copyright suit against OpenAI and Microsoft came out from under seal, quoting the companies' own employees about where the models' training data came from.
All four are checkable, and each publisher said where its own number stops being reliable.
DevelopmentWhat changedAction window
Astra for Law — 17 SeptemberOpenAI packaged GPT-6 Astra with a legal search index and legal instructions, and published a benchmark result for the complete setup. 1Available now through a gated program in ChatGPT and Codex; the API is not open yet.
Anthropic's pace measurements — 17 SeptemberAnthropic published three measurement methods and a snapshot of each: Claude leads 26% of its model R&D work, and about one in 47,000 agent decisions was blocked in August. 2A snapshot, not a trend; the company says it will report again.
Claude Code Projects — 17 SeptemberOne project conversation now coordinates parallel cloud sessions that keep running after you close your laptop. 3Public beta on Pro and Max; the enforced ceiling is 200 new threads a day.
Unredacted filings — 17 SeptemberMaterial unsealed in The New York Times' case quotes Microsoft's own staff calling the training-data practice "theft" and warning of a "doom loop". 4The exhibits behind the quotes remain sealed; the fair-use question is still open.
OpenAI introduced Astra for Law on 17 September: GPT-6 Astra, its latest and most powerful model, packaged with a legal search index and a set of instructions for legal analysis and writing, and offered to law firms and legal software companies as a foundation to build their own products on. 1
OpenAI's announcement page headed Introducing Astra for Law, dated September 17, 2026
OpenAI's own announcement of Astra for Law, published 17 September 2026. 1
The number OpenAI chose to publish is the one worth keeping. The company tested the complete setup on 200 US legal research questions drawn from the private validation set of Vals AI's Legal Research Bench, a benchmark that scores how well a system finds the relevant sources and passages and how well its answer meets the evaluation criteria. At the highest reasoning effort for both systems, Astra for Law passed the benchmark's overall correctness check on 54.0% of the questions. GPT-6 Astra using web search alone passed 38.7%. OpenAI describes that as a 40% relative improvement. On case-law questions the legal setup found 24% more reference cases, and it retrieved up to 54% more relevant passages from the correct court opinions when both systems ran at the same reasoning effort. 1
The comparison names its control, which is what makes it readable: one model, one set of questions, web search against a legal index. It also means the strongest configuration OpenAI sells got slightly more than half of the benchmark's questions right, and a cheaper setting will score lower. 1
The index is the part that changes the answer. It searches US case law, statutes, regulations, court rules and administrative decisions across more than 230 million URLs, with sources added daily, and it takes in the case-law collection of Free Law Project's CourtListener, which Free Law Project describes as covering more than 99.9% of published US precedential case law. 1
Access arrives gated. Astra for Law opens through a Trusted Access Program for eligible law firms. Firms that qualify get zero data retention on the API, and their ChatGPT Enterprise usage is excluded from human review by default. OpenAI says it is working with Latham & Watkins on information permissions, ethical walls, client instructions and firm oversight, and names Sullivan & Cromwell, Ropes & Gray and Cooley among the firms whose engineers built internal tools alongside OpenAI's. The model appears in the picker as GPT-6 Astra Law and in the API as gpt-6-astra-law, which is still to open. 1
Twenty-six partner-built plugins connect the model to software firms already run. iManage saves a negotiation brief to the matter file, Intapp surfaces activities that may need a time entry, DeepJudge pulls prior deals into a comparison, and Thomson Reuters is bringing HighQ matter context into ChatGPT and previewing a CoCounsel Legal connector. Harvey and Legora, two legal AI companies, will build on the model, and ChatGPT for Word became generally available the same day. 1
OpenAI also printed a comparison of its own making. Given the same litigation prompt, Astra for Law returned two closely matching precedents while Claude Fable 5.1 returned a holding that had been reversed on appeal; on a transactional prompt, the rival model reported finding no such case. OpenAI wrote and ran that test itself, and a buyer's own matters are the check that counts. 1

Anthropic measures how much of itself Claude builds

Anthropic published three measurement methods on 17 September, each with the snapshot its own method produced: how much of the company's model development Claude performs, how far its monitors reach into what its agents do, and where its computing power goes. The company frames all three as an answer to a world that cannot see inside a frontier lab. 2
The first is an index. Anthropic catalogues every kind of AI research and development task its staff do, then rates each one on the Automation Level scale that Epoch AI, an independent nonprofit that tracks the technology, proposed. The scale runs from AL0, where no AI is involved, to AL5, where a model operates fully autonomously with no human in the loop. AL3 means the model collaborates: it does large chunks of work under close human direction. AL4 means it leads: it completes most of a task from a high-level prompt while a human supervises the result. 2
As of August 2026, Anthropic reports that Claude leads 26% of its model R&D work, that the share of work at or above the collaborating level is above 90%, and that every piece of work measured still kept a human involved. The leading share was under 1% in February. 2
Stacked area chart titled Claude now leads 26% of model R&D work, showing six automation levels from August 2025 to August 2026
Anthropic's own chart of its R&D Automation Index, plotting the share of model R&D work at each level of Epoch AI's automation scale from August 2025 to August 2026. 2
Anthropic is explicit about how far the index can be trusted. For each week of July a Claude research agent reviewed a random 20% of staff in the departments that build models, which produced roughly 15,000 granular tasks; a Claude model organised them into a tree of 542 nodes; a Claude agent researched how each kind of work is done, and an independent Claude judge assigned the level. The judge agreed with the staff who own those work areas about as often as people agreed with each other — 59% exact agreement against 35% between humans — and the two landed within one level of each other 97% of the time. 2
Two further limits sit on the index. The basket of tasks is frozen at a July baseline, so a rising number says the work people were doing in July is being automated, and the question of whether new kinds of work have appeared is left to a separate comparison. Anthropic ran that comparison against a January basket and found no rise in novel tasks. The judge is also one of the company's own models, which Anthropic concedes could repeat the errors of the model it is checking. 2
Reuters reported the same figures from outside the company, and attributed the scale to Epoch AI, the independent nonprofit that developed it. 5
The second measurement covers the agents. Anthropic counted roughly 30,000 agents doing research and engineering work on its most-used internal platform at any one time in August. Every action those agents take passes through an online monitor before it runs, usually within seconds, and a second monitor ingests every action after the fact. Reviewing more than a billion decisions from that month, the company found 0.002% of them, or about one in 47,000, were blocked. The offline monitor flags around 100,000 transcripts a week; classifiers and an initial review resolve most of those, and roughly 50 a week reach a person. 2
Anthropic states its own limit on that one as well: the company has been monitoring agents for months, and describes its picture of what agents can do as likely to be incomplete. 2
The third measurement is computing power. In a snapshot week from 13 to 20 July, about 6% of the compute that went to AI research and development went to safety work, and about 12% of the compute behind AI-driven research did. Anthropic calls both figures conservative, since work that advanced capability and safety equally was counted as capability work. The company also notes what a compute share measures: a more efficient safety classifier lowers the safety portion without the lab doing less safety work. 2
Anthropic argues any frontier developer could publish the same three measurements, and says it plans to place independent third-party evaluators inside the company with access comparable to its internal risk teams. 2
The same day, in Scotland, the King convened representatives of Nvidia, Google DeepMind, OpenAI and Anthropic at Dumfries House, together with Britain's AI minister and the Ditchley Foundation, to consider whether a shared set of principles could guide how the technology is applied. The palace published the King's speech and its account of the discussion, which records delegates considering whether a shared set of principles could guide how the technology is applied; the account reports that consideration and no settled outcome. 6

One Claude Code project now runs its own sessions

Anthropic redesigned Projects in Claude Code on 17 September, turning what used to be a folder holding some files and one chat into a single long-running conversation that acts as a coordinator. You describe what needs doing, and Claude decides what becomes a thread. 3
Each thread is a full Claude Code cloud session with its own branch and its own copy of the repository. Threads run in parallel, open pull requests, run your tests and report back to the conversation, and they keep going after you close your laptop; you can check on them and steer them from a phone. When two threads touch the same code, the overlap surfaces as an ordinary merge conflict on the two branches. 37
A product card titled Start new threads listing three proposed tasks with Start buttons and a Start 3 threads button
A Claude Code project proposing a batch of threads; each proposal becomes its own cloud session with a branch and a copy of the repository. 3
Every new thread starts on the same footing: the project's repositories and uploaded files, its project instructions of up to 16,000 characters, and project memory that Claude writes and reads through an index file named MEMORY.md. A Library tab collects the files you added and the files the threads produced. 7
The cost model is where a project separates from a single session. Each running thread is a full session, so a project draws on the same plan limits as your other Claude Code work and reaches them faster; Anthropic's documentation warns that Pro subscribers should expect to hit their limit sooner on days they run one. The enforced ceiling is 200 new threads per day across your projects, and a project stays inside your plan limits unless you have turned usage credits on. An idle thread wakes and spends again when a check fails or a review comment lands on its pull request. A thread that reaches a limit waits and continues by itself when the limit resets, so work left running on a Friday can begin spending the next window without telling you. 7
A new project runs Opus everywhere, at high effort for threads and low effort for the coordinating conversation, which is the most expensive default the product offers. 7
Two behaviours are worth knowing before handing over real work. Permission rules, hooks and environment variables are read only from the settings file in the directory a thread starts in, so a single-repository project inherits them and a multi-repository project falls back on its project instructions. And approvals happen inside a thread: when a thread needs your sign-off it waits there, and a go-ahead given to the coordinator stays in the coordinator's conversation. 7
Access is narrow for now. Projects are in public beta on the Pro and Max plans, starting with accounts that already use cloud sessions and have no existing projects; Team and Enterprise follow later. Threads work on GitHub repositories and on files uploaded to the project; a local database, a device emulator or an API behind a VPN sits outside their reach. Threads run in Anthropic's cloud today, and running on your own machine is described as coming soon. 37

The filings quote the companies' own staff about the training data

Material that had been redacted in The New York Times' copyright suit against OpenAI and Microsoft was unsealed on 17 September, three years after the paper filed it. The unsealed passages describe how the companies are said to have gathered the publishers' work and what their own employees wrote about it at the time. 4
Microsoft's director of applied science, Brent Hecht, called mass scraping "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history" in a January 2023 memo, according to the filing. A presentation he wrote a year later described a "doom loop" in which AI answer engines erode the publishers whose work trains them, and said it was "highly unusual that an end-product threatens the economic foundations of its essential suppliers." Microsoft's own data, the filing says, showed its Copilot answer engine cutting click-through rates to nytimes.com by as much as 93% against traditional Bing search. 4
Other passages work on the fair-use defence from the publisher's side. Satya Nadella testified in a deposition this year that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training," and said he would have required OpenAI to retrain its models had he known the company trained on paywalled material. OpenAI's head of ChatGPT, Nick Turley, wrote internally that publishers face an "existential threat" from a product that is "largely substitutive." The filing describes a researcher telling OpenAI's president, Greg Brockman, about a "hack to get around nytimes paywall," and his reply: "ah nice." 4
The scale of the copying is the part with figures attached.
Dataset the filing describesWhat it is said to hold
OpenAI's mid-training datasetsMore than 91,692 copies of works published by The New York Times, the New York Daily News and the Center for Investigative Reporting. 4
A dataset derived from Common CrawlMore than 2 million documents from nytimes.com alone. 4
Project Mango, a joint data initiativeAt least 160,903 unique works from the news publishers. 4
Two limits decide how much weight the material carries. Much of the new information comes from The New York Times' own brief; the exhibits it draws on remain sealed, and the quotes appear without the surrounding context. TechCrunch's requests for comment to OpenAI and Microsoft went unanswered. Separately, judges have leaned toward the AI companies' side of the fair-use argument, and the Trump administration filed a brief on 2 September supporting OpenAI's position. 4

What to check before you rely on any of it

  • Ask for the score on work you already understand. OpenAI's 54.0% comes from 200 questions in someone else's validation set; a firm should ask for the retrieval rate on matters where it knows the right answer. 1
  • Find out which settings produced the number. Both the 54.0% and the 38.7% comparison come from the highest reasoning effort, and the index adds sources daily, so a cheaper setting and a later date will both read differently. 1
  • Check whether the vendor is grading itself. OpenAI ran and published the prompt that shows its legal model outperforming a rival's, and Anthropic's automation index is scored by Anthropic's own models; both companies name that weakness themselves. 12
  • Ask what a lab's index leaves out. Anthropic's basket of tasks is frozen at July 2026 and its compute figure covers a single week, so a reading a year from now will still be describing the work people did this summer. 2
  • Ask how much of your agent traffic a monitor sees. Coverage, review latency and block rate are the three figures Anthropic publishes for its own agents; the same three are the ones to request from any vendor whose agents act on your systems. 2
  • Price a project before leaving it running. Threads draw on plan limits faster than a single session; the ceiling is 200 new threads a day, and an idle thread wakes and spends again when a check fails or a review comment arrives. 7
  • Check what a multi-repository project drops. Permission rules, hooks and environment variables come only from the directory a thread starts in, so a multi-repository project falls back on its project instructions for the equivalent rules. 7
  • Read the sealed exhibits, or hold the question open. The quotes in circulation come from one party's brief and the documents behind them stay sealed, while the fair-use question is still being decided; both facts belong on the same page. 4

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

More from this channel