
Astra for Law, Claude's 26% of Anthropic's R&D, and a Cell cover won by a 9B model
A daily brief on OpenAI's law-firm product built on GPT-6 Astra, Anthropic's first running numbers on how much of its own research Claude performs, an open aging-biology toolkit winning a Cell cover with a 9-billion-parameter model, Huawei's accelerated AI chip, the UN's agent-ready statistics platform, and the day's compute funding and price increases.
This edition covers AI developments published between September 17 at 08:15 and September 18 at 08:15 in Dhaka. OpenAI turned its frontier model into a product for law firms. Anthropic published the first running numbers on how much of its own research its models perform. A peer-reviewed journal put an aging-biology toolkit on its cover. The UN made its statistics readable by AI agents, Huawei moved its next AI chip forward, and the companies that rent and build compute raised prices and money within hours of each other.
Quick scan
| Development | What changed | Why it matters | What to watch |
|---|---|---|---|
| Astra for Law | OpenAI wrapped GPT-6 Astra with a U.S. legal search index and legal-specific instructions, sold under new access controls. 1 | A frontier lab is now selling configuration, not just a model, into a profession that bills by the hour. 2 | Whether it reaches the general API, and what law firms get to verify about it. |
| Anthropic's R&D automation index | Anthropic said Claude "leads" 26% of its model R&D work, up from 1% in March, and published the method. 3 | The first recurring public number on AI building AI, which is the quantity the pacing debate turns on. 4 | Whether other labs publish comparable numbers, and whether third parties can check these. |
| Insilico's longevity toolkit | A Cell cover study released an open aging-biology benchmark, a family of 0.6B–9B models, and an agentic target-discovery platform. 5 | A 9-billion-parameter open model scored highest of the 26 systems tested. 6 | Whether the nominated target genes survive experimental follow-up. |
| Life Sciences Verification Program | Anthropic opened applications to let vetted biology teams past safeguards that block that work by default. 7 | Enforcement moved from blocking each request to reviewing behaviour after the fact. 7 | How many organisations the review admits, and what happens to a flagged case. |
| Huawei's Ascend 960DT | Huawei pulled the chip's launch forward from the third quarter of 2027 to the first. 8 | China's main domestic accelerator supplier is compressing its roadmap a week before a Trump–Xi meeting. | Whether the system Huawei showed alongside it matches the scale it described earlier. |
| UN System Data Commons | The UN replaced its statistics portal with a Google-built platform that supports agent connections. 9 | It answers a measured failure: six leading models averaged 21.2% accuracy on global development indicators. 9 | Whether the numbers hold up when someone else runs the test. |
| Compute capital and prices | Crusoe raised $3.9 billion, CoreWeave priced a $3 billion convertible sale, and Nebius raised GPU rental rates. 1011 | The price list is the one that reaches a customer directly. | Whether the increases hold when the next wave of capacity arrives. |
OpenAI packages its frontier model for law firms
Astra for Law combines GPT-6 Astra with a legal search index and instructions for legal analysis and writing. OpenAI introduced it on September 17 as a foundation for law firms and legal software companies to build on. 1 The index searches U.S. case law, statutes, regulations, court rules and administrative decisions across more than 230 million URLs, with material added daily. It includes the Free Law Project's CourtListener collection, which covers more than 99.9% of published U.S. precedential case law. 1
The measured gain comes from a comparison OpenAI ran against its own model without the configuration. On 200 questions from the private validation set of Vals AI's Legal Research Bench, Astra for Law passed the overall correctness check on 54.0% of questions at the highest reasoning effort, against 38.7% for GPT-6 Astra using web search alone. 1 On case-law questions it found 24% more reference cases. On the audited set of passages it retrieved up to 54% more relevant passages from the correct court opinions. 1 Vals AI's benchmark and OpenAI's configuration both belong to parties with an interest in the result, and OpenAI published the comparison itself rather than an independent reproduction.
Legal research is the first step in the product's own account of the work. OpenAI says custom instructions then carry that research into arguments and deal terms. It printed a side-by-side example in which Astra for Law and Claude Fable 5.1 answered the same memo prompt, and reports that the competing model cited a holding that had been reversed on appeal while its own model found two matching precedents. 1 That comparison is OpenAI's, on a prompt OpenAI chose.
Access arrives with the configuration. Firms join a Trusted Access Program that includes zero data retention on the API and excludes ChatGPT Enterprise usage from human review by default, and Latham & Watkins is working with OpenAI on information permissions, ethical walls and firm oversight. 1 The model appears in the picker as "GPT-6 Astra Law" and in the API as
gpt-6-astra-law, reaching selected law firms first; Harvey and Legora are building on it through the API. 1 OpenAI also shipped 26 partner-built plugins covering tools firms already use, among them iManage, Intapp, DeepJudge and Thomson Reuters, and made ChatGPT for Word generally available the same day. 1 Reuters framed the release as an escalation in the contest for law firm users. 2What to watch: whether the configuration reaches the general API, and whether a law firm or an outside evaluator reproduces the Legal Research Bench gap on documents the benchmark does not contain.
Anthropic publishes a running scoreboard for AI building AI
Anthropic's measurement of how much of its own AI research and development Claude performs is the most concrete number to arrive in the pacing debate, and the company published the method alongside it on September 17. 3 The measure catalogs every kind of AI R&D task at the company and rates how automated each one is. It then weights each task by the person-time that task consumes, producing an index built on roughly 15,000 tasks inside 542 task categories. 3
The rating scale comes from Epoch AI, an independent nonprofit that tracks AI progress. It runs from AL0, where no AI is involved, to AL5, where AI acts fully autonomously. At AL4 the model "leads": it completes most of a task from a high-level prompt while a human supervises. 3
As of August 2026, Anthropic says Claude leads 26% of its model R&D work, and that more than 90% of that work sits at or above the level where AI collaborates. Fully autonomous work, AL5, appears nowhere in the measured set. 3 Anthropic's own chart puts the AL4 share at 1% in March, 3% in April, 12% in May, 14% in June, 22% in July and 26% in August. The 90% measurement intervals widen to roughly four percentage points in the recent months. 3

Two further measurements cover oversight and money. Anthropic says roughly 30,000 agents were doing research and engineering work on its most-used internal platform at any one time in August, and that every action those agents took passed through an online monitor before execution. Of more than a billion decisions that month, about one in 47,000 was blocked. 3 Offline monitors flagged roughly 100,000 transcripts a week, most of which classifiers resolved, with about 50 a week reaching human review. In a sample week in July, about 6% of the compute that went to AI research went to safety work, rising to 12% for research the AI itself performed. 3
Anthropic states the limits of its own numbers. A Claude model judged the work, and Anthropic checked that judgment against staff ratings of areas those staff own. The model agreed with a human on an exact rating 59% of the time, while humans agreed with each other 35% of the time; the two landed within one level 97% of the time. 3 The task basket is frozen at July 2026, so a rising index means that the work people were doing then is being automated; Anthropic says it compared baskets from January and July and found the structure of the work stable. 3 No shared methodology exists between labs, which is what would make the numbers comparable across them. 4
What to watch: whether a second lab publishes the same three measures, and whether the independent evaluators Anthropic says it will embed get to publish their verification.
A Cell cover goes to small models trained on aging data
Insilico Medicine released an open toolkit for AI in aging research in a study published as the cover feature of Cell's September 17 issue, with collaborators from Liquid AI, the Buck Institute for Research on Aging, and Harvard Medical School and Brigham and Women's Hospital. 5
The benchmark exists because general-purpose models had gone largely untested on real aging data. LongevityBench evaluates reasoning across five domains: clinical data, genetics, epigenetics, transcriptomics and proteomics. 5 The researchers tested 18 frontier systems from OpenAI, Google, Anthropic, xAI, DeepSeek and Moonshot AI. The leading model changed with the data type, performance shifted with how a question was phrased, and predicting biological age from omics measurements defeated even the largest models. 5
The five Longevity-LLMs that followed range from 0.6 billion to 9 billion parameters. They were fine-tuned on aging-specific clinical and multi-omics data and built on open architectures from Liquid AI and Alibaba. 5 The best of them, L-Qwen3.5-9B, recorded the highest overall score of the 26 systems in the study, ahead of Google's Gemini 3.1-Pro, and the smallest model at about 0.6 billion parameters beat most of the frontier systems. 5 Curated training data and domain tuning produced that result on a fraction of the compute a frontier model consumes, which also means a research lab can run its own copy on local hardware.
The same study put the model to work. Longevity Claw, an open-source agentic platform, chains aging-specific tools together, including gene-set enrichment analysis, biological aging-clock calculation and evidence retrieval. It works through multiple steps on its own rather than answering one question at a time. 5 Run across 14 recognised hallmarks of aging, it nominated 328 genes as candidate targets. Those candidates showed enrichment of up to 5.6-fold against an independently published reference set of experimentally supported aging targets. 5 One nominee, KDM1A, had already been validated elsewhere as a target that extended lifespan in the worm C. elegans. 5 That overlap with a known target list is a check on the platform's output, and it is a check rather than proof: the nominated genes still need experimental work before they count as targets.
The models, the benchmark data, the evaluation code and the agent platform are published openly, on Hugging Face, GitHub and a public leaderboard. 6
Anthropic opens a vetted route around its biology safeguards
Anthropic introduced the Life Sciences Verification Program on September 17, giving verified life science teams access to its Mythos, Opus and Sonnet models with safeguards adjusted to allow drug discovery, research biology, clinical development and manufacturing work that its generally available models block. 7
Applicants go through a review of research credentials, security standards and ethical research oversight. 7 Two grant types follow. Standard Use covers most biology work for an entire team, is renewed annually, and applies to Mythos 5.1, Opus 5 and Sonnet 5 as well as future models. High-risk Use removes the remaining life-sciences blocks, applies to a single research project rather than a team, and must be renewed every six months. It is available today for Opus 5 and Sonnet 5, while high-risk access to Mythos stays limited to a small set of entities while Anthropic works with the U.S. government. Cyber safeguards remain in place under both grants. 7
The enforcement model changed with the access. Anthropic monitors program traffic after the fact against the use cases each organisation declared, rather than rejecting a request at the moment it is made, and it retains flagged data for 30 days to do that work. 7 The retention is compartmentalised and cannot be used for model training. Anthropic names three threats it designed against: access compromise through malware or account takeover, rogue or coerced insiders, and agents taking dangerous actions over long tasks or in swarms. 7 As a beta, the program is unavailable to organisations with a business associate agreement, which puts PHI data in separate non-HIPAA organisations, and it does not yet cover individual plans or third-party platforms. 7
What to watch: how many organisations clear verification, and what the first remediation case under the new monitoring looks like.
Huawei moves its next AI chip up a half-year
Huawei said at its Huawei Connect conference on September 17 that the next-generation Ascend 960DT AI chip will arrive in the first quarter of 2027, pulled forward from a plan for the third quarter of that year. 8 David Wang, the company's rotating and acting chairman, announced the revised timeline. 8 The announcement landed a week before U.S. President Donald Trump and Chinese President Xi Jinping meet in Washington on September 24. 8
Huawei's approach to scale is to run very large numbers of accelerators as one machine, under an architecture it calls Peerium and a processor-linking technology called UnifiedBus. 8 The Atlas 950 SuperPoD and SuperCluster are the first systems on it, and Huawei says an Atlas 950 SuperCluster can connect up to 256,000 accelerator cards. 8 China analyst Rui Ma noted on X that Huawei had earlier described its Atlas 960 SuperPoD as scaling to 15,488 Ascend 960 chips, while this week's announcement described a system with 4,096 chips. 8 The chip arrives earlier on paper while the system around it shrank.
Huawei's rotating chairman Eric Xu used the same conference to place Chinese labs on the risk curve. He told reporters that the mainstream U.S. model providers, with their computing power, can perceive risks of AI that providers in China cannot yet perceive. 12 His conclusion from that is to develop faster: he said Chinese developers may need to speed up to a level where they too can feel risks from AI development, while balancing development against managing those risks. 12 Xu also said Huawei cannot produce enough AI computing equipment to satisfy domestic demand. 12 China is drafting a mandatory national standard for AI agent safety while it presses ahead with deployment. 12
What to watch: whether the September 24 meeting touches chip export rules, and how Huawei explains the gap between the SuperPoD it described earlier and the one it announced this week.
The UN's statistics become something agents can query
The UN system launched the UN System Data Commons, an open platform built on Google's Data Commons. It replaces the older UNData portal, where visitors browsed a database interface instead of asking questions. 9 The new site takes natural-language queries, and it supports the Model Context Protocol, the standard that lets AI systems connect to external data sources, so an assistant can pull figures directly. 9 Every dataset is validated with UN system statisticians, and the platform tracks where each statistic comes from, so a figure retrieved by an AI system can be traced back to its source. 13

The reason the UN did this is a measurement of how badly models handled the material before. A UNICEF benchmark of six large language models, run across more than 133,000 responses to questions about global development indicators, produced an average accuracy score of 21.2%, UNICEF's chief statistician João Pedro Azevedo told reporters. About three in five responses contained no usable number at all. 9 When the same questions were run again on the same model versions about two days later, models that gave a number both times returned the same number only about half the time. 9 The test covered GPT-4o and GPT-4o-mini, Claude Sonnet 4.5 and Haiku 4.5, and Gemini 2.5 Flash and Gemini 2.0 Flash. 9 The study is a UNICEF working paper headed for journal submission and has not been peer-reviewed. 9
Demand is already arriving from that direction. Visits to UNICEF's data site from people clicking links in ChatGPT answers rose 67% year over year between January 1 and September 14, accounting for 6.4% of sessions this year, and UNICEF estimates AI assistants now account for about one in ten visits. 9 Google.org provided $2 million in capacity-building funding and technical support through the UN Foundation, and the platform runs on a UN-governed instance intended to be maintained by the UN itself. 9 Twenty-six UN entities have committed, data from nearly 20 is available at launch, and the target is 80% of the UN system's statistical datasets by 2027. 9
What to watch: whether the platform's answers hold up when an outside group runs the same questions, and whether the 2027 coverage target survives 26 agencies' publication schedules. Prem Ramaswami, who leads Google's Data Commons team, told TechCrunch that a human should review outputs before they are cited, because models misread nuance. 9
Compute got better funded and more expensive on the same day
Three statements about the same demand curve landed within hours, and they carry different kinds of information.
| Company | What changed | Amount or rate |
|---|---|---|
| Crusoe | Series F funding; valuation up from $10 billion ten months earlier | $3.9 billion raised at a $30.9 billion valuation 10 |
| CoreWeave | Launched a convertible debt offering and said it is signing contracts at higher prices | $3 billion offering 14 |
| Nebius | Raised pay-as-you-go rates for selected Nvidia chips, effective October 1; second increase in three months | GPU rates up 17% to 21%; some CPU-only instances up 25%; memory offerings up about 41% 11 |
Crusoe's round was co-led by Atreides Management, Mubadala Capital and Valor Equity Partners, with Founders Fund, GIC, Nvidia, the Qatar Investment Authority, Radical Ventures and TPG participating. The company also added three directors, among them Cloudflare finance chief Thomas Seifert and Redwood Materials founder JB Straubel. 10 The money funds existing projects, including a large site in Abilene, Texas used by OpenAI, and modular data centers that Crusoe calls Spark and ships by truck. 10 The company's reasoning for the modular design includes local resistance to large complexes, a fight now visible across Silicon Valley data center projects. 15
The Nebius price list is the item that reaches a customer. GPU rates rise between 17% and 21% from October 1, some CPU-only instances by 25%, and memory offerings by about 41%, with discounts for customers reserving clusters for several months. 11 The company signed four customer contracts averaging more than $1 billion each in its previous quarter. 11 A funding round and a debt sale say what investors expect demand to do. The price list says what demand is doing, and it is the one a buyer has to act on.
In brief
- Amazon gave its first public position on AI safety, telling Reuters that models should be released when they are ready and safe to use, which it says comes from rigorous testing and strong safeguards, and that progress and safety can be pursued together. 16 The statement arrived after Nvidia and Meta argued on September 15 that each company should assess and control the risks of its own AI. 16
- The U.S. Justice Department sees room for rivals to coordinate on safety. Stanley Woodward, the associate attorney general who leads the Antitrust Division, said coordinating on cybersecurity or security counts as legitimate to him, and that the division is reviewing whether to update guidance that already permits coordination on cybersecurity. No frontier lab has asked it for a meeting. 17 The European Commission's Teresa Ribera, speaking at the same event, called shared risk a case where cooperation serves everyone's interest. 17
- King Charles told technology leaders at a meeting in Cumnock, Scotland, that AI carries existential dangers, and urged them to protect humanity. 18
What to watch
- September 24. President Trump and President Xi meet in Washington, with Huawei's accelerated chip timeline and China's export-control environment on the table.
- Anthropic's evaluators. The company says it will embed independent third-party evaluators with access comparable to internal risk assessment teams. No names, terms or publication rights have been given yet. 3
- Astra for Law's general availability. The configuration reaches selected firms first and the API after that, which is when an outside party can test it on documents its benchmark does not contain. 1
- The UN platform's accuracy. The 21.2% figure comes from a working paper awaiting peer review, and the platform's value depends on whether grounded data changes what models answer. 9
- Crusoe's Spark factories. Modular data centers are a bet that compute can be sited where a community will accept it. The first shipped units will show whether that holds.
References
- 1Introducing Astra for Law
openai.com
- 2
- 3
- 4
- 5
- 6
- 7Introducing the Life Sciences Verification Program
anthropic.com
- 8
- 9
- 10
- 11
- 12
- 13Making global data easier to explore
blog.google
- 14
- 15
- 16
- 17
- 18
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
More from this channel›
- Four AI labs are sued over the slowdown, Trump promises an AI czar, and Ohio puts data centers on the ballot
- Gemini hacked three companies, Accenture moved inside Anthropic, and OpenAI's $278 billion burn
- DeepMind's AGI institute, Vera Rubin's MLPerf debut, and OpenAI's 166-page proof
- Z.AI's $5B raise, Anthropic's Nasdaq choice, Trump's slowdown pushback, and Astra at Perplexity
- Amodei's call to pace the frontier, OpenAI's 2026 IPO delay, Meta's AI reorg, and Hyundai's Nvidia bet
- Sakana's Fugu models, Nvidia's Anthropic IPO talks, Senate AI safety rules, and OpenAI's Habitat
