
Aug. 29 AI brief: Anthropic's TASTE benchmark, Gemini Notebook limits, and OpenAI's Thailand accelerator
A concise scan of Anthropic's new safety-research benchmark, Google's five-hour Gemini Notebook limits, and OpenAI's eight-week accelerator for 10 Thai startups.
Coverage window: Aug. 28 through the morning of Aug. 29, 2026. Three developments point to the work around AI becoming more operational: a benchmark asks whether models can judge safety research, Google changes how users spend Gemini Notebook compute, and OpenAI starts a government-backed accelerator for Thai startups.
| Development | What changed | Scale or status | Why it matters |
|---|---|---|---|
| Anthropic's TASTE benchmark 1 | Anthropic published a benchmark for model judgments on pairs of AI safety research proposals. | 92 proposal pairs; estimated expert agreement is 77%; the top tested model reached 60%. 1 | Safety research has many questions without an automatic right answer, so proposal selection itself needs a measurable test. |
| Gemini Notebook usage limits 2 | Google is replacing a single daily limit with compute-specific limits that refresh every five hours. | Rollout begins Sept. 2 for consumer accounts on web and mobile; heavy outputs can be deferred. 2 | Users can spend more of their allowance on complex tasks instead of exhausting it through a fixed daily bucket. |
| OpenAI x MHESI AI Accelerator 3 | OpenAI and Thailand's higher-education ministry launched an eight-week accelerator for local startups. | Ten teams split between medical/wellness AI and education; Demo Day is planned for Bangkok in November. 3 | The program tests whether model access, technical help, and local institutions can turn prototypes into products with evidence from real users. |
Anthropic asks whether models can choose good safety research
Anthropic published TASTE, short for The AI Safety Taste Evaluation, on Aug. 28. The benchmark presents models with two AI safety research proposals and asks which proposal is better. Experienced researchers created the preference labels through individual scoring, paired discussion, and revised scoring. The final benchmark contains 92 pairwise comparisons, with estimated human agreement of 77%. 1
The test targets a part of research that ordinary coding or question-answering benchmarks miss. A proposal about a future AI risk may have no quick, verifiable answer. A model can still help researchers by ranking directions, but only if the ranking tracks the judgments of people who understand the field.

Fable 5 led the tested models at 60% overall accuracy. On 74 pairs drawn from different prompts, Fable 5 reached 69%. Anthropic reports that almost every model fell within two standard deviations of chance, while individual model intervals were roughly plus or minus 10 percentage points. Those intervals and the 92-pair sample make the current result a starting point for measurement rather than a stable league table. 1
The next useful result is a larger, more varied evaluation with clearer agreement between independent human raters. Until then, TASTE gives safety teams a concrete question to ask before assigning models a role in research planning: can the model's preferences track expert judgment on the research direction itself?
Gemini Notebook turns usage into a five-hour budget
Google said on Aug. 28 that Gemini Notebook will move to flexible, compute-specific usage limits. The limit will account for prompt complexity, chat length, the number of sources, and the features used. The allowance will refresh every five hours, replacing the current daily reset. Google plans to begin the rollout on Sept. 2 for consumer accounts on the web and mobile apps. 2
The change matters most when one notebook mixes small questions with expensive outputs. Users can keep using the existing features, track their usage inside the notebook, and receive an alternative suggestion when a request would exceed the available compute. Google also says outputs such as Video Overviews and Slide Decks can be deferred and generated later, with notifications when they are ready. 2
The practical checkpoint is simple: whether the five-hour cycle makes long research sessions easier to plan, or merely moves the point at which users wait. Google has described the mechanics of the limit, while the source announcement leaves the allowance for each account unspecified.
OpenAI backs ten Thai startups for eight weeks
OpenAI and Thailand's Ministry of Higher Education, Science, Research and Innovation announced an eight-week accelerator on Aug. 28. The first cohort has 10 startups: five in medical and wellness AI and five in education. Each team receives $2,000 in API credits, technical guidance, access to OpenAI's frontier models, and a dedicated mentor. Weekly sessions cover product design, engineering, testing, evaluation, responsible AI, privacy, security, cost management, growth, and fundraising. 3
The ministry connects the cohort with Thailand's research and talent networks. The National Innovation Agency connects teams with funding and growth channels, while Mahidol University contributes academic expertise and mentors. Each startup sets a product, pilot, evaluation, or commercial milestone for the program. A Demo Day in Bangkok is planned for November, where the teams are expected to show user evidence, product evaluations, and deployment plans. 3
The program's useful test is the evidence after the credits are spent. Watch which teams move from a working prototype to a measured pilot, and whether the local university, government, and investor network keeps supporting the projects after Demo Day.
What to watch
- TASTE: A larger benchmark and independent human ratings that narrow the roughly 10-point model uncertainty. 1
- Gemini Notebook: The Sept. 2 rollout, account-level allowances, and whether deferred Video Overviews and Slide Decks reduce interruptions in long research sessions. 2
- Thailand accelerator: The November Demo Day, measured pilot results, and follow-on funding or deployments for the 10 startups. 3
References
- 1TASTE: Can AI Models Judge AI Safety Research Proposals?
alignment.anthropic.com
- 2More compute flexibility in Gemini Notebook
blog.google
- 3
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
More from this channel›
- Claude Fable 5.1, OpenAI Astra, and CrowdStrike's cyber models
- ChatGPT Ads hits $1B, Google opens AI Search controls, and a new test for visual hallucinations
- DALL-E GPT, Gemini Robotics ER 1.6, GitHub Spark, and the AI power bottleneck
- Aug. 30 AI brief: Sony Music and Warner sue Anthropic, Nvidia moves beyond the GPU
- Aug. 28 AI brief: Anthropic's hardware standard, a cyber-defense call, and a Pentagon blacklist blocked
- Aug. 27 AI brief: Hugging Face incident report, Z.ai Ox Alpha, Gemini 3.5 Transcribe, and compute deals
- Aug. 26 AI brief: Apple's on-device Macs, Jalapeño's first numbers, and a $5M wellbeing fund
- Aug. 25 AI brief: Thomson's domain model, GPT-5.6 in Kiro, and Wan3.0
