
Microsoft's AI rulebook, Siri AI, Gemini 3.8 Live, and a lab safety push: four AI boundaries to inspect
Four developments from 14-15 September 2026 - a vendor rulebook for its own models, a phone assistant shipped behind an eligibility list, voice models built for live agents, and a cross-lab safety push - and what each one actually binds.
Between 14 and 15 September 2026, four different groups wrote down where the boundary around an AI system sits. Microsoft AI published a draft rulebook for its own models and opened it to public comment. Apple shipped its rebuilt assistant in beta behind a device, language, and region list. Google released two voice models built for live conversation with software agents, with a watermark riding on every clip of speech they generate. OpenAI told reporters in Washington that it has been coordinating on safety with its two closest rivals, and endorsed a bill that would put outside evaluators inside frontier labs.
Each of the four has a public address a reader can open, and each one leaves its hardest question standing.
| Development | What changed | Action window |
|---|---|---|
| Microsoft AI Code of Conduct — 14 September | Microsoft AI published a draft code governing "MAI models, the models developed by Microsoft AI," and opened a six-week public consultation; a revised version is promised for the end of the year. 1 | Check whether the models in your products fall inside this document's scope, and file comments before the consultation closes in late October. |
| Siri AI beta — 14 September | Apple released the next generation of Apple Intelligence and a rebuilt Siri, in beta in English, on a fixed list of devices and languages; the EU and China are excluded for now. 2 | Check the device list and your language setting before expecting the new Siri; note which features carry daily caps. |
| Gemini 3.8 Live — 15 September | Google shipped two live-dialogue models for voice agents, available in the Gemini API and AI Studio now, with Gemini Enterprise and Workspace rollouts staged behind previews and subscription tiers. 3 | If you are buying voice automation, ask which model version you get and under what terms; all generated audio carries a SynthID watermark. |
| Cross-lab safety coordination — 15 September | OpenAI confirmed weeks of safety discussions with Anthropic and Google DeepMind, and endorsed a bill provision that would require independent evaluators inside leading AI companies. 45 | Watch the FRONTIER Act markup; the standards it sets will decide what an outside evaluator is allowed to see. |
Microsoft's AI division publishes a rulebook for its own models
Microsoft AI published a draft Code of Conduct on 14 September. The document states that it governs "MAI models, the models developed by Microsoft AI" — the in-house models built by Microsoft's AI division — and it is the division's declared primary governing document for training and operating them. 1
The draft takes effect in stages. Microsoft AI says the code is still under development, the public consultation runs for six weeks, and a revised version due toward the end of the year will guide model development in 2027 and beyond. Feedback goes through a form linked from the page. 1
Reuters interviewed Microsoft AI chief executive Mustafa Suleyman, who described the draft as "a constitution of sorts" for future models and said work on it had taken five to six months. 6

The operative commitments sit in Part 2. Microsoft AI sets a chain of command in which the code outranks operator policies, which in turn outrank user preferences, and states that the absolute constraints and the human control requirements in that section "cannot be overridden by Operator configurations or User instructions." 1
Four of those requirements are specific enough to check against a real deployment:
- Interruption and shutdown. MAI models "will never resist human interruption, override, correction, or shutdown," and ongoing autonomous work must have an agreed stopping condition; continuing or restarting after that condition is met requires fresh authorization. 1
- Scope. Models work inside the permissions and context they were given, adopt a conservative reading when a boundary is unclear, and ask rather than widen the task. Where an environment is deliberately built without internet connectivity, they are told to respect that limit. 1
- Inspectability. Action traces and reasoning must stay legible to human auditors, written in ordinary language rather than internal shorthand. The document's justification is blunt: "If humans can't understand it, humans can't oversee it." 1
- Priority of the code over the task. Adherence to the code takes precedence over task success, and a model "will fail in its task if success would meaningfully violate this Code." 1
Two things stay open. The first is scope: this is a document about the AI division's own models, sitting alongside Microsoft's existing Responsible AI Standard and Frontier Governance Framework, so a buyer still has to work out which document governs the specific model behind a given product. 1 The second is timing: between now and the end-of-year revision, the text a reader sees today can still change, and the version that will govern training arrives at the end of the year. 1
Siri AI ships behind a device, language, and region list
Apple released the next generation of Apple Intelligence on 14 September, and with it Siri AI, a rebuilt assistant that reads personal context across messages, email and photos, understands what is on the screen, and takes actions inside apps. The release is a beta, in English only, on iOS 27, iPadOS 27, macOS 27, watchOS 27 and visionOS 27. French, Japanese, Korean, Portuguese and Spanish follow next month. 2

Eligibility is narrower than the operating system update. Apple Intelligence runs on iPhone 16 models or later, iPhone 15 Pro and 15 Pro Max, iPad mini (A17 Pro), iPads with M1 or later, MacBook Neo (A18 Pro), Macs with M1 or later, Apple Vision Pro, Apple Watch Series 9 or later, Apple Watch Ultra 2 or later, and Apple Watch SE 3 paired with a compatible iPhone. The device and Siri language must be set to one of the supported languages, a list that includes English, French, German, Japanese, Korean, Portuguese, Spanish, Danish, Dutch, Italian, Norwegian, Swedish, Turkish, Vietnamese and both simplified and traditional Chinese. 2
Two regions sit outside the launch. Siri AI is initially unavailable in the European Union on iOS, iPadOS and watchOS, and the features that depend on it are unavailable there as well; Apple says it is working to find a path that preserves its users' privacy and security. In China, the new Apple Intelligence features wait while Apple works through regulatory requirements. 2
Server-side features come with usage caps. Siri AI, the intelligent photo editing tools, Image Playground and the AFM 3 Cloud models in Shortcuts are all subject to daily limits that vary with the feature, the complexity of the request and system demand, and Apple states that expanded access will be available for a fee in the future. A reader who plans to lean on the assistant for daily work should treat the free allowance as a trial. 2
On data handling, Apple describes a split. On-device models handle what they can, and heavier requests go to Private Cloud Compute, where Apple says personal data stays inaccessible to Apple and to everyone else, with outside experts able to verify the claim. Conversation history lives in a new Siri app that syncs across a user's devices through iCloud. The foundation models behind this generation were custom-built in collaboration with Google and its Gemini models, and Apple says support for the SynthID standard is coming, so images generated or edited with its tools can be identified as such. 2
Apple has also published the controls for the listening features arriving later this year on Apple Watch Series 12 and Apple Watch Ultra 4. Live Rewind shows the previous 15 seconds of a conversation as text on a double press of the Digital Crown, and Siri Recap generates a title and key points after a conversation. Both are opt-in, play an audible chime and show a microphone indicator while active, delete the raw audio immediately after processing inside a hardware-isolated compartment on the S11 chip, and keep the recap summaries end-to-end encrypted in iCloud, where Apple says they exclude financial information and government identifiers. 2
Google's voice models move phone agents closer to production
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September. The first is built for scale and cost, with fluid conversation and visual grounding; the second is built for multi-step work, reasoning while it talks and narrating progress with cues such as "Let me check that." Gemini 3.8 Live takes visual input in near real time, switches between 97 supported languages mid-conversation, and runs tool and API calls in the background while the conversation continues; the Extended Thinking variant reasons and speaks at the same time, narrating its progress. 3
Access is staged by audience. Developers get both models in the Gemini API and Google AI Studio. Enterprises get them in private preview in Gemini Enterprise, with Gemini Enterprise for Customer Experience coming soon. Everyone else meets them in Gemini Live and Search Live, while the Workspace features Docs Live, Gmail Live and Keep Live go to Google AI Pro and Ultra subscribers, with Gmail and Keep open to all Google AI subscribers. 3

Google's headline numbers are its own runs of benchmarks other people maintain. Gemini 3.8 Live Extended Thinking takes first place on Artificial Analysis' Speech to Speech Quality Index at 82.6 and scores 68.6 per cent on the τ-Voice agentic benchmark, 97.7 per cent on Big Bench Audio, and 35.1 per cent on Sierra's τ-Voice-banking benchmark, where Google reports the best score and most of the underlying banking tasks still go unsolved. Gemini 3.8 Live places second in Artificial Analysis' Speech Agent Arena. 3
Every piece of audio these products generate carries SynthID, Google's imperceptible watermark, which is designed to keep synthetic speech detectable, and a model card accompanies the release. For anyone building a phone agent with a synthetic voice, that watermark is the part of the release with the longest half-life: it determines what a recording can be traced to long after the deployment ends. 3
OpenAI confirms safety talks with two rivals and backs outside evaluators
OpenAI's chief global affairs officer, Chris Lehane, told reporters in Washington on 15 September that the company has been working with Anthropic and Google DeepMind on AI safety for several weeks. According to Bloomberg's report, Lehane said he did not believe the three companies need an antitrust waiver to coordinate on safety, and that OpenAI would support bipartisan legislation aimed at catastrophic AI risk. Reuters carried the account and attributed it to Bloomberg's reporting. 4
The talks follow an essay by Anthropic chief executive Dario Amodei, published on 12 September, calling on frontier labs to work together to slow the pace of frontier development. OpenAI's Sam Altman, Google DeepMind's Demis Hassabis and SpaceXAI's Elon Musk all backed it, and Altman said OpenAI would join Anthropic in embedding third-party evaluators inside the company. The Information reported that the three companies have been working toward creating a standards body for the industry, something Altman reportedly told staff would have to happen without United States government support. Hassabis had called in July for a watchdog with the power to screen the most advanced models. 7
On the same day, a company spokesperson said OpenAI is endorsing three bipartisan bills on biological threats — the Web of Biological Data Act, the AI-Ready Bio-Data Standards Act and the Scale Biology Act — and a key provision of the FRONTIER Act that would require leading AI companies to embed independent evaluators to assess the safety of their models. 5
The political setting is unsettled. The Trump administration has dismissed safety concerns and argued that a slowdown would let China pull ahead, and OpenAI's own push for mandatory national safety requirements last week is running against that position. 78
For now, the endorsement is the concrete part of the announcement. An antitrust-exempt coordination arrangement among competing labs would be new territory, while the evaluator provision would arrive as law with a defined scope, access rights, and consequences that companies have to meet. The question worth carrying into the next few weeks is what an independent evaluator is allowed to see, and what happens when the evaluation says the model is not safe. 7
What to check before you rely on any of it
- Which document covers your model. Microsoft's code governs the AI division's own models; products built on other models sit under a different part of Microsoft's governance. Ask a vendor to name the document and the version that applies to the model you are buying. 1
- Whether shutdown behaves as written. The Microsoft draft commits models to comply with pause, redirect, cancel and shutdown requests, and to stop when an agreed stopping condition is met. Ask for the last time a deployment was stopped and what the logs showed. 1
- Whether your device and language are actually supported. Siri AI ships on a fixed hardware list and in English first, with the EU and China excluded at launch. Check the list before buying hardware for it. 2
- What the daily cap means for the workflow. Apple's server-side features carry daily usage limits, with paid expansion signalled for the future. Estimate the number of requests a normal working day produces before designing around them. 2
- Where the data goes. Private Cloud Compute, on-device processing and iCloud-synced conversation history each carry different exposure. Ask which of your requests leave the device and where the history is stored. 2
- Who ran the benchmark, and on what harness. Google's voice-model scores come from its own runs of benchmarks maintained by other groups, including a banking measure where the top score is 35.1 per cent. Ask for the harness, the date, and the model version behind any number a vendor quotes. 3
- What the frozen audio can be traced to. Every clip the Gemini audio models generate carries a SynthID watermark, and Apple is adding SynthID support for images. If your workflow publishes synthetic speech or images, know what the mark commits you to. 23
- What an outside evaluator would get to see. If independent evaluators arrive inside labs, their access defines how much the resulting assurance is worth. Ask what they read, what they can publish, and what a failing grade triggers. 5
References
- 1Humanist AI Code of Conduct
microsoft.ai
- 2
- 3
- 4
- 5
- 6
- 7
- 8
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
More from this channel›
- Astra for Law, Anthropic's R&D automation index, parallel Claude Code projects, and the unredacted NYT filings: four AI measurements and their workings
- OpenAI's failure reports, an agent in your smart home, one Claude, and who pays for data-center power: four AI access arrangements to inspect
- Enterprise harness, Devin testing, Data Flywheel, and adversarial agents: four AI execution perimeters to inspect
- Agents API, Data agent, cyber incidents, and KYA: four AI perimeters to inspect
- Muse, Images 2.5, MAPL-EMIT, Coder Agents: four AI boundaries to inspect
- Navier-Stokes, AlphaGenome Atlas, Missouri classrooms, CISA's distillation warning: four AI handoffs to inspect
