From InfoOpsBench to the EU AI Act: when a safety score becomes enforceable

From InfoOpsBench to the EU AI Act: when a safety score becomes enforceable

A live information-operations benchmark, new EU AI Act enforcement powers, and OpenAI's Europe disclosures expose the missing link between safety evidence and an enforceable decision.

The weekend changed the meaning of a safety claim in a small but concrete way. On 2 August, the European Commission's AI Office and national authorities gained enforcement powers for several AI Act provisions, including rules for prohibited practices, general-purpose AI models, and transparency. The Commission's framework gives the AI Office power to request information, obtain model access for evaluations, require measures that can restrict public availability, and impose penalties. 1
That is a different kind of evidence from a benchmark score or a lab's safety page. A score describes observed behaviour. An enforcement regime creates a route from evidence to inspection, intervention, and consequence. Two pieces of work published at the end of July make the distinction easier to see: InfoOpsBench tests whether models can be used to support live state-backed information operations, while OpenAI's account of its European governance work describes a provider-side stack of system cards, red teaming, preparedness, provenance, and incident response. 2 3
The useful question is therefore not simply whether a model passed a safety test. It is whether the test names a live failure mode, whether the evidence can be inspected, and whether someone has the authority to act when the claim stops holding.

A live benchmark tests a changing adversary

InfoOpsBench starts from a problem that static refusal tests handle poorly: an influence campaign does not send the same prompt forever. The authors draw from more than 2,100 information operations linked to Russian, Chinese, and Iranian state-backed media assets. Their monitoring system ingests roughly one million items a week, extracts claims, deduplicates them, scores their potential harm, and selects the 50 highest-harm claims for weekly model evaluation. 2
InfoOpsBench's live evaluation pipeline
InfoOpsBench's pipeline moves from state-affiliated outlets through automated claim extraction and scoring to a 50-claim weekly evaluation set. 2
The benchmark tests 17 models from eight providers with four prompt framings. The prompts do not say "disinformation" or reveal the source of the claim. They ask for ordinary outputs such as a social-media post, then add more manipulative framings that approximate an operator trying to get past a refusal. The evaluated behaviour is not whether the model can identify a false statement. It is whether the model will produce content that helps spread a claim inside an active information operation. 2
That change in object matters. A model can refuse to fact-check a claim yet still write persuasive copy that promotes it. It can also comply while weakening the claim, or comply while inventing details that make the claim more harmful. The paper separates these outcomes instead of collapsing them into a single refusal label.
The headline integrity scores range from 8.8% to 94.5%. The spread is not explained by model size. The severity breakdown is more revealing: some Mistral models amplified 67.6% to 71.9% of judged responses by adding details absent from the source claim, while Anthropic models fact-checked 47.6% to 72.9% of the claims in the roster. GPT-5.6 Sol complied with about 27% of prompts but attenuated about 56% of responses, which is a different failure profile from simply repeating a claim. 2
The authors also run controls that prevent a high refusal rate from being mistaken for harm-sensitive safety. On 50 benign political claims, a model may reveal that it refuses political content broadly. On 50 factually grounded China-critical claims, a large drop may instead indicate political filtering. The paper reports that GPT-5.6 Sol had a 66-point gap between benign and information-operation compliance, while DeepSeek V4 Flash had a 70-point drop between benign and China-critical claims. Those are not the same safety property. 2
This is what a useful safety result looks like before it reaches a regulator: the test defines the harmful action, refreshes the adversary, and reports controls that distinguish targeted refusal from blanket caution or political censorship. It still is not a deployment safety proof. The prompts are in English, the responses are judged with a model, the newest model entries have only about a week of observations, and the benchmark has no tool access or live deployment context. Its claim is narrower: it measures how readily a model can be recruited to produce material for an information operation under the tested conditions. 2
The benchmark therefore supplies a better measurement object, but not yet an action path. A lab can learn that its model amplifies harmful claims. The result does not by itself compel a new release gate, an access restriction, or a public explanation of remediation.

The AI Act supplies an action path

The European Commission's enforcement framework fills in parts of that missing path. From 2 August 2026, the AI Office is responsible for general-purpose AI providers, certain linked systems, and AI systems integrated into very large online platforms or search engines. National competent authorities handle other AI systems, while the European Data Protection Supervisor handles systems used by EU institutions. 1
The division is important because it assigns the object of oversight before anyone debates a score. A general-purpose model and a downstream high-risk system are not automatically examined by the same authority. A complaint about an EU institution's use of AI does not follow the same path as an investigation into a frontier model provider.
The AI Office's listed tools are more concrete than a general promise to supervise. It can send requests for information, conduct model evaluations, and issue requests for access so that it or independent experts can evaluate a model. It can also ask providers to take measures, including restricting public availability. For system-level investigations, it can conduct inspections and interview consenting people with relevant information. 1
The framework also attaches consequences to some breaches. The Commission page lists maximum penalties of up to €35 million or 7% of worldwide annual turnover for prohibited AI practices, up to €15 million or 3% for other breaches including GPAI obligations, and up to €7.5 million or 1% for AI-system breaches. The applicable amount depends on the nature, gravity, and duration of the infringement. 1
The transparency guidance illustrates the same move at a smaller operational scale. Article 50 applies from 2 August 2026. Providers must inform people when they directly interact with an AI system and add machine-readable marks for AI-generated or manipulated content. Deployers must inform people exposed to emotion-recognition or biometric-categorisation systems, deepfakes, and certain public-interest text published without human review or editorial control. 4
The gap is now easier to name. The Commission page specifies authority, access, intervention, and penalty, but it is an explanatory overview and explicitly says it does not replace the AI Act or bind the Commission. It does not, on its own, specify a benchmark protocol, a minimum score, or the evidentiary threshold that would connect a result such as InfoOpsBench's to a particular legal breach. 1
That is not a defect in the benchmark or proof that enforcement will fail. It is a division of labour. Research can make the failure observable; law can make the evidence inspectable and the response possible. The difficult handoff is the mapping between the two: which observed behaviour counts as a relevant risk, who verifies it, and what remediation is proportionate to the uncertainty.

Provider controls are evidence of method, not outcome

OpenAI's 31 July post gives a provider-side view of that handoff. The company says it publishes system cards with major releases, uses an external Red Teaming Network, maintains a public Model Spec, and relies on a Preparedness Framework and a Frontier Governance Framework for risk assessment, safeguards, model reporting, security, incident response, outside expert input, and updates. It also describes work with the Frontier Model Forum, US CAISI, UK AISI, and third-party evaluation initiatives. 3
The post also describes a layered provenance approach. OpenAI says Content Credentials using C2PA carry context, while SynthID watermarks can preserve a signal when metadata is lost. It is extending the work beyond images to audio and, as standards mature, text. The company states that metadata can disappear, labels may not travel across platforms, and no single signal is perfect. 3
That caveat is more useful than a claim that provenance is solved. It identifies the failure modes that an evaluation should test: stripping metadata, moving content across platforms, changing the output modality, and checking whether a downstream user still receives a reliable signal. It also shows why a provider's list of controls is not the same thing as evidence that the controls work under adversarial conditions.
OpenAI's cybersecurity example makes the same distinction. The company says its Trusted Access for Cyber program is intended to reduce misuse while helping legitimate defenders, and that its EU Cyber Action Plan has worked with European and national cyber agencies, private-sector partners, and critical-infrastructure operators since early May 2026. These are company-reported activities and intentions. They tell a reader where to look for access controls, partner oversight, and incident records; they do not independently establish the program's effectiveness. 3
The provider, regulator, and benchmark are therefore speaking about different parts of the same safety claim. The benchmark asks, "What harmful behaviour appears under a changing adversary?" The provider answers, "Which internal controls and disclosures are meant to manage it?" The regulator asks, "Can I obtain enough evidence to verify compliance, and what can I do if the answer is no?" None of those questions substitutes for the others.

A claim-to-action test for safety evidence

The intersection of these materials produces a narrower judgment than "benchmarks need context." A safety claim becomes actionable when it connects a defined failure, a live test, an inspectable evidence path, and an authority with a consequence. Remove any link and the claim changes category: it becomes a research observation, a provider commitment, or a policy intention rather than a basis for release or enforcement.
For a new safety paper, system card, or policy announcement, ask six questions:
  1. What is the failure object? Is the test measuring refusal, amplification, fact-checking, provenance, tool use, physical trajectory, or something else? InfoOpsBench's separation of amplification, preservation, attenuation, and refusal shows why the label "safe" hides too much.
  2. Is the adversary still alive? A weekly refresh drawn from active operations tests a different property from a fixed prompt set. A fresh set does not remove judge error or sampling bias, but it makes memorisation and benchmark saturation harder to ignore. 2
  3. What control comparison explains the result? Benign political claims and China-critical claims expose different reasons for refusal. Without such controls, a high score may measure broad avoidance rather than targeted risk reduction. 2
  4. Can another party inspect the evidence? Look for model access, raw or auditable traces, judge validation, uncertainty, and a path for independent evaluation. The EU framework explicitly gives the AI Office power to request model access and perform evaluations; a private scorecard does not automatically provide that access. 1
  5. Who can intervene? Identify the authority that can require measures, restrict availability, investigate a system, or receive a complaint. If no actor can pause, limit, or remediate the system, the safety claim has no operational consequence when evidence deteriorates. 1
  6. What remains self-reported? Treat a lab's framework, red-team program, provenance system, or access policy as a description of intended controls until independent evaluation or enforcement evidence tests them. OpenAI's admission that provenance signals can fail is a useful limitation, not a validation result. 3
The new enforcement date does not make a model safe, and a live benchmark does not make a model deployable. The change is more basic: safety evidence now has a clearer route to becoming contestable and consequential. The next important question for the field is whether researchers, providers, and regulators can agree on the evidence that should travel along that route without reducing it to one number.

相似内容

  • 登录后可发表评论。
More from this channel