system: OPERATIONAL
← back to all hacks
GOVERNANCE LOW NEW

Summer 2026 AI Safety Index: nine labs graded on security and risk

The Future of Life Institute's July 2026 index grades nine frontier labs on 37 safety and security indicators. No company beats C+, and safety frameworks still lack teeth.

2026-07-21 // 7 min affects: anthropic-claude, openai-gpt, google-gemini, meta-llama, xai-grok, mistral, deepseek, alibaba-qwen, zhipu-glm

What is this?

On 7 July 2026, the Future of Life Institute (FLI) published the AI Safety Index — Summer 2026 Edition, the fourth in its series grading frontier AI companies on their safety and security practices. It is not a vulnerability report. It is a scorecard: an independent panel of seven researchers and governance experts assigns letter grades to nine leading labs — Anthropic, OpenAI, Google DeepMind, Meta, xAI, Mistral, DeepSeek, Alibaba Cloud and Z.ai — across 37 indicators grouped into six domains.

The headline is that nobody does well. The full report PDF, dated 14 July, puts Anthropic at the top with a C+ (2.66 on a 4.0 GPA scale), OpenAI at C (2.28) and Google DeepMind at C (2.01). Meta lands at D+, the two remaining Chinese labs with published grades (Z.ai and Alibaba Cloud) at D-, and xAI, DeepSeek and Mistral all receive failing grades. Evidence was collected up to 3 June 2026, so the grades reflect the state of play in mid-2026, not a live snapshot.

How the index scores labs

The six domains are Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing. For a security reader, the interesting indicators are concrete rather than aspirational: bug bounties for system vulnerabilities, pre-deployment external safety testing, independent review of safety evaluations, dangerous-capability evaluations with elicitation, protection of safeguards against fine-tuning, major-incident response, user privacy, and serious-incident reporting with government notification. These are the same controls a mature vendor-security program would ask about.

The scoring method matters if you want to cite it. The panel reviewed publicly available materials — model cards, research papers, benchmark results — supplemented by a targeted company survey aimed at transparency gaps such as whistleblower protection and external evaluation. Reviewers assigned domain-level grades against absolute standards, wrote justifications, and their scores were averaged; individual grades stay confidential. This is expert judgment, not an automated benchmark, so the value is in the reasoning and the recommendations rather than in a single number.

Anthropic leads five of six domains, largely on transparency (it publishes both model specs and system prompts) and a comparatively developed safety framework; OpenAI now leads Risk Assessment on the strength of broader external testing. The weakest domain across the whole industry is Existential Safety, where no company exceeds a C- and most sit at D or below.

Why it matters

Two findings should concern anyone who relies on these systems. First, the panel judged that safety frameworks have weak teeth: even the labs that published or updated frameworks as EU and US deadlines approached often lack quantitative thresholds, genuinely independent audits, and a clear internal authority that can actually halt a deployment. Second, several leaders have walked back earlier pledges to pause if capability redlines are approached, a pattern reviewers called “moving goalposts” that “undermined safety frameworks across the board.”

For defenders, the index is useful precisely because it is vendor-neutral and reasoned. It gives security and procurement teams an external reference to push back on marketing claims, and it names the gap between safety rhetoric and revealed behavior — the report notes that reassuring public messaging at several labs diverges from commercial and legislative conduct. As one panelist, Stuart Russell, put it, companies are increasingly “planning to release [systems] even if it’s demonstrably unsafe to do so.”

Defenses

The practical use of a report like this is as an input to vendor risk assessment. The through-line is: trust the evidence, not the pledge.

  • Map the domains to your own controls. Treat Risk Assessment and Information Sharing as proxies for a provider’s red-teaming maturity and incident transparency, and weight them for your use case.
  • Ask for what the index says is missing. Request quantitative capability thresholds, named independent auditors, and the internal body with authority to stop a launch — not just a framework document.
  • Check safeguard durability. Confirm how a provider protects safety guardrails against fine-tuning and API misuse, one of the index’s explicit indicators.
  • Verify incident and disclosure paths. Require a serious-incident reporting commitment and a working bug-bounty or vulnerability-disclosure channel before integrating a model into sensitive workflows.
  • Re-check each edition. Grades move between editions (Meta rose, xAI fell); use the trend, not a single snapshot, and revisit at renewal.

Status

AspectDetail
PublisherFuture of Life Institute (FLI)
Published7 July 2026 (full report PDF dated 14 July)
Scope9 labs, 37 indicators, 6 domains
Evidence cutoff3 June 2026
Top gradeAnthropic — C+ (2.66 / 4.0)
Failing gradesxAI, DeepSeek, Mistral
TypeIndependent safety assessment — not a vulnerability advisory

The takeaway is not a new attack but a measurement problem made explicit: the industry’s own safety commitments are, by the panel’s reading, an unreliable proxy for security in practice. The index gives buyers and defenders a structured, sourced way to interrogate that gap before they depend on a model.

Sources