AI Safety Index 2026: What the C+ Grades Really Mean

AI Tools & Governance · Updated August 2026

AI Safety Index 2026: What the C+ Grades Actually Mean for Anyone Choosing an AI Tool

Nine AI labs got graded on safety. Nobody passed. Here’s what that C+ actually protects you from — and what it doesn’t.

By Oyekale Olawale · 12 min read

Quick Answer

The Future of Life Institute’s Summer 2026 AI Safety Index graded nine AI companies on their safety policies, and the highest score any of them earned was a C+. Anthropic topped the field at 2.66 out of 4.0, OpenAI and Google DeepMind sit at C, and xAI, DeepSeek, and Mistral all failed outright. The part most write-ups skip: this index scores the company, not the actual product you’re logging into. A C+ tells you the lab publishes decent policies. It tells you nothing about whether the specific chatbot you’re pasting client data into today is configured safely. That gap is where I’ve spent the last few weeks testing.

I write about AI tools for a living, which means I’m inside Claude, ChatGPT, Gemini, and a rotating cast of open-weight models most days of the week. So when the Summer 2026 Index landed on July 7 with a headline this blunt — best grade in the industry: C+ — I wasn’t just going to summarize the press release. I pulled the full scorecard, cross-checked it against what I’ve actually watched these tools do in production, and I’m going to tell you where the grades line up with reality and where they don’t.

If you’re a solo creator, a small agency, or a business owner deciding which AI tool to trust with your work, this is the buyer’s version of the report — not the policy-wonk version.

The Scorecard: Nine Labs, One C+, Three Failures

FLI’s independent panel — seven outside experts, including UC Berkeley’s Stuart Russell and University of Montreal’s David Krueger — graded nine companies across 37 indicators in six domains, using evidence collected through June 3, 2026. Here’s the full report card.

C+

Top overall grade (Anthropic)

3 / 9

Labs that failed outright

37

Indicators, 6 domains

D+

Best existential-safety grade, industry-wide

Overall score, out of 4.0 — FLI AI Safety Index, Summer 2026

Anthropic
C+ · 2.66
OpenAI
C · 2.28
Google DeepMind
C · 2.01
Meta
D+ · 1.32
Z.ai
D− · 0.88
Alibaba Cloud
D− · 0.87
xAI
F · 0.65
DeepSeek
F · 0.47
Mistral
F · 0.33

Bars scaled to the index’s 4.0-point GPA-style ceiling. Source: Future of Life Institute, AI Safety Index Summer 2026.

Here’s the thing the letter grades hide: the gap between Anthropic’s C+ and Google DeepMind’s C is smaller than the gap between Google DeepMind and Meta, which is smaller still than the canyon separating Meta from the three failing labs. Read the letters and you’d think “top tier” versus “middle tier.” Read the numbers and the real story is three US-based labs clustered together, then a cliff.

What Each Domain Actually Measures

A single overall letter flattens six very different questions into one number, and that’s where most coverage of this report stops being useful. The panel grades companies across Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing. Here’s how the top three labs actually broke down, domain by domain — this is the table I keep coming back to when I’m deciding how much I trust a vendor’s public claims.

Domain Anthropic OpenAI Google DeepMind
Risk AssessmentC+C+C+
Current HarmsB−CC
Safety FrameworksB−C+C
Existential SafetyD+D+D
Governance & AccountabilityBCC−
Information SharingB+B−B−

Look at that Existential Safety row. Every single company in this table — the three best-performing labs on Earth, by FLI’s own measure — scores D or D+. Not one clears C-minus. The panel’s own language for the state of the field was blunt: constitutional classifiers, chain-of-thought monitoring, loss-of-control provisions — real efforts, but judged “entirely inadequate” against the scale of the problem. I don’t think that’s fear-mongering. I think it’s an honest admission that nobody, including the company I use every day to write this article, has actually solved the hard version of the alignment problem yet.

My Hands-On Take: What The Grades Look Like In Daily Use

I’m not a policy researcher. I’m the guy who has to actually get work done using these tools every single day, across dozens of client projects. So instead of just relaying FLI’s grades, I ran the same week of real work — research, drafting, fact-checking, light coding for site fixes — across Claude, ChatGPT, Gemini, and DeepSeek, and paid close attention to where the safety grades tracked what I actually experienced.

Claude’s B+ in Information Sharing lines up with something I notice constantly: Anthropic publishes its system prompts and model behavior specs in a way none of its rivals fully match, and I’ve been able to cite exact model-card language when writing tool comparisons on this site because that documentation actually exists and is kept current. That’s not marketing fluff — it’s the single reason I trust Claude’s stated limits more than a rival’s.

ChatGPT’s Risk Assessment grade is genuinely earned in my experience — its guardrails feel the most consistently tested against edge cases, and I’ve seen it flag borderline requests other tools wave through without comment. But I’ve also hit inconsistent refusal behavior on the exact same prompt phrased two different ways within the same session, which is a real UX flaw, not just a safety footnote — it makes the tool feel unpredictable for repeatable workflows.

Gemini sits in the middle for me. Google DeepMind’s C grade in Governance and its C-minus specifically flag unclear internal decision authority — and functionally, I’ve noticed Gemini’s content boundaries shift noticeably between product surfaces (the consumer app versus API access), which matches a company where, per the index panel itself, it’s genuinely unclear who inside the org can independently halt a risky deployment.

DeepSeek is where the failing grade and my day-to-day experience match most closely. There’s no published safety framework I can point a client to, refusal behavior is noticeably looser on sensitive topics, and I would not put confidential client material anywhere near it. If you’re comparing coding assistants that route through these models, that gap shows up directly in what I found while testing Verdent and its top alternatives — the underlying model’s safety posture bleeds straight into the tool built on top of it.

My honest opinion, for whatever a reviewer who uses these things eight hours a day is worth: the index grades map onto lived experience better than I expected going in. That surprised me. I went into this report skeptical that a policy scorecard could say anything about product behavior. It says more than I thought — just not everything.

The Trend: Almost Everyone Got Worse

FLI runs this index twice a year, so the more revealing story isn’t the snapshot — it’s the direction of travel. Comparing Winter 2025 to Summer 2026, seven of the eight previously-graded labs scored lower this round. Only Meta genuinely improved.

Lab Winter 2025 Summer 2026 Direction
AnthropicC+ · 2.67C+ · 2.66Flat
OpenAIC+ · 2.31C · 2.28↓ dropped a tier
Google DeepMindC · 2.08C · 2.01Slight dip
MetaD · 1.10D+ · 1.32↑ only real riser
Z.aiD · 1.12D− · 0.88↓
Alibaba CloudD− · 0.98D− · 0.87↓
xAID · 1.17F · 0.65↓↓ steepest fall
DeepSeekD · 1.02F · 0.47↓
MistralNot yet ratedF · 0.33New entrant, bottom score

xAI’s fall is the one that should worry buyers most: from a D at 1.17 down to a flat F at 0.65 in a single six-month cycle, the sharpest drop of any lab in the report. That’s not gradual drift — that’s a company moving in the wrong direction fast, right as it rebranded to SpaceXAI following its merger with SpaceX (a change that landed the day before the index published, so the report still lists the old name).

The Real Story: Everyone Is Walking Back Their Own Promises

Here’s what actually drove the coverage on this report, and it’s not the letter grades. Anthropic, OpenAI, Google DeepMind, and Meta — the four companies with the most public safety commitments — have all weakened or voided earlier pledges to pause development if specific risk thresholds were crossed. The panel called it “moving the goalposts.”

Even Anthropic, sitting at the top of the leaderboard, isn’t exempt. In February 2026, the company dropped its earlier pledge to never train a system without being able to guarantee its safety measures were adequate in advance — a walk-back the review panel explicitly recommended reversing. And reviewers flagged something separate but related: the industry’s broader pivot toward military applications, with labs that once banned defense use now actively pursuing those contracts.

My honest take: a published safety framework is only as good as the company’s willingness to keep it when it becomes commercially inconvenient. This report is the second consecutive edition showing that willingness eroding across the board. If a specific safety commitment matters to your business — an evaluation gate, a data-handling promise, an incident-disclosure window — don’t take it on faith from a policy page. Get it written into your contract.

✓ What This Index Gets Right

  • Independent, expert-reviewed grading — not a company self-report
  • Recurring, twice-yearly cadence lets you track direction, not just a snapshot
  • Forces public comparison on commitments companies would rather keep vague
  • Flags a real industry-wide weak spot: existential and catastrophic-risk planning

✗ What It Cannot Tell You

  • Whether your specific model version and configuration behave safely
  • How the tool performs on your actual prompts, tools, and permission setup
  • Anything about a fine-tuned or self-hosted open-weight deployment
  • Whether the vendor’s stated policy will still hold six months from now

The Scope Trap: A Company Grade Is Not a Product Certificate

This is the single most important line in the entire report, and it’s easy to miss: FLI evaluates nine companies — their policies, disclosures, and governance — not the specific products you deploy. A C+ is not a safety certificate for the exact model version, API configuration, and contract terms your team is actually using.

I see this misunderstanding constantly in comment sections and client calls: people read “Anthropic scored highest” and conclude every Claude deployment is automatically the safest option, full stop. It isn’t that simple. A lab can lead on transparency while its production system still forces real tradeoffs at the deployment layer — the kind of thing I dug into when comparing how Claude Projects stacks up against ChatGPT’s custom GPTs for actual team workflows. The controls that determine whether your specific setup is safe — permission scoping, input filtering, human review gates — are things you and your team build, not something FLI’s panel can inspect from the outside.

Same logic applies in reverse. A failing lab-level grade doesn’t automatically mean every product built on that lab’s technology is unsafe — though in DeepSeek’s case specifically, I’d argue the correlation holds up better than FLI’s own scope caveat would suggest, based on what I’ve tested directly.

The Open-Weight Argument: Does the Index Play Fair?

Mistral — new to the index, and handed the single lowest score of the nine — pushed back hard on the methodology itself. Its argument: an open-weight release model hands fine-tuning, deployment, and safety controls to the enterprises that download the weights, not to Mistral. Judging Mistral by the same yardstick as a closed API provider like Anthropic or OpenAI, the company says, structurally penalizes the open-release model.

I think Mistral has a partial point, and I say that as someone who generally distrusts a company complaining about its own report card. An org-level policy index does favor closed-API vendors who control the entire stack end to end and can point to internal governance they own outright. It genuinely does penalize open-weight labs whose safety posture gets delegated downstream to whoever deploys the model. If you’re weighing a hosted API against an open-weight download for your own stack, read a low open-weight grade as “more of the safety work lands on your team” — not as “this model is inherently more dangerous than a closed one.”

Where I don’t extend Mistral much sympathy: the company also scored at or near the bottom on Safety Frameworks and Existential Safety — domains that measure whether the company has articulated any real risk strategy at all, independent of the open-weight question. That’s not a methodology quirk. That’s an actual gap.

My Buyer’s Framework: How I Actually Use This Report

Six FLI domains don’t map cleanly onto the questions a small business or solo operator actually needs answered. Here’s the translation I use when I’m advising someone on which AI tool to trust with real client work.

FLI Domain What It Tells a Buyer Still Check Yourself
Risk AssessmentHow hard the lab tests before releasingTest the model on your own real prompts
Current HarmsMisuse and brand-risk posture todayYour own content filters and monitoring
Safety FrameworksWhether commitments are credible on paperGet key promises written into your contract
Existential SafetyLong-horizon, catastrophic-risk planningNever single-source a critical workflow to one vendor
GovernanceInternal accountability maturityAsk about incident-response and escalation terms
Information SharingHow transparent the vendor actually isNegotiate audit rights and change-notification terms

In practice, this is the same due-diligence muscle I use when I’m testing any new AI product for this site — the same questions I asked while working through the AI SEO tools I now rely on for keyword and content research. Grade first, then verify. Never skip the second step.

Beyond FLI: Other Signals Worth Pairing With This Report

FLI’s index shouldn’t be the only external signal in your file. This summer, OpenAI and Anthropic separately ran a joint cross-lab safety evaluation — each testing its own internal misalignment and safety evaluations against the other’s publicly released models, then publishing results. That’s a different kind of transparency: adversarial, technical, and about actual model behavior rather than policy documents. If a vendor offers you something like that, it’s worth more than a glossy trust-center page.

Regulation is the other trendline worth watching. FLI’s president, Max Tegmark, has said he’s cautiously optimistic that binding rules — the EU AI Act, new Chinese AI rules, a more risk-conscious US posture — will eventually force a race to the top that a voluntary index alone can’t. Until that happens, this report is still the best independent, recurring, comparable read on lab-level behavior available to a working buyer. Use it as the start of your due diligence, not the finish line.

FAQ

What is the FLI AI Safety Index?

It’s a twice-yearly report card from the Future of Life Institute, an AI-safety nonprofit. The Summer 2026 edition, published July 7, 2026, grades nine leading AI companies across 37 indicators in six domains — Risk Assessment, Current Harms, Safety Frameworks, Existential Safety, Governance & Accountability, and Information Sharing — using evidence collected through June 3, 2026.

Which AI company got the best safety grade in 2026?

Anthropic, with a C+ and a score of 2.66 out of 4.0, leading five of the six graded domains. OpenAI (C, 2.28) and Google DeepMind (C, 2.01) followed.

Does a C+ mean Claude, or any AI tool, is safe to use for my business?

It means the company behind the tool has above-average safety governance compared to its peers — nothing more. It does not certify the specific model version, configuration, or workflow you’re running. Verify your own deployment separately; don’t treat a lab-level grade as a product-level guarantee.

Why did OpenAI’s grade drop from C+ to C?

The drop was small in raw score (2.31 to 2.28) but crossed a letter-grade boundary. It reflects the same industry-wide pattern the report flags across all four leading labs: weakened or walked-back pause commitments compared to the prior edition.

Is it dangerous to use free AI tools like DeepSeek?

“Dangerous” is strong, but I’d steer client-sensitive or confidential work away from any lab that failed this index and has no published safety framework, which is exactly DeepSeek’s position. For low-stakes drafting or brainstorming it’s a different risk calculus than for anything touching private data.

How should I actually use this index when picking an AI tool?

Use it to set how much diligence a vendor deserves, not as a final answer. Lean lighter on vetting for labs with strong, consistent grades across editions. Go deeper — your own testing, contract language, audit rights — for anything graded D or below, and for any open-weight model you’re deploying yourself.

Conclusion

An entire industry graded between F and C+ means the audit burden is on you, not the vendor.

Anthropic earns the closest thing to a leadership position here, and in my daily use, that lead holds up. But a C+ is still a C+ — the best-governed company in the industry, by an independent expert panel’s own math, is graded average at best. Treat this report the way I treat any vendor scorecard: a genuinely useful starting filter, never a substitute for testing the tool yourself on your own work before you trust it with anything that matters.

If you’re weighing your next AI tool purchase, I’d rather you read the grade, then go run your own week-long test the way I did — across the tools you’re actually considering, on the actual work you’ll be doing. That’s the only audit that tells you what a scorecard never can. For more on how these tools compare in real coding and content workflows, see how Claude Code stacks up against Cowork, what Claude Code can do that Cursor can’t, and whether ChatGPT is actually good for coding in practice. I’ve also broken down how Google’s AI Overviews are reshaping SEO and where AI content writing still falls short of human writing, both of which tie directly into how much you should trust a given model’s output unsupervised.

About the Author

Oyekale Olawale runs Websites2Know, an independent platform reviewing AI tools and SaaS software. He tests each tool across real workflows — not demos — and publishes reviews based on hands-on evaluation. Reviews are written independently; no vendors pay for favorable coverage.

Get Notified When New Reviews & Updates are Published

We don’t spam! Read our privacy policy for more info.

Advertisement