Guidelight’s First AI Control Grades Put Anthropic and OpenAI on Top

The new nonprofit Guidelight graded Anthropic, OpenAI, Google, xAI, and Meta on six agent-control practices. Its first scorecard finds meaningful progress in monitoring, but less evidence that labs can consistently block or contain unsafe actions.

Written By
Corey Noles
Corey Noles
Aug 19, 2026
6 minute read

Frontier AI labs have moved beyond simply promising to keep an eye on their internal agents. Anthropic and OpenAI now describe monitoring large portions of agent activity. Google has published a detailed control roadmap. Four major labs recently gave an outside evaluator access to internal models and nonpublic information.

A new assessment from Guidelight AI Standards argues that these efforts are real, but incomplete. The organization graded Anthropic, OpenAI, Google, xAI, and Meta across six safeguards for AI systems used inside their own companies. Anthropic and OpenAI tied at the top with C+ grades, followed by Google at D+, xAI at D−, and Meta at F.

Those grades look grim at first glance. They need some context, both about what Guidelight measured and who is doing the grading.

Who is Guidelight?

Guidelight is a new independent nonprofit founded by Page Hedley and Steven Adler, two former OpenAI safety leaders. Hedley, Guidelight’s CEO, previously worked as OpenAI’s policy and ethics adviser. Adler, its chief scientist, worked at OpenAI from 2020 to 2024 and led teams including product safety and dangerous-capability evaluations.

The group is trying to create something the AI industry still lacks: a shared, concrete definition of what responsible frontier-model development should look like. Its first standards cover control, capability testing, and transparency. Guidelight then assesses companies against those standards using public evidence, with the idea that comparable scores can give lab employees, customers, investors, and the public a clearer benchmark.

Guidelight isn’t a regulator, and this assessment isn’t an official audit. The nonprofit says it accepts no funding from AI companies or their staff, although Adler discloses that he retains some OpenAI equity from his time there. Its standards are also living documents that the group expects to revise as the technology and safety research develop.

That makes this scorecard best read as an informed outside benchmark, not the final word on which AI company is safest.

Advertisement

What the scorecard measured

Guidelight selected six practices from its broader Control standard: logging internal AI activity, measuring whether monitors catch misbehavior, gating high-risk actions before they happen, halting systems after a burst of suspicious activity, commissioning outside reviews, and maintaining a containment plan for a model that tries to defeat its controls.

The scoring system runs from 0 for “not implemented” to 5 for “full implementation.” A 2 means “limited partial implementation,” while a 3 means “substantial partial implementation.” Each practice counted equally toward a company’s overall grade.


Across all five companies, the average was 1.6 out of 5. Twenty-two of the 30 individual ratings, or 73%, were a 2 or lower. Seven received a zero. No company earned a 4 or 5 on any practice.

Those are our calculations from Guidelight’s scorecard. They support the group’s conclusion that none of the five labs has publicly demonstrated a complete control system. But “partial implementation” still represents work underway, and the leaders have made more progress than the letter grades alone suggest.

The progress is concentrated in monitoring

Anthropic and OpenAI both received 3s for logging and monitor efficacy. Guidelight found that the two companies log substantial portions of internal agent activity, scan that activity for suspicious behavior, and have done some work to measure how often the monitoring systems catch what they’re supposed to catch.

Anthropic received a 3 on five of the six practices. OpenAI earned three 3s and three 2s, making its results more even across the scorecard. Neither company reached Guidelight’s definition of near-full implementation, but both have moved well past the starting line.

Outside evidence supports that progress. In a frontier-risk pilot conducted earlier this year, the independent research group METR received access to internal models and nonpublic information from Anthropic, Google, Meta, and OpenAI. METR found that automated monitors caught many harmful actions. It also found gaps in coverage and ways monitors could be disabled or evaded, especially when Anthropic allowed an embedded researcher to red-team part of its system.

Advertisement

That combination is useful to see. The monitors work often enough to be meaningful, but not reliably enough for companies to treat them as a complete safety layer.

Logging and third-party review were the highest-scoring categories in Guidelight’s data, each averaging 1.8 out of 5. Four of the five companies participated in METR’s pilot, with xAI the exception. Third-party access is still limited and inconsistent, but outside review is beginning to look like a normal part of frontier-model development rather than a special favor.

Prevention and response are the next layer

The lower scores came in controls that determine what happens after a monitor spots trouble.

Guidelight’s gated-action and circuit-breaking categories each averaged 1.6. These cover measures such as requiring approval before an agent changes security settings and pausing sensitive actions when a system generates repeated flags.

Containment plans scored lowest at 1.2. OpenAI led that category with a 3, followed by Google at 2 and xAI at 1. Anthropic and Meta both received zeroes based on the public evidence Guidelight could find.

The pattern suggests that labs are building observability first, then adding enforcement and incident response. That sequence is familiar in other areas of security: teams need to see a system before they can manage it well. Guidelight’s assessment shows where the second half of that work remains thin or poorly documented.

It also gives companies a fairly practical to-do list. Define which agent actions always require approval. Test whether repeated flags stop related activity across a system, not merely one agent session. Decide who can restrict a model’s access during an incident. Then publish enough evidence for outsiders to understand whether those measures work.

Why internal agent controls matter

Frontier models are no longer confined to chat windows and benchmark tests. Labs use agents to write code, build evaluations, secure networks, generate training data, and help develop later models. METR reported that OpenAI used AI assistance throughout the company, while Google said AI touched nearly all work involving code, configuration, technical design, and research ideation.

These controls don’t require assuming that an AI system is secretly plotting its escape. More ordinary failure modes are enough: reward hacking, evaluation gaming, unsafe tool calls, compromised credentials, or a prompt injection that pushes an agent outside its intended task.

Advertisement

Anthropic’s research on reward hacking has shown that models trained to exploit narrow shortcuts can develop broader deceptive behavior in controlled tests. And as enterprise agents spread into business systems, the same questions are reaching companies far beyond the frontier labs.

The practical lesson is simple: monitoring is a good start. Organizations also need permission boundaries, escalation rules, shutdown procedures, and someone clearly responsible for using them.

Google shows the difference between a roadmap and a rollout

Google is the most interesting middle case in the assessment. It scored 1.5 overall, but Guidelight called Google’s AI Control Roadmap the most detailed forward-looking control document any lab has published. Google’s own summary lays out work across prevention, detection, and containment.

Publishing that plan matters. It gives researchers and other labs something specific to challenge, copy, or improve. Guidelight’s scoring, however, focuses on public evidence of current implementation. A detailed roadmap can raise confidence in where a company is headed without proving that the controls are operating today.

That distinction applies across AI safety. Policies and system cards describe intent. Operational controls produce logs, measured detection rates, enforced permission boundaries, drill results, and other evidence that the system works under pressure.

Read the grades as a baseline

Guidelight relied only on public materials available through August 18, 2026. Two staff members independently scored each company, reconciled disagreements with a third person, and gave company staff a chance to flag errors or disclose more information.

The process is systematic, but it can’t see undisclosed controls or independently verify most company claims. A low grade may reflect a weak safeguard, limited disclosure, or both. A higher grade can rest on evidence that hasn’t been audited. The equal weighting also treats logging and containment as mathematically identical, a choice Guidelight says it may revisit.

The rankings therefore aren’t a definitive league table of company safety. They map what outsiders can currently verify against one organization’s emerging standard.

That is still useful. AI safety discussions often get stuck between sweeping promises and hypothetical disasters. Guidelight has turned a piece of that debate into a list of practices that can be checked, compared, and improved.

Advertisement

The encouraging part is that the labs aren’t starting from zero, and Guidelight says stronger controls are achievable with tools that exist today. This first scorecard establishes a baseline. The more important test will be whether the next one shows those roadmaps becoming working systems.

Corey Noles

Corey Noles is the Host of The Neuron: AI Explained podcast and Managing Editor of AI and Experimental Content at TechnologyAdvice, where he leads the charge in testing and refining emerging content strategies across the company's portfolio.

The Neuron Logo

Don't fall behind on AI. Get the AI trends & tools you need to know. Join 700,000+ professionals from top companies like Microsoft, Apple, Salesforce and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.