TechBriefe
Ai

OpenAI defines four core duties for external safety evaluators

Ana Maria Constantin 30.09.2026

Demanding Independent Verification of Model Risks

On Tuesday, OpenAI released a formal document outlining specific expectations for independent safety assessors. The company aims to clarify how these outside experts should interact with its internal teams. Lama Ahmad, who leads OpenAI’s external safety collaboration efforts, authored the guidance. The new framework seeks to standardize the review process for advanced AI models. It emphasizes the need for rigorous, third-party validation before major releases. This move signals a shift toward structured external oversight within the industry.

The primary goal is to grant assessors deep access to model capabilities and internal OpenAI states that this transparency allows reviewers to challenge fundamental assumptions made by the engineering team. Assessors must identify potential risks that internal staff might have overlooked due to proximity bias. They are expected to form independent conclusions regarding the effectiveness of current safeguards. This approach ensures that safety checks are not merely procedural but substantive. The company believes that fresh perspectives are critical for catching blind spots in complex systems.

The document specifies that assessors must verify whether the company’s risk mitigation strategies actually work in practice. Rather than accepting internal reports at face value, external parties must test the limits of the models themselves. This includes probing edge cases and stress-testing the system under unusual conditions. OpenAI wants reviewers to determine if the safeguards hold up against novel threats. The focus is on practical efficacy rather than theoretical compliance. By requiring this level of scrutiny, the company hopes to build greater public trust in its deployment timelines.

How External Reviewers Challenge Internal Assumptions

A key component of the new guidelines is the mandate for assessors to actively dispute internal narratives. Reviewers are encouraged to ask difficult questions about why certain risks were deemed acceptable. They must evaluate if the probability of harm was calculated correctly or if optimism bias influenced the decision-making process. This dynamic creates a necessary tension between the builders and the auditors. OpenAI argues that this friction improves the overall quality of the final product. The document highlights that assessors should feel empowered to push back against the development team’s preferred interpretations of data.

The implementation of these standards could reshape how major AI labs handle pre-release testing. If successful, this model may become a baseline for the broader technology sector. Companies will likely need to hire specialized external firms capable of executing these detailed reviews. The outcome depends on whether assessors can maintain true independence while working closely with the developers. Ultimately, the success of this initiative rests on the willingness of both parties to accept uncomfortable truths. A robust external check serves as a final line of defense before public release.

Frequently Asked Questions

Who authored the new safety assessment guidelines? Lama Ahmad wrote the document. She currently leads OpenAI’s work with external safety partners and assessors.

What is the main purpose of granting assessors deep access? Deep access enables reviewers to challenge internal assumptions and find missed risks. It allows them to reach independent conclusions about safeguard effectiveness.

When did OpenAI publish these expectations? The company released the document on Tuesday. It sets out the specific tasks required of outside safety evaluators.

Share:

More stories: