AEF-1 Standard

AEF-1: Minimum Operating Conditions for Independent Third Party AI Evaluations

Version 1, updated December 4, 2025

Independent third-party evaluations of AI systems are becoming increasingly central to AI governance approaches pursued by developers, governments, enterprises, and the public. However, the operating conditions of third-party evaluations can be opaque, even though they can significantly impact whether an evaluation is trustworthy and impartial. To address this, we present AEF-1, a standard and checklist that third-party evaluators can use to demonstrate how they achieved a set of operating conditions that support a baseline level of independence, access, and transparency during an evaluation.

The success of AEF-1 requires adoption by a range of stakeholders, including frontier developers.

The standard

The full standard, covering all five principles and the specific conditions under each.

The checklist

Evaluators demonstrate adherence by completing the checklist and publishing it alongside their results. Where a requirement cannot be met literally, the checklist is used to document which conditions were not met, and why.

A LaTeX template is coming soon.

AEF-1 in practice

Evaluations by Forum members carried out or documented under the conditions the standard describes.

Transluce logo

Mental Health Evaluation

Transluce · August 2026

An independent assessment of how 77 model variants respond to users experiencing mental health crises, published with a detailed disclosure of operating conditions consistent with AEF-1.

How AEF-1 was used: Included the AEF-1 checklist, disclosing the broader operating conditions the evaluation ran under and what may have affected its independence, access, and transparency.

See the evaluationSee the checklist

METR logo

Frontier Risk Report

METR · May 2026

An assessment of misalignment risk from AI agents deployed internally at Anthropic, Google, Meta, and OpenAI, carried out with more direct access to non-public information and more editorial independence than previous external evaluation engagements.

How AEF-1 was used: Included the AEF-1 reporting checklist in the published report.

See the evaluationSee the checklist (pg 67)

SecureBio logo

Biosecurity pre-release assessment

SecureBio · September 2026

Independent pre-release assessment of biosecurity-relevant capabilities in OpenAI's GPT-6 Astra, including knowledge and agentic benchmarks, safeguard refusals, and expert red-teaming.

How AEF-1 was used: Included the AEF-1 checklist in the published report, and uses AEF-1 as a foundation of its organizational principles.

See the reportSee the checklistSee their principles

European AI Office logo

Independence provisions under the General-Purpose AI Code of Practice

European AI Office

The EU AI Office has endorsed key provisions of the AEF-1 standard as a means for providers to comply with the independence provisions of the General-Purpose AI Code of Practice.

What the standard covers

1

Sufficient Access and Resources

Technical access · Information · Computational resources · Time · Safe harbor

2

Minimized Conflicts of Interest

Contingent compensation · Organizational control · CoI policy · CoI disclosure · Recusals · Separate agreements

3

Analytic Autonomy

Scoping · Evaluation autonomy · Direct access · Editorial control

4

Transparent Methods and Results

Methodological transparency · Disclosure rights · No contingent release · No misrepresentation · Timely disclosure · Redactions · Redaction disclaimer

5

Protection of Sensitive Information

Publication terms · Evaluation integrity · Protecting confidential information · Responsible disclosure policy