Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit

TechnicalLithuaniaUploaded on Oct 7, 2026
Agent Barn is an open-source control plane for running and governing AI agents on an organisation's own infrastructure. It was built by AAI Labs in Lithuania. Agents connect to team chat tools such as Slack and Microsoft Teams, and act on systems such as GitHub and Jira. Access roles control who manages each agent. Audit logs record conversations and tool calls, and costs are attributed to each agent. It is released under the Apache 2.0 licence.

ProceduralUnited KingdomUploaded on Oct 6, 2026
The CREST Accreditation Standards set requirements for organisations that provide cyber security services, verified through independent assessment. They are published by CREST, an international not-for-profit body. Three components address AI. Domain 7 requires all accredited organisations to govern their use of AI, including validating AI-supported outputs. Annex B sets requirements for using AI in penetration testing, with human validation of findings. A third standard covers security testing of generative AI and large language model-enabled systems.

TechnicalUploaded on Oct 7, 2026
Mandare is an open-source accountability system for fleets of AI agents. It gives each agent a signed identity and a signed mandate setting its spending caps, permissions and approval requirements. A gateway enforces these limits outside the agent, with an offline kill switch. Every action, including refusals, is recorded in a tamper-evident ledger that an independent witness can check. Raw data never leaves the user's machine. It is released under the AGPL-3.0 and Apache 2.0 licences.

TechnicalIndiaUploaded on Oct 7, 2026
AiEnsured is a commercial testing platform for AI products, designed to support the deployment of responsible AI. It is developed in India. The platform brings together fairness and bias evaluation, explainability, adversarial robustness testing and privacy and GDPR compliance checks. It also generates test cases automatically, including rare and extreme inputs. Further features cover experiment management, model comparison and performance testing. After deployment, it monitors models for concept drift, which can reduce accuracy over time.

TechnicalJapanUploaded on Oct 7, 2026
The VeritasChain Protocol (VCP) is an open specification for cryptographically verifiable audit trails in algorithmic and AI-driven trading. It is maintained by the VeritasChain Standards Organization. VCP lets regulators and auditors verify mathematically that records of trading decisions, orders, risk controls and AI model use are complete and unaltered. A privacy module allows personal data to be erased while keeping the audit trail intact. VCP is designed to support MiFID II, EU AI Act and GDPR compliance.

TechnicalUploaded on Oct 2, 2026
DelusionEval is a dataset for evaluating how AI chatbots behave when conversations reinforce users' delusional beliefs. It was developed by Stanford University. It contains 725 excerpts from real, anonymised conversations, each labelled with one of 18 chatbot behaviours. These include harmful behaviours, such as endorsing delusions or facilitating self-harm, and protective ones, such as discouraging violence. Researchers use it to audit chatbot safety and develop detection methods. Access is restricted to non-commercial research.

TechnicalProceduralUploaded on Oct 2, 2026
Noisegate is an open-source demonstration of a gateway that lets AI agents query sensitive data without exposing individual records. Its privacy protections do not depend on the AI being trustworthy. A trusted layer checks every query against a policy, adds calibrated noise to answers and limits each user's privacy budget. Working attack demonstrations show its protections in action. The developer describes it as a demonstration, not a production product. It is released under the Apache 2.0 licence.

TechnicalIrelandUploaded on Oct 2, 2026
Diffprivlib is an open-source Python library for differential privacy, developed by IBM Research in Dublin. It adds calibrated random noise to data analysis and machine learning, so that no individual can be identified. The library provides privacy mechanisms, machine learning models with built-in privacy, data analysis tools and a privacy budget tracker. Its models work like those of scikit-learn, making them easy to adopt. It is intended for research and education, and was archived in September 2026.

TechnicalUploaded on Oct 2, 2026
pyPANTERA is an open-source Python package for obfuscating text using differential privacy. It was developed at the University of Padua. The package replaces words in a text with substitutes, so the original cannot be directly recovered, while keeping the text useful for tasks such as search. It implements seven published mechanisms in one framework, allowing users to reproduce and compare them. It won the Best Resource Paper Award at CIKM 2024 and is released under the GPL-3.0 licence.

Objective(s)


EducationalProceduralUnited KingdomUploaded on Sep 30, 2026
AI Compliancy is an online EU AI Act risk assessment, obligation and report tool for UK small and medium-sized businesses that use AI built into everyday software. Users check whether the Act can reach them, record the software they use, assess what they use each AI feature for, work through the obligations that follow, and produce a dated report.

TechnicalUploaded on Sep 23, 2026
Inner Warden is an open-source security agent for Linux and macOS servers. It detects attacks such as brute-force attempts and privilege escalation, and alerts operators in real time. It can use AI models to recommend responses, but AI remains advisory unless operators allow automatic action. Inner Warden can also monitor autonomous AI agents and block risky commands. All actions are reversible and recorded in an audit trail. It runs locally and is released under the MIT licence.

TechnicalUnited StatesUploaded on Sep 23, 2026
Sponsio is an open-source runtime safety tool for AI agents. It checks every agent action against deterministic rules, called agent contracts, before the action is executed. Contracts are based on formal methods and can allow, block, escalate or redirect actions. Every decision is logged in an audit trail. Users can apply ready-made contract bundles or draft rules in plain English. Sponsio works with major agent frameworks in Python and TypeScript, under the Apache 2.0 licence.

TechnicalKoreaUploaded on Sep 23, 2026
KSAFE-MM is a Korean-language benchmark for evaluating safety risks in multimodal large language models. It was developed by K-intelligence at KT in South Korea. The benchmark contains 14,135 query-image pairs across 11 risk categories, including hate, violence, privacy and weaponisation. One subset tests Korean culture-specific risks and applies jailbreak strategies such as role-play. Researchers use it to check whether models respond safely. It is available for research under a CC BY-NC 4.0 licence.

TechnicalUploaded on Sep 23, 2026
Doberman is an open-source security layer for AI coding agents such as Claude Code and Codex. It checks every action an agent takes before it runs. Routine actions pass, sensitive actions need human approval and dangerous actions are blocked. Any uncertainty results in the action being denied. Protections can tighten automatically, but weakening them requires human approval and is logged. Doberman is written in Python and released under the Apache 2.0 licence.

TechnicalUploaded on Sep 25, 2026
IndoBias is a benchmark for evaluating social bias in large language models in Indonesian and three local languages. It was developed by researchers at the Mohamed bin Zayed University of Artificial Intelligence and Universitas Indonesia. One track tests whether models prefer stereotypical sentences over neutral ones. A second track tests whether models describe local groups and institutions more positively or negatively. The results show strong stereotypical bias in current models, especially regarding ideology and religion in local languages.

TechnicalUploaded on Oct 2, 2026
F2Bench is a benchmark for evaluating fairness in large language models that also considers factual accuracy. It was developed by researchers at Inner Mongolia University and Xiamen University. The benchmark contains 2 568 test instances across ten demographic categories, including intersectional combinations. One task tests whether models accept biased conclusions after gradual stereotypical suggestions in a conversation. A second task tests whether models avoid stereotypes while still reflecting real statistics accurately. The code and dataset are publicly available.

Objective(s)

Related lifecycle stage(s)

Operate & monitorVerify & validate

TechnicalUnited StatesUploaded on Sep 26, 2026
M4 Bias Eval FairFace is a dataset for evaluating social bias in vision-language models. It was created by Hugging Face to assess its IDEFICS models. It contains around 11,000 face images labelled by perceived gender, ethnicity and age. For each image, two models wrote a résumé, a dating profile and a news article about an arrest. Comparing these texts across demographic groups reveals stereotypes the models have learned. The dataset is released under a CC BY 4.0 licence.

TechnicalKoreaUploaded on Sep 26, 2026
FLEX is a benchmark for testing whether large language models stay fair when prompts are designed to induce bias. It was developed by researchers at Korea University. The benchmark contains 3,145 multiple-choice questions from established fairness datasets, each combined with an adversarial prompt. These include assigning negative personas, forbidding refusals and introducing small text changes. Results show that models appearing fair on standard benchmarks can still be easily manipulated. The data and code are available for research purposes.

TechnicalUploaded on Sep 26, 2026
Gate AI: LLM Security Benchmark Evaluation Methodology and Results is a technical report on evaluating detectors of prompt injection and jailbreak attacks. It was published by Constellation Network. The method tests a detector across 16 public benchmarks with more than 12 000 samples. It uses a single detection threshold for all benchmarks and prevents near-duplicate prompts from inflating results. Competing detectors are compared at the same false-alarm rate. The report applies the method to Gate AI, Constellation Network's commercial security gateway.

TechnicalUnited StatesUploaded on Sep 26, 2026
CEB (Compositional Evaluation Benchmark for Fairness) is a fairness evaluation benchmark containing about 11 000 samples spanning multiple bias types, social groups, and tasks. It enables consistent measurement of compositional fairness behavior to identify failures that occur when fairness conditions combine across groups and bias factors. It is used by developers and auditors to support the fairness and accountability objectives, and to assess robustness of fairness performance under varied evaluation settings.

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.