A python library built on top of Inspect AI to support Control research and evaluations, by Redwood and EquiStamp
-
Updated
Sep 9, 2026 - Python
A python library built on top of Inspect AI to support Control research and evaluations, by Redwood and EquiStamp
Research developing AI control protocols using task decomposition
🚦🗺️ UrbanFlow AI is a web app for generating 3D traffic simulations from real OpenStreetMap areas. It builds SUMO scenarios, runs microscopic vehicles, pedestrians, buses and trams, edits road events, controls real traffic lights with TraCI, trains JSON AI policies, saves models, and shows live metrics, charts, and notebooks.
An intelligent traffic management system that dynamically adjusts highway lane configurations using AI-powered congestion detection and a movable median barrier.
Judge-first framework where LLM outputs must converge under explicit, adversarial oracles.
In-depth exploration of Large Language Models (LLMs), their potential biases, limitations, and the challenges in controlling their outputs. It also includes a Flask application that uses an LLM to perform research on a company and generate a report on its potential for partnership opportunities.
Python client for Aegis — stabilize AI systems instantly with a simple API call.
Kho lưu trữ này chứa tài liệu, bài tập, và mã nguồn liên quan đến môn Trí tuệ nhân tạo trong điều khiển. Môn học tập trung vào ứng dụng AI trong các hệ thống điều khiển tự động, bao gồm lý thuyết và thực hành.
An attacker × monitor factorial in ControlArena: which model you pick as your trusted monitor matters more than its capability tier.
Independent AI governance and control standard
AI Governance — human-controlled code injection with oversight
White-box detection of collusion in an untrusted monitor: model organisms, linear probes, and a control evaluation that prices what they buy.
Benchmark for detecting insider threats by AI agents in a simulated frontier AI lab
Lichtarbeit
v0.8 preview · 2,288 multi-step agent trajectories · 513 hand-authored gold + 1,775 provenance-flagged augmented. A per-step benchmark that scores whether a verifier catches drift inside an agent's trajectory — not whether a prompt is harmful.
GG Tank Watch - frozen public-information archive of a resolved May 2026 chemical emergency. Conduit-only design; responsible-AI safety patterns enforced in code and tests.
Control protocols robust to a monitor-aware adversary — empirical AI-control eval on the APPS backdoor testbed (open-weight models).
A compact AI-control benchmark for trusted monitoring of untrusted coding agents.
Is cross-lingual chain-of-thought oversight failure monitor-side or model-side? A control-style deception eval on open-weight reasoning models.
Trusted monitors emit an integer suspiciousness score that discards information they already computed; reading the logit distribution instead is free and better (+0.031 AUROC, 36x more usable audit depth). Plus a negative result: linear probes on monitor activations are confounded by dataset artefacts, and a control that catches it.
To associate your repository with the ai-control topic, visit your repo's landing page and select "manage topics."