Stars
Our library for RL environments + evals
Collection of evals for Inspect AI
Code for 'The Linear Representation Hypothesis and the Geometry of Large Language Models' (ICML 2024)
Code for 'The Geometry of Categorical and Hierarchical Concepts in Large Language Models' (ICLR 2025, Oral)
A resource repository for representation engineering in large language models
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity.
Improving Alignment and Robustness with Circuit Breakers
Representation Engineering: A Top-Down Approach to AI Transparency
Steering vectors for transformer language models in Pytorch / Huggingface
Algebraic value editing in pretrained language models
Code release for the paper "Style Vectors for Steering Generative Large Language Models", accepted to the Findings of the EACL 2024.
A library for making RepE control vectors