This repository contains the code for the paper "RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?" [arxiv].
To setup the environment in a virtual machine, you can use the docker setup present in cybershiptrooper/obfuscated_rl_env
Alternatively, you could run the following commands:
# DEFINE HF_TOKEN AND WANDB_TOKEN BEFORE RUNNING THIS
export HF_TOKEN="<enter_token>"
export WANDB_TOKEN="<enter_token>" # or source bashrc/env after adding it there
git clone https://github.com/cybershiptrooper/RL_Obfuscation.git
git clone https://github.com/huggingface/trl.git
python -m venv .venv
source .venv/bin/activate
if ! echo $(nvcc --version) | grep -q "12.1"; then
echo "CUDA version is not 12.1" # assert cuda version is 12.1
exit 1
fi
pip install torch==2.2.2 torchvision==0.17.2 torchaudio==2.2.2 --index-url https://download.pytorch.org/whl/cu121
cd obfuscated_backdoors
source installation.sh
cd ..
cd trl
git checkout 49711efab9e0cc3762d3228c9fd5a8064d489503
pip install -e .
huggingface-cli login --token $HF_TOKEN
# put token in .bashrc
echo "export HF_TOKEN=$HF_TOKEN" >> ~/.bashrc
wandb login $WANDB_TOKEN
echo "source ~/.venv/bin/activate" >> ~/.bashrc
apt-get update
apt-get install tmux -y
apt-get install git-lfs -y
# optional installs for visualizations
pip install circuitsvis
pip install plotly nbformat>=4.2.0
pip install -U kaleido
source ~/.bashrcThere are three main scripts:
Fine-tunes a model on the Badllama dataset to create a helpful-only model. To run the script, please use
accelerate launch --multi_gpu sft_bad_llama.py \
--model_id Qwen/Qwen2.5-14B-Instruct \
--dataset_name palisaderesearch/Badllama-ICLR24 \
--num_train_epochs 3 \
--push_to_hubCommon overrides: --learning_rate, --per_device_train_batch_size, --gradient_accumulation_steps, --hf_username.
Trains and/or evaluates a probe on a model.
- Run with a config file (recommended):
python eval_probe.py --config experiments/eval_probe/7b_base.yaml- Or run with CLI flags only:
python eval_probe.py \
--model_name <model_name> \
--dataset_name <dataset> \
--probe_type linear \
--probe_layers 10 \
--aggregate_method mean \
--load_from_disk false \
--probe_save_path datasets/helpful_only_linear_probes.pthMore example configs can be found in experiments/eval_probe.
Performs RL-Obfuscation on a model.
- Run with a config file:
python grpo_script.py --config experiments/8b_static/linear_mean.yaml- To enable curriculum training (threshold ramping), use:
python grpo_script.py --config experiments/base.yaml --curriculumA template is provided at experiments/template_config.yaml. Concrete examples: experiments/base.yaml, experiments/backdoored.yaml, and experiments/hyperparam_plots/base.yaml.
- Visualize completions and losses via a local server:
python log_server.pyAll training/evaluation settings are defined in YAML config files. Configs support inheritance via base_config.
- Inheritance: when a config contains
base_config: path/to/parent.yaml, the parent is loaded first and then deeply merged with the child (child values override parent; nested dictionaries merge; lists are replaced, not merged).
Example:
# experiments/base.yaml
model:
name: model_a
run_desc: grpo_linear_mean
training:
num_train_epochs: 3
beta: 0.05
whitebox_models:
- type: threshold
probe_threshold: 0.3
aggregate_method: mean
probes:
type: linear
load_from_disk: true
path: /root/obfuscated_backdoors/datasets/helpful_only_probes.pth
layers: [10]# experiments/derived.yaml
base_config: experiments/base.yaml
training:
num_train_epochs: 5 # overrides parent
whitebox_models:
- type: threshold # replaces the whole list
probe_threshold: 0.45
aggregate_method: meanRun with:
python grpo_script.py --config experiments/derived.yamlBelow are the top-level sections expected by grpo_script.py and their fields.
- Model
| Field | Type | Description |
|---|---|---|
| name | str | Hugging Face model id to train (policy). |
| hf_token | str or null | HF token; if null, taken from HF_TOKEN env var (required for loading/pushing private models). |
| revision | str (optional) | Model revision/commit/tag to load. |
| load_in_8bit | bool (optional) | Load policy in 8-bit; disables requires_grad_ if true. |
- Dataset
| Field | Type | Description |
|---|---|---|
| name | str | HF dataset id to load when load_from_disk is false. |
| load_from_disk | bool | Whether to load a prepared dataset dict from disk. |
| local_path | str | Path passed to DatasetDict.load_from_disk when load_from_disk is true. |
| train_splits | list[Split] | Training splits (see Split below). |
| test_splits | list[Split] | Evaluation splits (see Split below). |
- Split (items inside
train_splits/test_splits)
| Field | Type | Description |
|---|---|---|
| split_name | str | Name of split in the dataset (must exist). |
| n_examples | int or "all" | Number of examples to sample (or all). |
| label | 0 or 1 | Label for examples in the split. |
| backdoored | bool | Whether examples are backdoored. |
| prepend_str | str (optional) | String to prepend to prompt. |
| replace_str | str (optional) | String to remove from prompt. |
- Probe dataset (
probe_dataset)
| Field | Type | Description |
|---|---|---|
| name | str | HF dataset id for probe-training examples. |
| load_completions_from_disk | bool | If true, load pre-generated completions from disk. |
| completions_path | str | Path to prepared dataset with completions (when loading from disk). |
| jailbreak_dataset_name | str (optional) | HF dataset for benign/jailbreak items when needed. |
- Probes (
probes)
| Field | Type | Description |
|---|---|---|
| type | str | One of linear, non_linear, maxpool, attention. |
| load_from_disk | bool | Load probes from path if true, otherwise train then save. |
| path | str | File path for loading/saving probe weights. |
| layers | list[int] | Layers to probe (must be subset of saved probes if loading). |
| training.n_epochs | int | Epochs when training new probes. |
| training.batch_size | int | Batch size when training new probes. |
- Whitebox reward models (
whitebox_models: list of items)
| Field | Type | Description |
|---|---|---|
| type | str | threshold, continuous, or gaussian (reward function type). |
| probe_threshold | float | Threshold (used by threshold and as center for others). |
| negative_is_1 | bool | Whether the negative class is 1 in datasets. |
| max_length | int | Max tokens for activation caching. |
| cache_batch_size | int (optional) | Activation cache batch size. |
| verbose | bool | Log layer-level rewards and details. |
| aggregate_method | str | mean, max, min over generation tokens. |
| weight | float | Weight multiplier for this model’s reward. |
- Blackbox reward models (
blackbox_models: list of items)
| Field | Type | Description |
|---|---|---|
| name | str | Built-in reward name: equivalent, strongreject, or length. |
| weight | float | Weight multiplier for this model’s reward. |
| config | dict (optional) | Model-specific params. For equivalent: true_token, false_token. For length: target_length, scale_factor, `mode: token |
- Training (
training)
| Field | Type | Description |
|---|---|---|
| num_train_epochs | int | Number of RL epochs. |
| logging_steps | int | Logging frequency. |
| save_strategy | str | steps or no. |
| save_steps | int | Checkpoint frequency (if steps). |
| push_to_hub | bool | Push checkpoints to Hub. |
| temperature | float | Sampling temperature for GRPO. |
| beta | float | KL penalty weight. |
| learning_rate | float | Optimizer learning rate. |
| warmup_steps | int | LR warmup steps. |
- Curriculum (
curriculum_args)
| Field | Type | Description |
|---|---|---|
| num_curriculum_steps | int | Number of threshold increments. |
| curriculum_start | float | Starting threshold for curriculum. |
- LoRA (optional top-level
lorasection)
| Field | Type | Description |
|---|---|---|
| r | int | LoRA rank. |
| lora_alpha | int | LoRA scaling. |
| lora_dropout | float | Dropout for LoRA adapters. |
| target_modules | list[str] or null | Names of modules to adapt (model-dependent). |
| bias | str | none (recommended). |
| layers | list[int] or null | Specific layers to transform (if supported). |
Notes:
- If
whitebox_modelsis omitted, training proceeds with only blackbox rewards (a warning is printed). - Legacy
reward_model(single whitebox model) is supported for backward compatibility but usewhitebox_modelsinstead.
Fields accepted by eval_probe.py (via config or CLI flags):
| Field | Type | Description |
|---|---|---|
| model_name | str | HF model id whose representations the probes read. |
| model_revision | str (optional) | Model revision/commit/tag. |
| device_map | str | Device map for loading (e.g., auto). |
| dataset_name | str | HF dataset id to use. |
| load_from_disk | bool | Whether to load dataset with completions from disk. |
| dataset_disk_path | str | Path to DatasetDict.load_from_disk with completions. |
| jailbreak_dataset_name | str | HF dataset name for jailbreak/benign splits (if needed). |
| jailbreak_split_name | str | Split name inside the jailbreak dataset. |
| probe_type | str | One of linear, non_linear, maxpool, attention. |
| probe_layers | list[int] | List of layers to probe. |
| train_new_probes | bool | Train new probes if true; otherwise load from probe_save_path. |
| batch_size | int | Batch size for probe train/eval. |
| n_epochs | int | Epochs for probe training. |
| probe_save_path | str | Path to save/load probe weights ({probe_type} is formatted). |
| max_length | int | Max length for tokenization during eval. |
| aggregate_method | str | mean, max, topk_mean, median, last_token. |
| num_bins | int | Bins for histograms. |
| log_yaxis | bool | Log scale for histogram y-axis. |
| save_plots | bool | Save figures to plot_path. |
| plot_path | str | Output path for plots ({probe_type} and {aggregate_method} formatted). |
| verbose | bool | Verbose logging. |
| show_plots | bool | Display plots interactively. |
| fprs | list[float] | FPR points for threshold/TPR tables. |
See experiments/eval_probe/7b_base.yaml and experiments/eval_probe/weak_to_strong/7b_to_8b.yaml for examples.