Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

RL-Obfuscation

This repository contains the code for the paper "RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?" [arxiv].

rl_obfuscation

Setup

To setup the environment in a virtual machine, you can use the docker setup present in cybershiptrooper/obfuscated_rl_env

Alternatively, you could run the following commands:

# DEFINE HF_TOKEN AND WANDB_TOKEN BEFORE RUNNING THIS 
export HF_TOKEN="<enter_token>"
export WANDB_TOKEN="<enter_token>" # or source bashrc/env after adding it there

git clone https://github.com/cybershiptrooper/RL_Obfuscation.git
git clone https://github.com/huggingface/trl.git
python -m venv .venv
source .venv/bin/activate

if ! echo $(nvcc --version) | grep -q "12.1"; then
    echo "CUDA version is not 12.1" # assert cuda version is 12.1
    exit 1
fi
pip install torch==2.2.2 torchvision==0.17.2 torchaudio==2.2.2 --index-url https://download.pytorch.org/whl/cu121
cd obfuscated_backdoors
source installation.sh
cd ..

cd trl
git checkout 49711efab9e0cc3762d3228c9fd5a8064d489503
pip install -e .
huggingface-cli login --token $HF_TOKEN
# put token in .bashrc
echo "export HF_TOKEN=$HF_TOKEN" >> ~/.bashrc
wandb login $WANDB_TOKEN
echo "source ~/.venv/bin/activate" >> ~/.bashrc
apt-get update 
apt-get install tmux -y
apt-get install git-lfs -y

# optional installs for visualizations
pip install circuitsvis
pip install plotly nbformat>=4.2.0 
pip install -U kaleido

source ~/.bashrc

Usage

There are three main scripts:

Fine-tunes a model on the Badllama dataset to create a helpful-only model. To run the script, please use

accelerate launch --multi_gpu sft_bad_llama.py \
  --model_id Qwen/Qwen2.5-14B-Instruct \
  --dataset_name palisaderesearch/Badllama-ICLR24 \
  --num_train_epochs 3 \
  --push_to_hub

Common overrides: --learning_rate, --per_device_train_batch_size, --gradient_accumulation_steps, --hf_username.

Trains and/or evaluates a probe on a model.

  • Run with a config file (recommended):
python eval_probe.py --config experiments/eval_probe/7b_base.yaml
  • Or run with CLI flags only:
python eval_probe.py \
  --model_name <model_name> \
  --dataset_name <dataset> \
  --probe_type linear \
  --probe_layers 10 \
  --aggregate_method mean \
  --load_from_disk false \
  --probe_save_path datasets/helpful_only_linear_probes.pth

More example configs can be found in experiments/eval_probe.

Performs RL-Obfuscation on a model.

  • Run with a config file:
python grpo_script.py --config experiments/8b_static/linear_mean.yaml
  • To enable curriculum training (threshold ramping), use:
python grpo_script.py --config experiments/base.yaml --curriculum

A template is provided at experiments/template_config.yaml. Concrete examples: experiments/base.yaml, experiments/backdoored.yaml, and experiments/hyperparam_plots/base.yaml.

  • Visualize completions and losses via a local server:
python log_server.py

Configs: how they work

All training/evaluation settings are defined in YAML config files. Configs support inheritance via base_config.

  • Inheritance: when a config contains base_config: path/to/parent.yaml, the parent is loaded first and then deeply merged with the child (child values override parent; nested dictionaries merge; lists are replaced, not merged).

Example:

# experiments/base.yaml
model:
  name: model_a
run_desc: grpo_linear_mean
training:
  num_train_epochs: 3
  beta: 0.05
whitebox_models:
  - type: threshold
    probe_threshold: 0.3
    aggregate_method: mean
probes:
  type: linear
  load_from_disk: true
  path: /root/obfuscated_backdoors/datasets/helpful_only_probes.pth
  layers: [10]
# experiments/derived.yaml
base_config: experiments/base.yaml
training:
  num_train_epochs: 5   # overrides parent
whitebox_models:
  - type: threshold     # replaces the whole list
    probe_threshold: 0.45
    aggregate_method: mean

Run with:

python grpo_script.py --config experiments/derived.yaml

Config reference (RL-Obfuscation)

Below are the top-level sections expected by grpo_script.py and their fields.

  • Model
Field Type Description
name str Hugging Face model id to train (policy).
hf_token str or null HF token; if null, taken from HF_TOKEN env var (required for loading/pushing private models).
revision str (optional) Model revision/commit/tag to load.
load_in_8bit bool (optional) Load policy in 8-bit; disables requires_grad_ if true.
  • Dataset
Field Type Description
name str HF dataset id to load when load_from_disk is false.
load_from_disk bool Whether to load a prepared dataset dict from disk.
local_path str Path passed to DatasetDict.load_from_disk when load_from_disk is true.
train_splits list[Split] Training splits (see Split below).
test_splits list[Split] Evaluation splits (see Split below).
  • Split (items inside train_splits / test_splits)
Field Type Description
split_name str Name of split in the dataset (must exist).
n_examples int or "all" Number of examples to sample (or all).
label 0 or 1 Label for examples in the split.
backdoored bool Whether examples are backdoored.
prepend_str str (optional) String to prepend to prompt.
replace_str str (optional) String to remove from prompt.
  • Probe dataset (probe_dataset)
Field Type Description
name str HF dataset id for probe-training examples.
load_completions_from_disk bool If true, load pre-generated completions from disk.
completions_path str Path to prepared dataset with completions (when loading from disk).
jailbreak_dataset_name str (optional) HF dataset for benign/jailbreak items when needed.
  • Probes (probes)
Field Type Description
type str One of linear, non_linear, maxpool, attention.
load_from_disk bool Load probes from path if true, otherwise train then save.
path str File path for loading/saving probe weights.
layers list[int] Layers to probe (must be subset of saved probes if loading).
training.n_epochs int Epochs when training new probes.
training.batch_size int Batch size when training new probes.
  • Whitebox reward models (whitebox_models: list of items)
Field Type Description
type str threshold, continuous, or gaussian (reward function type).
probe_threshold float Threshold (used by threshold and as center for others).
negative_is_1 bool Whether the negative class is 1 in datasets.
max_length int Max tokens for activation caching.
cache_batch_size int (optional) Activation cache batch size.
verbose bool Log layer-level rewards and details.
aggregate_method str mean, max, min over generation tokens.
weight float Weight multiplier for this model’s reward.
  • Blackbox reward models (blackbox_models: list of items)
Field Type Description
name str Built-in reward name: equivalent, strongreject, or length.
weight float Weight multiplier for this model’s reward.
config dict (optional) Model-specific params. For equivalent: true_token, false_token. For length: target_length, scale_factor, `mode: token
  • Training (training)
Field Type Description
num_train_epochs int Number of RL epochs.
logging_steps int Logging frequency.
save_strategy str steps or no.
save_steps int Checkpoint frequency (if steps).
push_to_hub bool Push checkpoints to Hub.
temperature float Sampling temperature for GRPO.
beta float KL penalty weight.
learning_rate float Optimizer learning rate.
warmup_steps int LR warmup steps.
  • Curriculum (curriculum_args)
Field Type Description
num_curriculum_steps int Number of threshold increments.
curriculum_start float Starting threshold for curriculum.
  • LoRA (optional top-level lora section)
Field Type Description
r int LoRA rank.
lora_alpha int LoRA scaling.
lora_dropout float Dropout for LoRA adapters.
target_modules list[str] or null Names of modules to adapt (model-dependent).
bias str none (recommended).
layers list[int] or null Specific layers to transform (if supported).

Notes:

  • If whitebox_models is omitted, training proceeds with only blackbox rewards (a warning is printed).
  • Legacy reward_model (single whitebox model) is supported for backward compatibility but use whitebox_models instead.

Config reference (Probe training/eval)

Fields accepted by eval_probe.py (via config or CLI flags):

Field Type Description
model_name str HF model id whose representations the probes read.
model_revision str (optional) Model revision/commit/tag.
device_map str Device map for loading (e.g., auto).
dataset_name str HF dataset id to use.
load_from_disk bool Whether to load dataset with completions from disk.
dataset_disk_path str Path to DatasetDict.load_from_disk with completions.
jailbreak_dataset_name str HF dataset name for jailbreak/benign splits (if needed).
jailbreak_split_name str Split name inside the jailbreak dataset.
probe_type str One of linear, non_linear, maxpool, attention.
probe_layers list[int] List of layers to probe.
train_new_probes bool Train new probes if true; otherwise load from probe_save_path.
batch_size int Batch size for probe train/eval.
n_epochs int Epochs for probe training.
probe_save_path str Path to save/load probe weights ({probe_type} is formatted).
max_length int Max length for tokenization during eval.
aggregate_method str mean, max, topk_mean, median, last_token.
num_bins int Bins for histograms.
log_yaxis bool Log scale for histogram y-axis.
save_plots bool Save figures to plot_path.
plot_path str Output path for plots ({probe_type} and {aggregate_method} formatted).
verbose bool Verbose logging.
show_plots bool Display plots interactively.
fprs list[float] FPR points for threshold/TPR tables.

See experiments/eval_probe/7b_base.yaml and experiments/eval_probe/weak_to_strong/7b_to_8b.yaml for examples.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages