- Linux + NVIDIA GPU (recommended)
- Recent NVIDIA drivers + CUDA toolkit compatible with your PyTorch build (example below uses CUDA 11.8)
mkdir -p ~/miniconda3
wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh -O ~/miniconda3/miniconda.sh
bash ~/miniconda3/miniconda.sh -b -u -p ~/miniconda3
rm -rf ~/miniconda3/miniconda.sh
~/miniconda3/bin/conda init bashRestart your shell or:
source ~/.bashrcconda create -n pfp python=3.10 -y
conda activate pfpgit clone https://github.com/a-r-j/ProteinWorkshop
cd ProteinWorkshop
pip install -e .Install PyTorch (example: CUDA 11.8 wheels):
pip install torch==2.1.2+cu118 torchvision==0.16.2+cu118 torchaudio==2.1.2+cu118 \
--index-url https://download.pytorch.org/whl/cu118Additional utilities:
pip install rootutils cafaeval
pip install umap-learn pandas matplotlib datashader bokeh holoviews scikit-image colorcetIf ProteinWorkshop provides an example env file:
# If present:
# cp .env.example .env
# otherwise:
touch .envEdit .env:
vi .envExample:
ROOT_DIR="/home/user/ProteinWorkshop"
RUNS_PATH="/home/user/ProteinWorkshop/runs"
DATA_PATH="/home/user/ProteinWorkshop/data"
WANDB_API_KEY="" # optional
WANDB_ENTITY="gokceuludogan" # optional
WANDB_PROJECT="puffin" # optionalThis project relies on the ProteinWorkshop GO datasets (GO-MF) and (optionally) InterPro annotations for unit-boundary evaluation.
go-mf
data/GeneOntology/test_interpro.json(InterPro annotations)
All datasets are expected to live under the directory specified by:
DATA_PATH=/path/to/dataMake sure this is correctly set in your .env file (see Installation).
data/
└── GeneOntology/
GO datasets are downloaded via ProteinWorkshop.
Configure your .env:
DATA_PATH="/path/to/ProteinWorkshop/data"Then run:
workshop download go-mf
workshop download go-bp
workshop download go-ccThis will populate:
data/GeneOntology/
├── nrPDB-GO_train.txt
├── nrPDB-GO_valid.txt
├── nrPDB-GO_test.txt
├── nrPDB-GO_annot.tsv
└── ...
These datasets are used for:
- GO prediction (MF / BP / CC)
- Clustering evaluation
- Unit discovery benchmarks
To produce IA scores for PDB with different versions of GO ontology
python src/data/ia.py --annot .\data\GeneOntology\ground_truth\terms.tsv \
--graph .\data\go-basic-2020-06-01.obo --prop \
--outfile data\IA-nrPDB-go-basic-2020-06-01.txt- Train a model (PUFFIN / Protygus / mincut variants)
- Extract protein units (segments) from trained models
- Characterize units (size, structure, connectivity, random baselines)
- Evaluate unit function (GO neighborhood tests)
- Analyze unit–InterPro correspondence (optional, deeper analysis)
Each step is modular and can be run independently.
Training uses Hydra (src/train.py).
PUFFIN is typically trained with a dual objective:
- protein-level GO prediction
- unit discovery
python src/train.py \
name=puffin \
encoder=puffin \
encoder.gnn_type=GAT \
encoder.hidden_dim=512 \
encoder.num_clusters=64 \
encoder.num_res_gnn_layers=2 \
encoder.num_seg_gnn_layers=2 \
encoder.proj_layer=true \
encoder.fuse_lm_method=sum \
objective_type=dual \
function_weight=1.0 \
unit_weight=0.5What this does
- Trains a PUFFIN model with 64 latent units
- Uses a GAT-based residue and unit GNN
- Optimizes both GO prediction and unit discovery
- Saves checkpoints under
models/puffin/
After training, units are extracted using src/cluster.py.
python src/cluster.py \
encoder=puffin \
encoder.num_clusters=64 \
ckpt_path models/puffin/epoch_*.ckpt \
cluster.input_file data/GeneOntology/nrPDB-GO_test.txt \
cluster.split test \
output_dir units/puffin/test/Outputs
test_residue_assignments.csvtest_segment_embeddings.npytest_segment_metadata.csv
These define residue → unit assignments and unit embeddings.
This step analyzes extracted units:
- size distributions
- structural compactness
- contact-based coherence
- comparison to size-matched random baselines
python src/segment_characterize.py \
--cluster_dir ismb26/segments/puffin_K64/test \
--output_dir ismb26/results/segment_reports/puffin_K64/test \
--prefix test \
--structure_dir data/pdb_chain \
--contact_cutoff 10.0 \
--random_baselineUnits are evaluated by checking whether nearby units in embedding space share GO functions.
python src/unit_func_eval.py \
--segments_root units/ \
--model_name puffin_K64 \
--split valid \
--annotation_dir data/GeneOntology \
--go_aspect MF \
--out_root results/unit_func_reports \
--k_neighbors 50 \
--n_queries 5000For deeper biological validation, units can be compared to InterPro annotations:
- map InterPro regions to best-matching units (IoU)
- propagate InterPro2GO mappings
- evaluate GO ranking quality (Hit@K, MRR)
This analysis combines:
-
unit assignments
-
cluster enrichment results
-
InterPro JSON + InterPro2GO mappings
-
Associated script:
src/interpro_proto_go_term_comparison.py
The same pipeline applies to:
- Mincut pooling (unsupervised learned units)
- ESM + k-means (structure-agnostic baseline)
All baselines produce compatible outputs under:
ismb26/segments/<model_name>/<split>/
so they can be evaluated identically.
ismb26/
├── models/ # trained checkpoints
├── units/ # extracted units
│ └── puffin/
│ ├── train/
│ ├── valid/
│ └── test/
├── results/
│ ├── unit_reports/ # structural characterization
│ ├── unit_func_reports/ # GO neighborhood eval
│ └── func_eval/ # protein-level eval logs
If you use this repository, please cite the following related paper:
@article{10.1093/bioinformatics/btag265,
author = {Uludoğan, Gökçe and Giledereli, Buse and Ozkirimli, Elif and Özgür, Arzucan},
title = {PUFFIN: protein unit discovery with functional supervision},
journal = {Bioinformatics},
volume = {42},
number = {Supplement_1},
pages = {btag265},
year = {2026},
month = {07},
issn = {1367-4811},
doi = {10.1093/bioinformatics/btag265},
url = {https://doi.org/10.1093/bioinformatics/btag265},
}This code base is licensed under the MIT license. See LICENSE for details.