SE(3)-Equivariant GNN Scoring for Molecular Docking
Installation | Quick Start | Documentation | Benchmark | Citation
PandaDock is a molecular docking suite combining a flexible-ligand conformational search with an SE(3)-equivariant Graph Neural Network scoring function that reaches Pearson R = 0.88 against experimental binding affinities on PDBbind.
The search treats every rotatable bond as an explicit degree of freedom. Position is sampled uniformly over the docking box, orientation uniformly over SO(3), and each Monte Carlo step is relaxed to a local minimum with a quasi-Newton optimizer using fully analytic gradients. Scoring runs against precomputed affinity grids, so a full search costs seconds to minutes rather than hours.
- Flexible-ligand search — all rotatable bonds searched, uniform SO(3) orientation sampling, L-BFGS relaxation with analytic gradients
- Grid-accelerated scoring — AutoDock Vina functional form, precomputed per atom type
- PandaDock-GNN — SE(3)-equivariant scoring, R = 0.88 on PDBbind
- Hybrid workflow — search with the empirical function, rank with the GNN
- Universal rescorer — rescore poses from any docking tool (Vina, Glide, GOLD)
- Rich output — poses with bond orders, complexes, interaction analysis, plots, PandaMap 2D diagrams and an HTML report
- Specialized modes — induced-fit, metal coordination, tethered docking
- Reproducible —
--seedfixes the search; every run records its parameters
Python 3.8+ is required. RDKit is a hard requirement and is most reliably installed from conda-forge.
conda create -n pandadock python=3.10
conda activate pandadock
conda install -c conda-forge rdkit
pip install pandadockFrom source:
git clone https://github.com/pritampanda15/PandaDock.git
cd PandaDock
pip install -e .Optional extras:
pip install -e ".[gnn]" # PandaDock-GNN (PyTorch, PyTorch Geometric)
pip install pandamap # 2D protein-ligand interaction diagrams# 1. Prepare structures (adds hydrogens, assigns protonation)
pandadock-prepare -r receptor.pdb -l ligand.sdf -o prepared/
# 2. Define the binding site
pandadock-gridbox -r receptor.pdb -m similarity \
--reference-ligand known_ligand.sdf -o box.json
# 3. Dock
pandadock dock -r receptor.pdb -l ligand.sdf -g box.json -o results/Or give coordinates directly:
pandadock dock -r receptor.pdb -l ligand.sdf \
--center 10 12 8 --box 22 22 22 -o results/| Option | Default | Description |
|---|---|---|
-r, --receptor |
required | Receptor PDB file |
-l, --ligand |
required | Ligand file (SDF/MOL2/PDB) |
-g, --grid-config |
— | Grid box JSON from pandadock-gridbox |
--center X Y Z |
— | Box centre in Å (alternative to -g) |
--box X Y Z |
— | Box dimensions in Å |
-s, --scoring |
vina |
vina or physics_based |
-n, --num-poses |
20 |
Maximum binding modes returned |
-e, --exhaustiveness |
auto | Independent search runs |
--seed |
random | Seed for reproducible runs |
--rigid-ligand |
off | Disable torsional search |
--grid-spacing |
0.375 |
Affinity grid spacing in Å |
--rescoring |
none |
none or mmgbsa |
-o, --output-dir |
docking_output |
Output directory |
--fast |
off | Reduced sampling for smoke tests only |
Unless set explicitly, exhaustiveness scales with the number of rotatable bonds:
| Rotatable bonds | 0 | 4 | 8 | 12+ |
|---|---|---|---|---|
| Independent runs | 8 | 16 | 24 | 32 |
A budget that is ample for a rigid fragment leaves a flexible ligand's search
space under-explored, and the failure is silent — the run returns a
confident-looking pose from a local minimum well above the global one. Raise -e
for large or highly rotatable ligands.
Two defaults worth knowing: runs are not reproducible without --seed, and
--fast drops to exhaustiveness 2, so its poses should not be reported as
results.
--num-poses is an upper bound. Poses are clustered at 2 Å heavy-atom RMSD, so a
pocket supporting only a few distinct modes returns fewer than requested. The
pose_diversity.png plot shows which case you are in.
A snug box outperforms a large one. Measured on confirmed search failures, reducing padding around the ligand from 8 Å to 5 Å lowered median best-of-N RMSD from 5.11 Å to 2.48 Å — and beat quadrupling exhaustiveness in a larger box, while running faster. Search volume matters more than sampling budget.
Every run writes:
| File | Contents |
|---|---|
report.html |
Run parameters and all plots in one page |
poses.sdf |
All poses with bond orders, rank, score, confidence |
pose{N}.pdb |
Individual ligand poses |
complex{N}.pdb |
Receptor plus ligand, ligand as HETATM chain L |
*_poses.json |
Full coordinates and per-term energies |
*_summary.json |
Run parameters, ensemble ΔG, runtime |
interaction_analysis.json |
Detected interactions, per pose |
pose_scores.png |
Scores by rank, with the rank-1/rank-2 gap |
energy_components.png |
Energy terms, favourable versus penalty |
pose_diversity.png |
Pairwise RMSD between returned poses |
interaction_fingerprint.png |
Residues contacted by each pose |
pandamap_2d_*.png |
2D interaction diagram (requires pandamap) |
Prefer poses.sdf for downstream work: PDB cannot represent bond orders or
formal charges, so viewers infer bonds from distance and routinely mis-assign
aromatic rings.
Interactions are detected with explicit chemistry — donor/acceptor matching including the protein backbone, hydrophobic typing, electrostatics, π-stacking, π-cation and metal coordination — not distance thresholds alone.
The score is an empirical docking score in kcal/mol. It ranks poses. It is not a measured binding free energy and should not be converted to a Kd or IC50 and reported as a potency prediction; use PandaDock-GNN for affinity.
The recommended workflow: search with the empirical function, rank with the GNN.
pandadock gnn download-model # ~82 MB
pandadock hybrid -r receptor.pdb -l ligand.sdf \
--center 10 12 8 --box 22 22 22 \
-m models/pandadock_gnn.pt -o results/Produces the same output set as dock, ranked by predicted pEC50 with the
empirical score retained alongside for comparison.
| Command | Description |
|---|---|
pandadock gnn download-model |
Download the pre-trained model |
pandadock gnn predict |
Predict binding affinity for a complex |
pandadock gnn rescore |
Rescore poses from any docking tool |
pandadock gnn train |
Train on ULVSH, PDBbind or a combined set |
pandadock gnn benchmark |
Evaluate on a test set |
pandadock gnn compare |
Compare against baseline scoring methods |
# Induced-fit: refines receptor side chains around the ligand
pandadock-flex -r receptor.pdb -l ligand.sdf --center 10 12 8 --radius 12 -o results/
# Metal coordination: metal-aware scoring with geometry constraints
pandadock-metal dock -r receptor.pdb -l ligand.sdf --center 10 12 8 --box 22 22 22 -o results/
# Tethered: restrains the ligand centroid near a reference pose
pandadock-tethered dock -r receptor.pdb -l ligand.sdf \
--ref reference.sdf -t 3.0 -o results/Tethered docking applies a flat-bottom centroid restraint inside the objective: free movement within the radius, harmonic cost beyond it.
Induced-fit docking redocks into every refined receptor, so cost scales with the
number of poses carried into refinement: budget hours rather than minutes for a
single ligand. Use --initial-poses-to-retain to control that trade-off. Output
matches dock: poses.sdf, ligand poses, and complexes with the refined
receptor. Refined receptors are written to a temporary directory and removed
when the run ends, including on failure. The IFD score is the binding energy plus
a weighted receptor-strain penalty, so a pose requiring more side-chain
rearrangement scores worse than an equivalent one that requires none.
Metal parameters fall back to built-in approximations when no AutoDock-format parameter file is supplied. Those are adequate for identifying coordination geometry but not for quantitative metal binding energies; pass a parameter file for that.
| Command | Description |
|---|---|
pandadock-prepare |
Add hydrogens, assign protonation, generate 3D |
pandadock-gridbox |
Define binding sites and write grid box configs |
pandadock-report |
Regenerate plots and reports from a results directory |
| Mode | Use for |
|---|---|
similarity --reference-ligand LIG |
Box centred on a known ligand (redocking) |
cavities |
Blind docking — detects pockets without a ligand |
residues --residues A:123,A:145 |
Box around specified residues |
manual --center X Y Z --box X Y Z |
Explicit coordinates |
For redocking or any case where a ligand pose is known, use
similarity --reference-ligand. Cavity detection may centre the box on a
neighbouring pocket, placing the true site near the box edge where sampling is
poorer.
pandadock is the flexible-ligand search. Legacy names are retained for
backwards compatibility and all resolve to the same algorithm:
| Name | Status |
|---|---|
pandadock / pandacore |
Current flexible-ligand Monte Carlo search |
monte_carlo_cpu |
Deprecated alias |
genetic_algorithm_cpu |
Deprecated alias |
enhanced_hierarchical_cpu |
Deprecated alias |
Docking runs on CPU. The workload parallelises across ligands rather than within a single search, so throughput scales with core count.
benchmarking/ contains a redocking harness. These scripts are development
tooling and are not shipped in the PyPI package — clone the repository to use
them.
# Split whole PDB entries into receptor/ligand pairs
python benchmarking/prepare_complexes.py --input complexes/ --output prepared/
# Check for ligand leakage before spending compute
python benchmarking/validate_prepared.py --manifest prepared/manifest.csv
# Redock and measure
python benchmarking/redock_benchmark.py --manifest prepared/manifest.csv \
--output results/ --padding 5 --seed 42 -j 8
# Publication tables (Markdown and LaTeX)
python benchmarking/make_report.py results/redock_results.csv --output results/reportRMSD is symmetry-corrected and computed without superposition. Top-1 and best-of-N success rates are reported separately: quoting best-of-N as though it were top-1 substantially overstates accuracy.
analyze_redock.py separates search failures from ranking failures. A complex
where a sub-2 Å pose was generated but not ranked first is a scoring problem that
more sampling cannot fix, and it is the quantity that tells you whether a learned
rescorer is worth applying.
These are PandaDock-GNN scoring results. For pose-prediction accuracy see Measuring Pose Accuracy above.
| Metric | Value |
|---|---|
| Pearson R | 0.88 |
| Spearman R | 0.88 |
| RMSE | 0.93 pK units |
| MAE | 0.68 pK units |
| Within 1.0 pK | 77.5% |
| Within 1.5 pK | 90.5% |
ULVSH is primarily an activity classification benchmark. Its affinity labels are heavily censored: inactive compounds are assigned a floor value, and in the 95-compound test split 89 of 95 pEC50 values are exactly 4.0. A Pearson correlation computed over that distribution is dominated by separating actives from the constant floor rather than by ranking affinities, so it is reported here as classification performance.
| Metric | Value | N |
|---|---|---|
| Activity classification AUC | 0.94 | 95 (test) |
| Within-target affinity Pearson R (CASR) | 0.77 | 12 |
CASR is the only test-split target with enough compounds at distinct affinities to support a within-target correlation. For continuous affinity prediction see the PDBbind and BindingDB results above and below, whose labels span a genuine range.
| Training Configuration | Test Pearson R | Test RMSE | N (train) |
|---|---|---|---|
| BindingDB Only | 0.81 | - | 7,113 |
| BindingDB + ULVSH | 0.79 | 0.96 | 7,866 |
| BindingDB + ULVSH + PDBbind | 0.49 | 1.37 | 12,118 |
Note: Combined training with PDBbind shows reduced performance due to affinity scale differences (pKd vs pEC50). For best results, train on datasets with compatible affinity measurements.
Key Results:
- PandaDock-GNN achieves R = 0.88 on PDBbind (5,316 complexes)
- R = 0.81 on BindingDB test set (889 complexes)
- AUC = 0.94 for activity classification on the ULVSH test split
Full documentation available at pandadock.readthedocs.io:
If you use PandaDock in your research, please cite:
@article{panda2024pandadock,
title={PandaDock: SE(3)-Equivariant Graph Neural Network Scoring for Molecular Docking},
author={Panda, Pritam Kumar},
journal={bioRxiv},
year={2024},
note={Manuscript in preparation}
}We welcome contributions! Please see CONTRIBUTING.md for guidelines.
PandaDock is released under the MIT License. See LICENSE for details.
Author: Pritam Kumar Panda Affiliation: Stanford University Email: pritampanda@stanford.edu GitHub: @pritampanda15
PandaDock builds upon excellent open-source projects:
- AutoDock Vina (scoring function inspiration)
- PyTorch and PyTorch Geometric (GNN framework)
- RDKit (molecular handling)
- E(n)-Equivariant GNN (Satorras et al. 2021)
Star this repository if you find it useful!