Replication Package for "TrajDeleter: Enabling Trajectory Forgetting in Offline Reinforcement Learning Agents", NDSS 2025.
Please check our agents' parameters in this anonymous link:
Please download the model from this link, and please move these folders to the ``unlearning" folder.
The descriptions of folds are as follows:
| fold_name | descriptions |
|---|---|
| Shadow Models | Shadow agents for auditing |
| Exact unlearning agents | Retraining agents from scratch |
| Fully trained agents | Agents Trained using the entire dataset |
| algorithm | discrete control | continuous control |
|---|---|---|
| Batch Constrained Q-learning (BCQ) | ✓ | ✓ |
| Bootstrapping Error Accumulation Reduction (BEAR) | x | ✓ |
| Conservative Q-Learning (CQL) | ✓ | ✓ |
| Implicit Q-learning (IQL) | ✓ | ✓ |
| Policy in the Latent Action Space with Perturbation (PLAS-P) | ✓ | ✓ |
| Twin Delayed Deep Deterministic Policy Gradient plus Behavioral Cloning (TD3PlusBC) | ✓ | ✓ |
The structure of this project in `unlearning' folder is as follows:
Unlearning
-- env_test.py ------------------ download and test the dataset
-- efficacy_evaluation.py ------------------ the codes of trajauditor for RQ1.
-- mujoco_exact_unlearning.py ------------------ unlearning with retraining the agents from scratch.
-- mujoco_finetuning.py ------------------ unlearning with fine-tuning method.
-- mujoco_fully_training.py ------------------ train the original agents using the whole dataset.
-- mujoco_random_reward.py ------------------ unlearning with random_reward method.
-- mujoco_auditor.py ------------------ the codes of trajauditor for RQ2.
-- mujoco_trajdeleter.py ------------------ unlearning using our proposed method.
-- performance_test.py ------------------ evaluate the average cumulative rewards obtained by agents.
-- poisoning_retrain.py ------------------ unlearning the poisoned trajectories using retraining method.
-- poisoning_training.py ------------------ training agents using the poisoned dataset.
-- script-aditor.py ------------------ script files of trajauditor and achieve the results in RQ1.
-- script-exact-unlearning.py ------------------ script files of retraining.
-- script-mujoco-fine-tune.py ------------------ script files of fine tuning.
-- script-mujoco-random-reward.py ------------------ script files of random reward method.
-- script-mujoco-test.py ------------------ script files of agents' performance testing.
-- script-mujoco-trajdeleter.py ------------------ script files of our proposed unlearning method.
-- script-performance-test.py ------------------ script files of testing the cumulative returns of agents.
-- script-poisoning-fine-tuning.py ------------------ script files of fine-tuning the poisoned dataset.
-- script-poisoning-retrain.py ------------------ script files of retraining the poisoned dataset.
-- script-poisoning.py ------------------ script files of poisoning agents.
-- script-trajauditor.py ------------------ script files of trajauditor.
packages.txt ------------------ the list of our env.
params ------------------ the files of hype-parameters settings of offline reinforcement learning algorithms.
This code was developed with python 3.7.11.
The version of Mujoco is Mujoco 2.1.0.
Our CUDA version is 12.4.
pip install d3rlpy==1.0.0
pip install mujoco-py==2.1.2.14
pip install gym==0.22.0
pip install scikit-learn==1.0.2
pip install Cython==0.29.36
mkdir ~/.mujoco
Download mujoco210-linux-x86_64.tar.gz from Mujoco 2.1.0, and unzip it to ~/.mujoco. Then, you can find the folder ~/.mujoco/mujoco210.
Please move the lib folder in our repository to the ~/.mujoco/ folder.
Then replace the d3rlpy folder in the your path of environment with the given d3rlpy folder, which can be downloaded here D3RLPY or extract it from the d3rlpy.zip file. For example, the path can be /anaconda3/envs/<the-name-of-environment>/lib/python3.7/site-packages/.
pip install dm_control==0.0.425341097
git clone https://github.com/aravindr93/mjrl.git
cd mjrl
pip install -e .pip install patchelf
git clone https://github.com/rail-berkeley/d4rl.git
cd d4rl
git checkout 71a9549f2091accff93eeff68f1f3ab2c0e0a288 from distutils.core import setup
from platform import platform
from setuptools import find_packages
setup(
name='d4rl',
version='1.1',
install_requires=['gym',
'numpy',
'mujoco_py',
'pybullet',
'h5py',
'termcolor',
'click'],
packages=find_packages(),
package_data={'d4rl': ['locomotion/assets/*',
'hand_manipulation_suite/assets/*',
'hand_manipulation_suite/Adroit/*',
'hand_manipulation_suite/Adroit/gallery/*',
'hand_manipulation_suite/Adroit/resources/*',
'hand_manipulation_suite/Adroit/resources/meshes/*',
'hand_manipulation_suite/Adroit/resources/textures/*',
]},
include_package_data=True,
)
Then:
pip install -e .
You can also pull our image from Docker Hub:
docker pull liuzzyg/trajdeleter:latest
Then enter the image:
docker run --gpus all -it liuzzyg/trajdeleter:latest /bin/bash
README.md under the unlearning folder.