This repository contains the official codebase for the paper [Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning]
Please first clone the repository and install the required environment by following the steps below:
conda create -n satenv python=3.8
conda activate satenv
# Torch with CUDA 12.1
conda install pytorch torchvision torchaudio pytorch-cuda=12.1 -c pytorch -c nvidia
# Install PyG packages
conda install pyg -c pyg
# Install required libraries
pip install -r requirements.txtSAT consists of two main stages: (1) hierarchical knowledge alignment and (2) structural instruction tuning.
Download the original datasets: FB15k-237N and CoDeX-S from here FB15k-237 and YAGO3-10 from here
Construct query subgraphs using the code available here.
Extract entity descriptions from this tool.
Note that we have provided pre-processed data (e.g., FB15k-237N) within the repository.
To train the aligner model, run the following command. The trained model will be saved to ./checkpoints/{data_name}/gt-xxx.pkl.
cd aligner
python3 model/main.py --data_name FB15k-237NPlace the pre-trained graph encoder model in the specified directory, and obtain the graph data containing node features, edge indices, and related information.
Prepare the base model (Llama2) by downloading its weights from this Hugging Face page.
Prepare the instruction tuning data. We provide the pre-processed data in the data_llm_lp directory.
To tune the predictor model and evaluate its performance, run the following commands:
cd predictor
bash ./scripts/run_llm_lp.sh
bash ./scripts/run_llm_lp_eval.sh