Delve into Guard Models as Guardrails for LLM Agents

This repository contains the code and data for the evaluation and fine-tuning of models presented in the paper "Delve into Guard Models as Guardrails for LLM Agents". We open-source our resources to facilitate further research in using guard models to enhance the safety and reliability of LLM agents.

Repository Structure

agent_evaluation/ Contains the code and datasets used for evaluating the performance and safety of LLM agents.
guard_model_evaluation/ Includes the code and data for testing and benchmarking the guard models.
finetune/ Provides the codebase for fine-tuning models, along with the synthetic data generated for this process.
result/ Stores the experimental results and logs obtained from our studies.

Getting Started

Please refer to the README files within each subdirectory for specific instructions on how to run the code and reproduce the results.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Delve into Guard Models as Guardrails for LLM Agents

Repository Structure

Getting Started

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Name		Name	Last commit message	Last commit date
Latest commit History 2 Commits
agent_evaluation		agent_evaluation
finetune		finetune
guard_model_evaluation		guard_model_evaluation
result		result
README.md		README.md

Folders and files

Latest commit

History

Repository files navigation

Delve into Guard Models as Guardrails for LLM Agents

Repository Structure

Getting Started

About

Resources

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

Packages