This document provides a summary of the safety evaluation datasets that been used in Libra Leaderboard.
| Dataset Name | Include | Year | Publication Name | Purpose Type | Language | Entries | GitHub | HF Link | Notes |
|---|---|---|---|---|---|---|---|---|---|
| AART | ✅ | 2023 | AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications | broad safety | English | 3,269 prompts | Link | No | - contains examples for specific geographic regions - prompts also change up use cases and concepts |
| AdvBench | ✅ | 2023 | Universal and Transferable Adversarial Attacks on Aligned Language Models | broad safety | English | 1,000 prompts | Link | No | - focus of the work is adversarial / to jailbreak LLMs - AdvBench tests whether jailbreaks succeeded |
| AnthropicHarmlessBase | ✅ | 2024 | Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback | broad safety | English | 44,849 conversational turns | Link | Link | - most prompts created by 28 US-based crowdworkers |
| AnthropicRedTeam | ✅ | 2024 | Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned | broad safety | English | 38,961 conversations | Link | Link | - created by 324 US-based crowdworkers - ca. 80% of examples come from ca. 50 workers |
| BAD | ✅ | 2022 | Bot-Adversarial Dialogue for Safe Conversational Agents | broad safety | English | 78,874 conversations | Link | No | - download only via ParlAI - approximately 40% of all dialogues are annotated as offensive, with a third of offensive utterances generated by bots |
| BBQ | ✅ | 2023 | BBQ: A Hand-Built Bias Benchmark for Question Answering | bias | English | 58,492 examples | Link | Link | - focus on stereotyping behaviour - covers 9 categories of bias: age, disability status, gender identity, nationality, physical appearance, race/ethnicity, religion, socioeconomic status, sexual orientation - 25+ templates per category |
| BeaverTails | Todo | 2022 | BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset | broad safety | English | 333,963 conversations | Link | Link | - 16,851 unique prompts sampled from AnthropicRedTeam - covers 14 harm categories (e.g. animal abuse) - annotated for safety by 3.34 crowdworkers on average |
| CDialBias | ❌ | 2019 | Towards Identifying Social Bias in Dialog Systems: Framework, Dataset, and Benchmark | bias | Chinese | 28,343 conversations | Link | No | - cover 4 categories of bias: race, gender, religion, occupation |
| CoNA | ✅ | 2024 | Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions | narrow safety | English | 178 prompts | Link | No | - Focus: harmful instructions (e.g. hate speech) - All prompts are (meant to be) unsafe |
| ConfAIde | ✅ | 2023 | Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory | narrow safety | English | 1,326 prompts | Link | No | - The benchmark is split into 4 tiers with different prompt formats - tier 1 contains 10 prompts, tier 2 298, tier 3 4270, tier 4 50 |
| ControversialInstructions | ✅ | 2023 | Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions | narrow safety | English | 40 prompts | Link | No | - Focus: controversial topics (e.g. immigration) - All prompts are (meant to be) unsafe |
| CyberattackAssistance | ✅ | 2024 | Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models | narrow safety | English | 1,000 prompts | Link | No | - instructions are split into 10 MITRE categories - The dataset comes with additional LLM-rephrased instructions |
| DecodingTrust | ✅ | 2022 | DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models | broad safety | English | 243,877 prompts | Link | Link | - split across 8 'trustworthiness perspectives': toxicity, stereotypes, adversarial and robustness, privacy, ethics and fairness |
| DICES350 | ✅ | 2023 | DICES Dataset: Diversity in Conversational AI Evaluation for Safety | value alignment | English | 350 conversations | Link | No | - 104 ratings per item - annotators from US - annotation across 24 safety criteria |
| DoAnythingNow | ❌ | 2023 | 'Do Anything Now': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models | narrow safety | English | 6,387 prompts | Link | No | - there are 666 jailbreak prompts among the 6,387prompts |
| DoNotAnswer | ✅ | 2023 | Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs | broad safety | English | 939 prompts | Link | Link | - split across 5 risk areas and 12 harm types - authors prompted GPT-4 to generate questions |
| DoAnythingNow | ✅ | 2022 | 'Do Anything Now': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models | broad safety | English | 46,800 prompts | Link | No | - 13 scenarios with 30 questions each, then expanded by combination with 8 communities, 3 prompt setups and 5 repetations |
| GandalfIgnoreInstructions | ✅ | 2023 | Gandalf Prompt Injection: Ignore Instruction Prompts | other | English | 1,000 prompts | No | Link | - focus of the work is adversarial / prompt extraction - not all prompts are attacks |
| HackAPrompt | Todo | 2023 | Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition | other | mostly English | 601,757 prompts | No | Link | - focus of the work is adversarial / prompt hacking - prompts were written by ca. 2.8k people from 50+ countries |
| HarmBench | ✅ | 2022 | HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal | broad safety | English | 400 prompts | Link | No | - The dataset covers 7 semantic categories of behaviour: Cybercrime & Unauthorized Intrusion, Chemical & Biological Weapons/Drugs, Copyright Violations, Misinformation & Disinformation, Harassment & Bullying, Illegal Activities, and General Harm - The dataset also includes 110 multimodal prompts |
| HarmfulQ | ✅ | 2021 | On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning | broad safety | English | 200 prompts | Link | No | - focus on 6 attributes: "racist, stereotypical, sexist, illegal, toxic, harmful" - authors do manual filtering for overly similar questions |
| RedEval | ✅ | 2023 | Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment | broad safety | English | 1,960 prompts | Link | Link | - split into 10 topics (e.g. "Mathematics and Logic") - similarity across prompts is quite high - not all prompts are unsafe / safety-related |
| HExPHI | ✅ | 2024 | Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! | broad safety | English | 330 prompts | No | Link | - Main focus is on finetuning models - Prompts cover 11 harm areas |
| HolisticBias | ❌ | 2023 | I’m sorry to hear that: Finding New Biases in Language Models with a Holistic Descriptor Dataset | bias | English | 459,758 prompts | Link | No | - 26 sentence templates - covers 13 categories of bias: ability, age, body type, characteristics, culturural, gender/sex, nationality, nonce, political, race/ethnicity, religion, sexual orientation, socioeconomic |
| HypothesisStereotypes | ✅ | 2023 | Analyzing Stereotypes in Generative Text Inference Tasks | bias | English | 2,098 prompts | Link | No | - uses 103 context situations as templates - covers 6 categories of bias: gender, race, nationality, religion, politics, socio - task for LLM is to generate hypothesis based on premise |
| LatentJailbreak | ✅ | 2024 | Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models | narrow safety | English | 416 prompts | Link | No | - focus of the work is adversarial / to jailbreak LLMs - 13 prompt templates instantiated with 16 protected group terms and 2 posititional types - main exploit focuses on translation |
| MaliciousInstruct | ✅ | 2022 | Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation | broad safety | English | 100 prompts | Link | No | - covers ten 'malicious intentions': psychological manipulation, sabotage, theft, defamation, cyberbullying, false accusation, tax fraud, hacking, fraud, and illegal drug use. |
| MaliciousInstructions | ✅ | 2024 | Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions | broad safety | English | 100 prompts | Link | No | - Focus: malicious instructions (e.g. bombmaking) - All prompts are (meant to be) unsafe |
| MIC | ❌ | 2024 | The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems | value alignment | English | 38,000 conversations | Link | No | - based on the RoT paradigm introduced in SocialChemistry - 38k prompt-reply pairs come with 99k rules of thumb and 114k annotations |
| MoralChoice | ✅ | 2023 | Evaluating the Moral Beliefs Encoded in LLMs | value alignment | English | 1,767 binary-choice question | Link | Link | - 687 scenarios are low-ambiguity, 680 are high-ambiguity - three Surge annotators choose the favourable action for each scenario |
| DialogueSafety | ✅ | 2021 | Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack | broad safety | English | 90,000 prompts | Link | No | - download only via ParlAI - the "single-turn" dataset provides a "standard" and "adversarial" setting, with 3 rounds of data collection each |
| PersonalInfoLeak | ✅ | 2024 | Are Large Pre-Trained Language Models Leaking Your Personal Information? | narrow safety | English | 3,238 entries | Link | No | - main task is to predict email given name |
| PhysicalSafetyInstructions | ✅ | 2023 | Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions | narrow safety | English | 1,000 prompts | Link | No | - Focus: commonsense physical safety - 50 safe and 50 unsafe prompts |
| PromptExtractionRobustness | ✅ | 2023 | Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game | narrow safety | English | 569 samples | Link | Link | - filtered from larger raw prompt extraction dataset - collected using the open Tensor Trust online game |
| PromptHijackingRobustness | ✅ | 2021 | Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game | narrow safety | English | 775 samples | Link | Link | - filtered from larger raw prompt extraction dataset - collected using the open Tensor Trust online game |
| QHarm | Todo | 2023 | Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions | broad safety | English | 100 prompts | Link | No | - Wider topic coverage due to source dataset - Prompts are mostly unsafe |
| RedditBias | ❌ | 2024 | RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models | bias | English | 11,873 Reddit comments | Link | No | - covers 4 categories of bias: religion, race, gender, queerness - evaluation by perplexity and conversation |
| RuLES | ❌ | 2024 | Can LLMs Follow Simple Rules? | narrow safety | English | 862 prompts | Link | Link | - The dataset covers 19 rules across 14 scenarios |
| SafeText | ❌ | 2022 | SafeText: A Benchmark for Exploring Physical Safety in Language Models | narrow safety | English | 367 prompts | Link | No | - 5 ratings for relevance per item during filtering - advice format most often elicits yes/no answer |
| SafetyPrompts | ❌ | 2023 | Safety Assessment of Chinese Large Language Models | broad safety | Chinese | 100,000 prompts | Link | Link | - covers 8 safety scenarios and 6 types of adv attack - do not release 'sensitive topics' scenario |
| SimpleSafetyTests | ✅ | 2024 | SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models | broad safety | English | 100 prompts | Link | Link | - The dataset is split into ten types of prompts |
| SPMisconceptions | ✅ | 2022 | Can Large Language Models Provide Security & Privacy Advice? Measuring the Ability of LLMs to Refute Misconceptions | narrow safety | English | 122 prompts | Link | No | - misconceptions all relate to security and privacy - uses templates to turn misconceptions into prompts - covers six categories (e.g. crypto and blockchain, law and regulation) |
| StrongREJECT | ✅ | 2023 | A StrongREJECT for Empty Jailbreaks | broad safety | English | 346 prompts | Link | No | - The focus of the work is adversarial / to jailbreak LLMs - The 6 question categories are: disinformation/deception, hate/harassment/discrimination, illegal goods/services, non-violent crimes, sexual content, violence |
| SycophancyEval | ✅ | 2024 | Towards Understanding Sycophancy in Language Models | narrow safety | English | 20,956 prompts | Link | Link | - Dataset uses four different task setups to evaluate sycophancy: answer (7268 prompts), are_you_sure (4888 prompts), feedback (8500 prompts), mimicry (300 prompts) |
| TDCRedTeaming | ✅ | 2022 | TDC 2023 (LLM Edition): The Trojan Detection Challenge | broad safety | English | 100 prompts | Link | No | - covers 7 categories: Bigotry and Abusive Language, Violent Content and Conduct, Illegal Activities, Malware and Exploits, Scams, Misinformation and Disinformation, Other Undesirable Content |
| ToxiGen | ✅ | 2023 | ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection | narrow safety | English | 260,851 prompts | Link | Link | - covers 13 target groups - seed prompts are used to generate implicit hate - evaluating generative LLMs is not the focus |
| TruthfulQA | Todo | 2022 | TruthfulQA: Measuring How Models Mimic Human Falsehoods | narrow safety | English | 817 prompts | Link | Link | - covers 38 categories (e.g. health and politics) - comes with multiple choice expansion |
| XSTest | ✅ | 2023 | XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models | narrow safety | English | 450 prompts | Link | Link | - split into ten types of prompts - 250 safe prompts and 200 unsafe prompts |
- Year: 2023
- Publication Name: AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 3,269 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/google-research-datasets/aart-ai-safety-dataset
- HF Link: [not available](not available)
- Notes:
- contains examples for specific geographic regions
- prompts also change up use cases and concepts
- Year: 2023
- Publication Name: Universal and Transferable Adversarial Attacks on Aligned Language Models
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 1,000 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/llm-attacks/llm-attacks/tree/main/data/advbench
- HF Link: [not available](not available)
- Notes:
- focus of the work is adversarial / to jailbreak LLMs
- AdvBench tests whether jailbreaks succeeded
- Year: 2024
- Publication Name: Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Purpose: broad safety
- LLMDev Usage: train and eval
- Entries: 44,849 conversational turns
- Entries Type: chat
- Language: English
- GitHub: https://github.com/anthropics/hh-rlhf
- HF Link: https://huggingface.co/datasets/Trelis/hh-rlhf-dpo
- Notes:
- most prompts created by 28 US-based crowdworkers
- Year: 2024
- Publication Name: Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Purpose: broad safety
- LLMDev Usage: train and eval
- Entries: 38,961 conversations
- Entries Type: chat
- Language: English
- GitHub: https://github.com/anthropics/hh-rlhf
- HF Link: https://huggingface.co/datasets/Anthropic/hh-rlhf
- Notes:
- created by 324 US-based crowdworkers
- ca. 80% of examples come from ca. 50 workers
- Year: 2022
- Publication Name: Bot-Adversarial Dialogue for Safe Conversational Agents
- Purpose: broad safety
- LLMDev Usage: train and eval
- Entries: 78,874 conversations
- Entries Type: chat
- Language: English
- GitHub: https://github.com/facebookresearch/ParlAI/tree/main/parlai/tasks/bot_adversarial_dialogue
- HF Link: [not available](not available)
- Notes:
- download only via ParlAI
- approximately 40% of all dialogues are annotated as offensive, with a third of offensive utterances generated by bots
- Year: 2023
- Publication Name: BBQ: A Hand-Built Bias Benchmark for Question Answering
- Purpose: bias
- LLMDev Usage: eval only
- Entries: 58,492 examples
- Entries Type: chat
- Language: English
- GitHub: https://github.com/nyu-mll/BBQ
- HF Link: https://huggingface.co/datasets/heegyu/bbq
- Notes:
- focus on stereotyping behaviour
- covers 9 categories of bias: age, disability status, gender identity, nationality, physical appearance, race/ethnicity, religion, socioeconomic status, sexual orientation
- 25+ templates per category
- Year: 2022
- Publication Name: BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
- Purpose: broad safety
- LLMDev Usage: train and eval
- Entries: 333,963 conversations
- Entries Type: chat
- Language: English
- GitHub: https://github.com/PKU-Alignment/beavertails
- HF Link: https://huggingface.co/datasets/PKU-Alignment/BeaverTails
- Notes:
- 16,851 unique prompts sampled from AnthropicRedTeam
- covers 14 harm categories (e.g. animal abuse)
- annotated for safety by 3.34 crowdworkers on average
- Year: 2019
- Publication Name: Towards Identifying Social Bias in Dialog Systems: Framework, Dataset, and Benchmark
- Purpose: bias
- LLMDev Usage: eval only
- Entries: 28,343 conversations
- Entries Type: chat
- Language: Chinese
- GitHub: https://github.com/para-zhou/CDial-Bias
- HF Link: [not available](not available)
- Notes:
- cover 4 categories of bias: race, gender, religion, occupation
- Year: 2024
- Publication Name: Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 178 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/vinid/instruction-llms-safety-eval
- HF Link: [not available](not available)
- Notes:
- Focus: harmful instructions (e.g. hate speech)
- All prompts are (meant to be) unsafe
- Year: 2023
- Publication Name: Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 1,326 prompts
- Entries Type: other
- Language: English
- GitHub: https://github.com/skywalker023/confAIde/tree/main/benchmark
- HF Link: [not available](not available)
- Notes:
- The benchmark is split into 4 tiers with different prompt formats
- tier 1 contains 10 prompts, tier 2 298, tier 3 4270, tier 4 50
- Year: 2023
- Publication Name: Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 40 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/vinid/instruction-llms-safety-eval
- HF Link: [not available](not available)
- Notes:
- Focus: controversial topics (e.g. immigration)
- All prompts are (meant to be) unsafe
- Year: 2024
- Publication Name: Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 1,000 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/facebookresearch/PurpleLlama/tree/main/CybersecurityBenchmarks/datasets/mitre
- HF Link: [not available](not available)
- Notes:
- instructions are split into 10 MITRE categories
- The dataset comes with additional LLM-rephrased instructions
- Year: 2022
- Publication Name: DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 243,877 prompts
- Entries Type: multiple choice
- Language: English
- GitHub: https://github.com/AI-secure/DecodingTrust
- HF Link: https://huggingface.co/datasets/AI-Secure/DecodingTrust
- Notes:
- split across 8 'trustworthiness perspectives': toxicity, stereotypes, adversarial and robustness, privacy, ethics and fairness
- Year: 2023
- Publication Name: DICES Dataset: Diversity in Conversational AI Evaluation for Safety
- Purpose: value alignment
- LLMDev Usage: eval only
- Entries: 350 conversations
- Entries Type: chat
- Language: English
- GitHub: https://github.com/google-research-datasets/dices-dataset/
- HF Link: [not available](not available)
- Notes:
- 104 ratings per item
- annotators from US
- annotation across 24 safety criteria
- Year: 2023
- Publication Name: 'Do Anything Now': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 6,387 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/verazuo/jailbreak_llms
- HF Link: [not available](not available)
- Notes:
- there are 666 jailbreak prompts among the 6,387prompts
- Year: 2023
- Publication Name: Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 939 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/Libr-AI/do-not-answer
- HF Link: https://huggingface.co/datasets/LibrAI/do-not-answer
- Notes:
- split across 5 risk areas and 12 harm types
- authors prompted GPT-4 to generate questions
- Year: 2022
- Publication Name: 'Do Anything Now': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 46,800 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/verazuo/jailbreak_llms
- HF Link: [not available](not available)
- Notes:
- 13 scenarios with 30 questions each, then expanded by combination with 8 communities, 3 prompt setups and 5 repetations
- Year: 2023
- Publication Name: Gandalf Prompt Injection: Ignore Instruction Prompts
- Purpose: other
- LLMDev Usage: other
- Entries: 1,000 prompts
- Entries Type: chat
- Language: English
- GitHub: [not available](not available)
- HF Link: https://huggingface.co/datasets/Lakera/gandalf_ignore_instructions
- Notes:
- focus of the work is adversarial / prompt extraction
- not all prompts are attacks
- Year: 2023
- Publication Name: Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition
- Purpose: other
- LLMDev Usage: other
- Entries: 601,757 prompts
- Entries Type: chat
- Language: mostly English
- GitHub: [not available](not available)
- HF Link: https://huggingface.co/datasets/hackaprompt/hackaprompt-dataset
- Notes:
- focus of the work is adversarial / prompt hacking
- prompts were written by ca. 2.8k people from 50+ countries
- Year: 2022
- Publication Name: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 400 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/centerforaisafety/HarmBench/tree/main/data/behavior_datasets
- HF Link: [not available](not available)
- Notes:
- The dataset covers 7 semantic categories of behaviour: Cybercrime & Unauthorized Intrusion, Chemical & Biological Weapons/Drugs, Copyright Violations, Misinformation & Disinformation, Harassment & Bullying, Illegal Activities, and General Harm
- The dataset also includes 110 multimodal prompts
- Year: 2021
- Publication Name: On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 200 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/SALT-NLP/chain-of-thought-bias
- HF Link: [not available](not available)
- Notes:
- focus on 6 attributes: "racist, stereotypical, sexist, illegal, toxic, harmful"
- authors do manual filtering for overly similar questions
- Year: 2023
- Publication Name: Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 1,960 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/declare-lab/red-instruct/tree/main/harmful_questions
- HF Link: https://huggingface.co/datasets/declare-lab/HarmfulQA
- Notes:
- split into 10 topics (e.g. "Mathematics and Logic")
- similarity across prompts is quite high
- not all prompts are unsafe / safety-related
- Year: 2024
- Publication Name: Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 330 prompts
- Entries Type: chat
- Language: English
- GitHub: [not available](not available)
- HF Link: https://huggingface.co/datasets/LLM-Tuning-Safety/HEx-PHI
- Notes:
- Main focus is on finetuning models
- Prompts cover 11 harm areas
- Year: 2023
- Publication Name: I’m sorry to hear that: Finding New Biases in Language Models with a Holistic Descriptor Dataset
- Purpose: bias
- LLMDev Usage: eval only
- Entries: 459,758 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/facebookresearch/ResponsibleNLP/tree/main/holistic_bias
- HF Link: [not available](not available)
- Notes:
- 26 sentence templates
- covers 13 categories of bias: ability, age, body type, characteristics, culturural, gender/sex, nationality, nonce, political, race/ethnicity, religion, sexual orientation, socioeconomic
- Year: 2023
- Publication Name: Analyzing Stereotypes in Generative Text Inference Tasks
- Purpose: bias
- LLMDev Usage: eval only
- Entries: 2,098 prompts
- Entries Type: other
- Language: English
- GitHub: https://github.com/AnnaSou/stereotypes_generative_inferences
- HF Link: [not available](not available)
- Notes:
- uses 103 context situations as templates
- covers 6 categories of bias: gender, race, nationality, religion, politics, socio
- task for LLM is to generate hypothesis based on premise
- Year: 2024
- Publication Name: Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 416 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/qiuhuachuan/latent-jailbreak/tree/main
- HF Link: [not available](not available)
- Notes:
- focus of the work is adversarial / to jailbreak LLMs
- 13 prompt templates instantiated with 16 protected group terms and 2 posititional types
- main exploit focuses on translation
- Year: 2022
- Publication Name: Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 100 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/Princeton-SysML/Jailbreak_LLM/tree/main/data
- HF Link: [not available](not available)
- Notes:
- covers ten 'malicious intentions': psychological manipulation, sabotage, theft, defamation, cyberbullying, false accusation, tax fraud, hacking, fraud, and illegal drug use.
- Year: 2024
- Publication Name: Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 100 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/vinid/instruction-llms-safety-eval
- HF Link: [not available](not available)
- Notes:
- Focus: malicious instructions (e.g. bombmaking)
- All prompts are (meant to be) unsafe
- Year: 2024
- Publication Name: The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems
- Purpose: value alignment
- LLMDev Usage: eval only
- Entries: 38,000 conversations
- Entries Type: chat
- Language: English
- GitHub: https://github.com/SALT-NLP/mic
- HF Link: [not available](not available)
- Notes:
- based on the RoT paradigm introduced in SocialChemistry
- 38k prompt-reply pairs come with 99k rules of thumb and 114k annotations
- Year: 2023
- Publication Name: Evaluating the Moral Beliefs Encoded in LLMs
- Purpose: value alignment
- LLMDev Usage: eval only
- Entries: 1,767 binary-choice question
- Entries Type: multiple choice
- Language: English
- GitHub: https://github.com/ninodimontalcino/moralchoice
- HF Link: https://huggingface.co/datasets/ninoscherrer/moralchoice
- Notes:
- 687 scenarios are low-ambiguity, 680 are high-ambiguity
- three Surge annotators choose the favourable action for each scenario
- Year: 2021
- Publication Name: Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack
- Purpose: broad safety
- LLMDev Usage: train and eval
- Entries: 90,000 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/facebookresearch/ParlAI/tree/main/parlai/tasks/dialogue_safety
- HF Link: [not available](not available)
- Notes:
- download only via ParlAI
- the "single-turn" dataset provides a "standard" and "adversarial" setting, with 3 rounds of data collection each
- Year: 2024
- Publication Name: Are Large Pre-Trained Language Models Leaking Your Personal Information?
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 3,238 entries
- Entries Type: other
- Language: English
- GitHub: https://github.com/jeffhj/LM_PersonalInfoLeak
- HF Link: [not available](not available)
- Notes:
- main task is to predict email given name
- Year: 2023
- Publication Name: Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 1,000 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/vinid/instruction-llms-safety-eval
- HF Link: [not available](not available)
- Notes:
- Focus: commonsense physical safety
- 50 safe and 50 unsafe prompts
- Year: 2023
- Publication Name: Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 569 samples
- Entries Type: chat
- Language: English
- GitHub: https://github.com/HumanCompatibleAI/tensor-trust-data
- HF Link: https://huggingface.co/datasets/qxcv/tensor-trust?row=0
- Notes:
- filtered from larger raw prompt extraction dataset
- collected using the open Tensor Trust online game
- Year: 2021
- Publication Name: Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 775 samples
- Entries Type: chat
- Language: English
- GitHub: https://github.com/HumanCompatibleAI/tensor-trust-data
- HF Link: https://huggingface.co/datasets/qxcv/tensor-trust?row=0
- Notes:
- filtered from larger raw prompt extraction dataset
- collected using the open Tensor Trust online game
- Year: 2023
- Publication Name: Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 100 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/vinid/instruction-llms-safety-eval
- HF Link: [not available](not available)
- Notes:
- Wider topic coverage due to source dataset
- Prompts are mostly unsafe
- Year: 2024
- Publication Name: RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models
- Purpose: bias
- LLMDev Usage: eval only
- Entries: 11,873 Reddit comments
- Entries Type: other
- Language: English
- GitHub: https://github.com/umanlp/RedditBias
- HF Link: [not available](not available)
- Notes:
- covers 4 categories of bias: religion, race, gender, queerness
- evaluation by perplexity and conversation
- Year: 2024
- Publication Name: Can LLMs Follow Simple Rules?
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 862 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/normster/llm_rules
- HF Link: https://huggingface.co/datasets/normster/RuLES
- Notes:
- The dataset covers 19 rules across 14 scenarios
- Year: 2022
- Publication Name: SafeText: A Benchmark for Exploring Physical Safety in Language Models
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 367 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/sharonlevy/SafeText
- HF Link: [not available](not available)
- Notes:
- 5 ratings for relevance per item during filtering
- advice format most often elicits yes/no answer
- Year: 2023
- Publication Name: Safety Assessment of Chinese Large Language Models
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 100,000 prompts
- Entries Type: chat
- Language: Chinese
- GitHub: https://github.com/thu-coai/Safety-Prompts
- HF Link: https://huggingface.co/datasets/thu-coai/Safety-Prompts
- Notes:
- covers 8 safety scenarios and 6 types of adv attack
- do not release 'sensitive topics' scenario
- Year: 2024
- Publication Name: SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 100 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/bertiev/SimpleSafetyTests
- HF Link: https://huggingface.co/datasets/Bertievidgen/SimpleSafetyTests
- Notes:
- The dataset is split into ten types of prompts
- Year: 2022
- Publication Name: Can Large Language Models Provide Security & Privacy Advice? Measuring the Ability of LLMs to Refute Misconceptions
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 122 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/purseclab/LLM_Security_Privacy_Advice
- HF Link: [not available](not available)
- Notes:
- misconceptions all relate to security and privacy
- uses templates to turn misconceptions into prompts
- covers six categories (e.g. crypto and blockchain, law and regulation)
- Year: 2023
- Publication Name: A StrongREJECT for Empty Jailbreaks
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 346 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/alexandrasouly/strongreject/tree/main
- HF Link: [not available](not available)
- Notes:
- The focus of the work is adversarial / to jailbreak LLMs
- The 6 question categories are: disinformation/deception, hate/harassment/discrimination, illegal goods/services, non-violent crimes, sexual content, violence
- Year: 2024
- Publication Name: Towards Understanding Sycophancy in Language Models
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 20,956 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/meg-tong/sycophancy-eval/tree/main
- HF Link: https://huggingface.co/datasets/meg-tong/sycophancy-eval
- Notes:
- Dataset uses four different task setups to evaluate sycophancy: answer (7268 prompts), are_you_sure (4888 prompts), feedback (8500 prompts), mimicry (300 prompts)
- Year: 2022
- Publication Name: TDC 2023 (LLM Edition): The Trojan Detection Challenge
- Purpose: broad safety
- LLMDev Usage: eval only
- Entries: 100 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/centerforaisafety/tdc2023-starter-kit/tree/main/red_teaming/data
- HF Link: [not available](not available)
- Notes:
- covers 7 categories: Bigotry and Abusive Language, Violent Content and Conduct, Illegal Activities, Malware and Exploits, Scams, Misinformation and Disinformation, Other Undesirable Content
- Year: 2023
- Publication Name: ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 260,851 prompts
- Entries Type: other
- Language: English
- GitHub: https://github.com/microsoft/TOXIGEN
- HF Link: https://huggingface.co/datasets/skg/toxigen-data
- Notes:
- covers 13 target groups
- seed prompts are used to generate implicit hate
- evaluating generative LLMs is not the focus
- Year: 2022
- Publication Name: TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 817 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/sylinrl/TruthfulQA
- HF Link: https://huggingface.co/datasets/truthful_qa
- Notes:
- covers 38 categories (e.g. health and politics)
- comes with multiple choice expansion
- Year: 2023
- Publication Name: XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
- Purpose: narrow safety
- LLMDev Usage: eval only
- Entries: 450 prompts
- Entries Type: chat
- Language: English
- GitHub: https://github.com/paul-rottger/exaggerated-safety
- HF Link: https://huggingface.co/datasets/natolambert/xstest-v2-copy
- Notes:
- split into ten types of prompts
- 250 safe prompts and 200 unsafe prompts