Skip to content

Latest commit

 

History

History
684 lines (631 loc) · 51 KB

File metadata and controls

684 lines (631 loc) · 51 KB

Safety Datasets used in Libra Leaderboard

This document provides a summary of the safety evaluation datasets that been used in Libra Leaderboard.

Overview

Dataset Name Include Year Publication Name Purpose Type Language Entries GitHub HF Link Notes
AART 2023 AART: AI-Assisted Red-Teaming with Diverse Data Generation for New LLM-powered Applications broad safety English 3,269 prompts Link No
- contains examples for specific geographic regions
- prompts also change up use cases and concepts
AdvBench 2023 Universal and Transferable Adversarial Attacks on Aligned Language Models broad safety English 1,000 prompts Link No
- focus of the work is adversarial / to jailbreak LLMs
- AdvBench tests whether jailbreaks succeeded
AnthropicHarmlessBase 2024 Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback broad safety English 44,849 conversational turns Link Link
- most prompts created by 28 US-based crowdworkers
AnthropicRedTeam 2024 Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned broad safety English 38,961 conversations Link Link
- created by 324 US-based crowdworkers
- ca. 80% of examples come from ca. 50 workers
BAD 2022 Bot-Adversarial Dialogue for Safe Conversational Agents broad safety English 78,874 conversations Link No
- download only via ParlAI
- approximately 40% of all dialogues are annotated as offensive, with a third of offensive utterances generated by bots
BBQ 2023 BBQ: A Hand-Built Bias Benchmark for Question Answering bias English 58,492 examples Link Link
- focus on stereotyping behaviour
- covers 9 categories of bias: age, disability status, gender identity, nationality, physical appearance, race/ethnicity, religion, socioeconomic status, sexual orientation
- 25+ templates per category
BeaverTails Todo 2022 BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset broad safety English 333,963 conversations Link Link
- 16,851 unique prompts sampled from AnthropicRedTeam
- covers 14 harm categories (e.g. animal abuse)
- annotated for safety by 3.34 crowdworkers on average
CDialBias 2019 Towards Identifying Social Bias in Dialog Systems: Framework, Dataset, and Benchmark bias Chinese 28,343 conversations Link No
- cover 4 categories of bias: race, gender, religion, occupation
CoNA 2024 Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions narrow safety English 178 prompts Link No
- Focus: harmful instructions (e.g. hate speech)
- All prompts are (meant to be) unsafe
ConfAIde 2023 Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory narrow safety English 1,326 prompts Link No
- The benchmark is split into 4 tiers with different prompt formats
- tier 1 contains 10 prompts, tier 2 298, tier 3 4270, tier 4 50
ControversialInstructions 2023 Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions narrow safety English 40 prompts Link No
- Focus: controversial topics (e.g. immigration)
- All prompts are (meant to be) unsafe
CyberattackAssistance 2024 Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models narrow safety English 1,000 prompts Link No
- instructions are split into 10 MITRE categories
- The dataset comes with additional LLM-rephrased instructions
DecodingTrust 2022 DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models broad safety English 243,877 prompts Link Link
- split across 8 'trustworthiness perspectives': toxicity, stereotypes, adversarial and robustness, privacy, ethics and fairness
DICES350 2023 DICES Dataset: Diversity in Conversational AI Evaluation for Safety value alignment English 350 conversations Link No
- 104 ratings per item
- annotators from US
- annotation across 24 safety criteria
DoAnythingNow 2023 'Do Anything Now': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models narrow safety English 6,387 prompts Link No
- there are 666 jailbreak prompts among the 6,387prompts
DoNotAnswer 2023 Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs broad safety English 939 prompts Link Link
- split across 5 risk areas and 12 harm types
- authors prompted GPT-4 to generate questions
DoAnythingNow 2022 'Do Anything Now': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models broad safety English 46,800 prompts Link No
- 13 scenarios with 30 questions each, then expanded by combination with 8 communities, 3 prompt setups and 5 repetations
GandalfIgnoreInstructions 2023 Gandalf Prompt Injection: Ignore Instruction Prompts other English 1,000 prompts No Link
- focus of the work is adversarial / prompt extraction
- not all prompts are attacks
HackAPrompt Todo 2023 Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition other mostly English 601,757 prompts No Link
- focus of the work is adversarial / prompt hacking
- prompts were written by ca. 2.8k people from 50+ countries
HarmBench 2022 HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal broad safety English 400 prompts Link No
- The dataset covers 7 semantic categories of behaviour: Cybercrime & Unauthorized Intrusion, Chemical & Biological Weapons/Drugs, Copyright Violations, Misinformation & Disinformation, Harassment & Bullying, Illegal Activities, and General Harm
- The dataset also includes 110 multimodal prompts
HarmfulQ 2021 On Second Thought, Let’s Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning broad safety English 200 prompts Link No
- focus on 6 attributes: "racist, stereotypical, sexist, illegal, toxic, harmful"
- authors do manual filtering for overly similar questions
RedEval 2023 Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment broad safety English 1,960 prompts Link Link
- split into 10 topics (e.g. "Mathematics and Logic")
- similarity across prompts is quite high
- not all prompts are unsafe / safety-related
HExPHI 2024 Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! broad safety English 330 prompts No Link
- Main focus is on finetuning models
- Prompts cover 11 harm areas
HolisticBias 2023 I’m sorry to hear that: Finding New Biases in Language Models with a Holistic Descriptor Dataset bias English 459,758 prompts Link No
- 26 sentence templates
- covers 13 categories of bias: ability, age, body type, characteristics, culturural, gender/sex, nationality, nonce, political, race/ethnicity, religion, sexual orientation, socioeconomic
HypothesisStereotypes 2023 Analyzing Stereotypes in Generative Text Inference Tasks bias English 2,098 prompts Link No
- uses 103 context situations as templates
- covers 6 categories of bias: gender, race, nationality, religion, politics, socio
- task for LLM is to generate hypothesis based on premise
LatentJailbreak 2024 Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models narrow safety English 416 prompts Link No
- focus of the work is adversarial / to jailbreak LLMs
- 13 prompt templates instantiated with 16 protected group terms and 2 posititional types
- main exploit focuses on translation
MaliciousInstruct 2022 Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation broad safety English 100 prompts Link No
- covers ten 'malicious intentions': psychological manipulation, sabotage, theft, defamation, cyberbullying, false accusation, tax fraud, hacking, fraud, and illegal drug use.
MaliciousInstructions 2024 Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions broad safety English 100 prompts Link No
- Focus: malicious instructions (e.g. bombmaking)
- All prompts are (meant to be) unsafe
MIC 2024 The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems value alignment English 38,000 conversations Link No
- based on the RoT paradigm introduced in SocialChemistry
- 38k prompt-reply pairs come with 99k rules of thumb and 114k annotations
MoralChoice 2023 Evaluating the Moral Beliefs Encoded in LLMs value alignment English 1,767 binary-choice question Link Link
- 687 scenarios are low-ambiguity, 680 are high-ambiguity
- three Surge annotators choose the favourable action for each scenario
DialogueSafety 2021 Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack broad safety English 90,000 prompts Link No
- download only via ParlAI
- the "single-turn" dataset provides a "standard" and "adversarial" setting, with 3 rounds of data collection each
PersonalInfoLeak 2024 Are Large Pre-Trained Language Models Leaking Your Personal Information? narrow safety English 3,238 entries Link No
- main task is to predict email given name
PhysicalSafetyInstructions 2023 Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions narrow safety English 1,000 prompts Link No
- Focus: commonsense physical safety
- 50 safe and 50 unsafe prompts
PromptExtractionRobustness 2023 Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game narrow safety English 569 samples Link Link
- filtered from larger raw prompt extraction dataset
- collected using the open Tensor Trust online game
PromptHijackingRobustness 2021 Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game narrow safety English 775 samples Link Link
- filtered from larger raw prompt extraction dataset
- collected using the open Tensor Trust online game
QHarm Todo 2023 Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions broad safety English 100 prompts Link No
- Wider topic coverage due to source dataset
- Prompts are mostly unsafe
RedditBias 2024 RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models bias English 11,873 Reddit comments Link No
- covers 4 categories of bias: religion, race, gender, queerness
- evaluation by perplexity and conversation
RuLES 2024 Can LLMs Follow Simple Rules? narrow safety English 862 prompts Link Link
- The dataset covers 19 rules across 14 scenarios
SafeText 2022 SafeText: A Benchmark for Exploring Physical Safety in Language Models narrow safety English 367 prompts Link No
- 5 ratings for relevance per item during filtering
- advice format most often elicits yes/no answer
SafetyPrompts 2023 Safety Assessment of Chinese Large Language Models broad safety Chinese 100,000 prompts Link Link
- covers 8 safety scenarios and 6 types of adv attack
- do not release 'sensitive topics' scenario
SimpleSafetyTests 2024 SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models broad safety English 100 prompts Link Link
- The dataset is split into ten types of prompts
SPMisconceptions 2022 Can Large Language Models Provide Security & Privacy Advice? Measuring the Ability of LLMs to Refute Misconceptions narrow safety English 122 prompts Link No
- misconceptions all relate to security and privacy
- uses templates to turn misconceptions into prompts
- covers six categories (e.g. crypto and blockchain, law and regulation)
StrongREJECT 2023 A StrongREJECT for Empty Jailbreaks broad safety English 346 prompts Link No
- The focus of the work is adversarial / to jailbreak LLMs
- The 6 question categories are: disinformation/deception, hate/harassment/discrimination, illegal goods/services, non-violent crimes, sexual content, violence
SycophancyEval 2024 Towards Understanding Sycophancy in Language Models narrow safety English 20,956 prompts Link Link
- Dataset uses four different task setups to evaluate sycophancy: answer (7268 prompts), are_you_sure (4888 prompts), feedback (8500 prompts), mimicry (300 prompts)
TDCRedTeaming 2022 TDC 2023 (LLM Edition): The Trojan Detection Challenge broad safety English 100 prompts Link No
- covers 7 categories: Bigotry and Abusive Language, Violent Content and Conduct, Illegal Activities, Malware and Exploits, Scams, Misinformation and Disinformation, Other Undesirable Content
ToxiGen 2023 ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection narrow safety English 260,851 prompts Link Link
- covers 13 target groups
- seed prompts are used to generate implicit hate
- evaluating generative LLMs is not the focus
TruthfulQA Todo 2022 TruthfulQA: Measuring How Models Mimic Human Falsehoods narrow safety English 817 prompts Link Link
- covers 38 categories (e.g. health and politics)
- comes with multiple choice expansion
XSTest 2023 XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models narrow safety English 450 prompts Link Link
- split into ten types of prompts
- 250 safe prompts and 200 unsafe prompts

Datasets in Detail

AART


AdvBench


AnthropicHarmlessBase


AnthropicRedTeam


BAD


BBQ


BeaverTails


CDialBias


CoNA


ConfAIde


ControversialInstructions


CyberattackAssistance


DecodingTrust


DICES350


DoAnythingNow


DoNotAnswer


DoAnythingNow


GandalfIgnoreInstructions


HackAPrompt


HarmBench


HarmfulQ


RedEval


HExPHI


HolisticBias


HypothesisStereotypes


LatentJailbreak


MaliciousInstruct


MaliciousInstructions


MIC


MoralChoice


DialogueSafety


PersonalInfoLeak


PhysicalSafetyInstructions


PromptExtractionRobustness


PromptHijackingRobustness


QHarm


RedditBias


RuLES


SafeText


SafetyPrompts


SimpleSafetyTests


SPMisconceptions


StrongREJECT

  • Year: 2023
  • Publication Name: A StrongREJECT for Empty Jailbreaks
  • Purpose: broad safety
  • LLMDev Usage: eval only
  • Entries: 346 prompts
  • Entries Type: chat
  • Language: English
  • GitHub: https://github.com/alexandrasouly/strongreject/tree/main
  • HF Link: [not available](not available)
  • Notes:
    - The focus of the work is adversarial / to jailbreak LLMs
    - The 6 question categories are: disinformation/deception, hate/harassment/discrimination, illegal goods/services, non-violent crimes, sexual content, violence

SycophancyEval


TDCRedTeaming


ToxiGen


TruthfulQA


XSTest