Skip to content

Dataset to be considered: CLERC #16

Description

@JulienGaumez

A Dataset for Legal Case Retrieval and Retrieval-Augmented Analysis Generation. We work with legal professionals to transform a large open-source legal corpus into a dataset supporting two important backbone tasks: information retrieval (IR) and retrieval-augmented generation (RAG). This dataset CLERC (Case Law Evaluation Retrieval Corpus), is constructed for training and evaluating models on their ability to (1) find corresponding citations for a given piece of legal analysis and to (2) compile the text of these citations (as well as previous context) into a cogent analysis that supports a reasoning goal.

Dataset: https://huggingface.co/datasets/jhu-clsp/CLERC
Paper: https://arxiv.org/abs/2406.17186

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions