This dataset is designed to tackle the core weaknesses of today's large language models when it comes to processing long documents and performing complex reasoning. It consists of 7,500 high-quality training examples across three languages—Chinese, English, and Korean. Each instance is built around a long-text passage and includes questions that require synthesizing information across paragraphs and documents, while following multi-step logical chains. The goal is to offer a thorough and rigorous evaluation framework that tests a model's ability to perceive long-range context, retrieve relevant information, construct sound reasoning paths, and trace evidence back to its source.
For more details, please refer to the link: https://www.nexdata.ai/datasets/llm/2121?source=Github
Long-document Multi-hop Reasoning QA Dataset
7,500
id、context、file_count、question、answer、reasoning_chain、supporting_evidence、hops
ZH,EN,KO
JSON
Commercial License