Find the right dataset.

Compare three resources for fact-checking speech and dialogue.

Compare datasets

Scroll the table horizontally to compare all three datasets.

Comparison of TRILOGUE, Lost in Speech, and MAD2 based on their linked dataset cards.
TRILOGUEDialogue claims + evidenceLost in SpeechArticle alteration detectionMAD2Sentence check-worthiness
Collection size11,957 dialogues12,013 article-derived samples3,712 unaltered · 8,301 altered1,000 conversations
Use it to studyWhich dialogue claims need checking, whether they are true, and what evidence supports them.Whether an article-derived sample contains a factual or contextual alteration.Which sentences in a conversation are worth fact-checking.
LanguagesEnglish · Russian · KazakhEnglish · Russian · KazakhEnglish
Audio released8,036 recordings · 242.5 hRussian & Kazakh; English withheld.3,978 recordingsSynthetic Kazakh; English & Russian not released.1,000 recordings · 11.04 hReplacement MoonCast audio; different from the paper’s experimental recordings.
Source & downloadTRILOGUE on Hugging FaceLost in Speech on Hugging FaceMAD2 on Hugging Face

Read the linked dataset cards for current availability, definitions, and exceptions. “Released” refers to included files, subject to each repository’s access conditions.

Source relationships & article mapping

Generated dialogues

TRILOGUE

Multi-turn dialogues with check-worthiness, factuality, evidence, and paired speech.

Verified source mapping
Article-derived samples

Lost in Speech

Unaltered articles and controlled rewrites, represented as text, synthesized speech, and ASR.

The verified mapping links 11,956 dialogue inputs and 3,665 article families. Original article texts remain withheld upstream.

Mapping methods and unresolved cases

Cite the datasets

TRILOGUE

BibTeX
@inproceedings{chun-etal-2026-trilogue,
  title     = {{TRILOGUE}: A Trilingual Spoken Dialogue Fact-Checking
               Benchmark with Evidence and Paired Audio},
  author    = {Chun, Chaewan and Aristombayeva, Meruyert and
               Choi, Jiyoung and Nahar, Mahjabin and
               Zhang, Delvin Ce and Lee, Dongwon},
  booktitle = {Findings of the Association for Computational Linguistics:
               EMNLP 2026},
  year      = {2026}
}

Gated research access. Source articles are linked rather than redistributed. English TTS audio is withheld; seven Kazakh human recordings have incomplete final audio.

Lost in Speech

BibTeX
@inproceedings{aristombayeva-etal-2026-lost,
  title     = {Lost in Speech: Trilingual Spoken Hallucination Detection
               Across Audio and Transcripts},
  author    = {Aristombayeva, Meruyert and Lucas, Jason S. and
               Chun, Chaewan and Lee, Dongwon},
  booktitle = {Proceedings of the Second Workshop on Speech and Audio
               Language Models (SALMA)},
  year      = {2026}
}

Gated research access. Original texts are withheld even where ASR is available. Exclude Russian IDs 8333, 8412, and 8456 pending upstream review; their labels are not verified.

MAD2

BibTeX
@inproceedings{chun-etal-2026-context-aware,
  author    = {Chun, Chaewan and Zhang, Delvin Ce and Lee, Dongwon},
  title     = {Context-Aware Multimodal Claim Verification in Spoken Dialogues},
  booktitle = {The 2nd Speech and Audio Language Models Workshop
               ({SALMA}), {EMNLP}},
  year      = {2026}
}

The supplied audio is a replacement corpus. Published model results and WER concern the original recordings. Current sentence labels indicate check-worthiness, not factual correctness.

The current card says the release version remains to be finalized.