Skip to content

Translate Hugging Face blog post: The Open ASR Leaderboard Adds Its First Global South Language - #200

Open
Jwaminju wants to merge 1 commit into
mainfrom
translate/open-asr-leaderboard-global-south
Open

Translate Hugging Face blog post: The Open ASR Leaderboard Adds Its First Global South Language#200
Jwaminju wants to merge 1 commit into
mainfrom
translate/open-asr-leaderboard-global-south

Conversation

@Jwaminju

Copy link
Copy Markdown
Collaborator

Source: https://huggingface.co/blog/open-asr-leaderboard-global-south

This PR adds a Korean translation draft for open-asr-leaderboard-global-south.

Downstream handoff:

  • SEO review should use the translation-flow manifest.
  • Quality review should use the translation-flow manifest.

@Jwaminju Jwaminju added the hf-agent:managed Opt PR into HF Agent review automation label Aug 29, 2026
@github-actions

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

🚀 View preview at
https://hugging-face-krew.github.io/pr-preview/pr-200/

Built to branch gh-pages at 2026-08-29 07:01 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@Jwaminju

Jwaminju commented Aug 29, 2026

Copy link
Copy Markdown
Collaborator Author

HF Agent Review

Gate Result
Quality ❌ Fail
SEO ✅ Pass

Head SHA: 845ef4c71d96d0ea47fa78a3d05292de6873fbd2

Quality report — ❌ Fail

Quality Report

  • Status: reject
  • Quality Score: 0.0
  • Hard failures: 1
  • Issues: 92
  • Source available: True
  • Source changed: False
  • Source segments: 95
  • Target segments: 95

Scorecard

Dimension Score
adequacy 0.0
technical_accuracy 60.0
completeness 0.0
terminology 0.0
fluency 0.0
publishing_integrity 80.0
style_locale 60.0

Metrics

  • qe_metric: heuristic
  • qe_average: 0.9224
  • qe_min: 0.4775
  • embedding_similarity_average: 0.8502
  • embedding_similarity_min: 0.2372
  • cache_hits: 0
  • cache_misses: 190

MQM Judge

  • Enabled: True
  • Provider: openai
  • Model: gpt-5.6-luna
  • Reasoning effort: none
  • Prompt: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/judges/mqm_prompt.md
  • Prompt hash: 887d2931aa289213f0bdce4a917a0ac8364dad8e470011069b8f8bf758d69a90
  • Style guide hash: 937d8cd893578d30e716a3eb513cdf5f10d6fd3ad8f5e77068b57f96e160de12
  • Requested segments: 95
  • Evaluated segments: 95
  • MQM errors: 49
  • Cache hits: 0
  • Cache misses: 95
  • Severity counts: {'major': 3, 'minor': 46}
  • adequacy_average: 0.9587
  • technical_average: 0.9901
  • fluency_average: 0.9429

Style Guide

  • Enabled: True
  • Guide: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/style/hf-blog-ko-translation-guide.md
  • Policy: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/configs/style_policy.yml
  • Style score: 60.0
  • Rule hits: {'alt_text_caption': 6, 'information_addition': 4, 'link_text_translation': 15, 'list_consistency': 1, 'modal_strength': 10}

Style Guide Findings

Rule Severity Segment Current Suggested
list_consistency minor phrase, phrase, phrase, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence Use either sentence-style endings or phrase-style endings consistently within one list.
modal_strength major p_037 English 세트는 표준 문자열 참조를 사용하며, 리더보드의 정규화기는 대부분의 철자 변형을 하나로 통합합니다. Hindi에는 이러한 변형이 훨씬 많고, 정규화기로 해결할 수 없습니다. 변형이 두 관례 사이의 고정된 매핑이 아니기 때문입니다. 따라서 Hindi 세트는 lattice를 제공합니다. 전사의 각 구간마다 정답으로 허용되는 철자 목록이 포함됩니다. Preserve the strength of can using: 수 있습니다.
modal_strength major l_050 지역이나 휴대전화도 점수를 좌우하지 않습니다: Indian English 공개 세트는 30개 주 및 연방 직할지에 걸친 428개의 고유 지구를 포함합니다. Hindi 세트는 Hindi belt 언어이므로 더 좁게 집중되어 있지만, 여전히 202개 및 295개의 지구를 아우릅니다. 녹음에는 315개에서 582개의 서로 다른 기기 모델이 사용되었으며, 어떤 하위 집합에서도 단일 모델이 세그먼트의 2.1%를 초과하지 않습니다. 표준화된 하드웨어로 수집한 말뭉치는 하나의 마이크 응답에 과적합되지만, 이 데이터셋은 그럴 수 없습니다. Preserve the strength of can using: 수 있습니다.
modal_strength major p_061 두 언어는 서로 다른 지리적 분포를 보이며, 이 분포는 유의미한 정보를 제공합니다. Hindi 세트는 Hindi belt에 집중되어 있으며, 화자의 약 40%를 Uttar Pradesh가 차지합니다. 이는 인구 비례로 표본을 추출한 Hindi 말뭉치에서 예상되는 모습입니다. Indian English 세트의 분포는 훨씬 평평합니다. 어떤 주도 13%를 초과하지 않으며, 화자의 3분의 1은 규모가 가장 큰 8개 주 밖에서 왔습니다. 공개 및 비공개 절반은 두 언어 모두에서 매우 유사한 분포를 보입니다. Preserve the strength of should using: 좋습니다, 해야 합니다.
modal_strength major l_067 발화 유도: 대규모로 즉흥 발화를 유도하는 일은 그 자체로 어렵습니다. 구조화된 안내가 없으면 참여자는 짧고 내용이 빈약한 응답을 하는 경향이 있기 때문입니다. 따라서 각 대화는 개방형 서술 단서로 시작하고, 여행, 의료, 농업, 교육, 디지털 서비스 등을 포함하는 영역의 후속 질문을 점진적으로 공개했습니다. 이를 통해 대화를 대본 없이도 긴 설명으로 이끌었습니다. 후보 주제는 대규모 언어 모델로 생성한 뒤 모국어 화자인 언어학자가 검토하고 현지화했습니다. Preserve the strength of can using: 수 있습니다.
modal_strength major p_072 리더보드의 8개 모델은 이 세트에서 4.81에서 4.99 WER 사이의 점수를 기록합니다. 최고와 최저의 차이는 0.18포인트로, 5시간의 데이터가 구분할 수 있는 범위 안에 있습니다. 말뭉치 전체에서 순위를 매기면 이들은 동일한 모델입니다. Preserve the strength of can using: 수 있습니다.
modal_strength major p_073 화자를 지역별로 묶으면 다른 이야기가 나타납니다. 각 화자의 출신 지구를 인도 주를 분류하는 내무부의 구역 위원회 기준에 따라 상위 구역으로 집계하면, 표본이 충분한 5개 구역이 됩니다. openai/whisper-large-v3-turbo의 차이는 이들 사이에서 0.46포인트입니다. 말뭉치 전체에서는 mistralai/Voxtral-Mini-3B-2507보다 14/100포인트 뒤처졌지만, 구역별 차이는 1.68포인트로, 중부 구역에서는 4.38, 동부 구역에서는 6.06을 기록합니다. 리더보드에서는 구분할 수 없는 두 시스템이 화자의 출신 지역에 따라 정확도가 얼마나 달라지는지에서는 거의 네 배의 차이를 보입니다. Preserve the strength of up to using: 최대.
modal_strength major p_078 English의 표기 변이는 한정적입니다. 영국식과 미국식 철자, 구두점, 대소문자, 숫자와 단어의 차이 등은 정규화기가 대부분 하나의 형식으로 매핑할 수 있으며, 리더보드의 정규화기도 그렇게 합니다. Hindi는 같은 방식으로 한정되지 않습니다. 일상적인 발화에는 코드 믹싱이 많이 나타나고, English에서 유래한 단어에는 정착된 Devanagari 철자가 없으며, 합성어는 선호에 따라 붙여 쓰거나 띄어 씁니다. 하나의 구절에 유효한 표기 형식이 10개 이상 존재할 수 있으며, 이를 하나로 통합할 고정된 매핑도 없습니다. 매핑할 표준적인 한쪽 형식이 없기 때문입니다. Preserve the strength of can using: 수 있습니다.
modal_strength major p_081 따라서 Hindi 세트는 lattice를 제공합니다. 전사의 각 구간마다 정답으로 허용되는 표기 형식의 집합이 포함됩니다. 이를 구축하는 일은 수작업입니다. 동일한 오디오에 대한 여러 ASR 전사에서 후보 변형을 추출하고 언어 모델로 확장한 다음, 모국어 화자인 언어학자가 해당 발화에 유효한 형식을 결정하고 나머지를 제거합니다. 이를 통해 실제 발화 내용과 일치하는 형식만 허용됩니다. Preserve the strength of can using: 수 있습니다.
modal_strength major p_086 또한 모든 사람이 이 세트의 결과를 직접 재현할 수 있도록 구현체 voi-oiwer를 오픈 소스로 공개합니다. Preserve the strength of can using: 수 있습니다.

Issues

QL-001 formatting / critical

  • Message: Front matter key authors changed or is missing.
  • Source: user: vanshikachhabra-voicearena
  • Target: user: bezzam user: Shobhitbanga
  • Suggested fix: Preserve front matter authors exactly.

QL-002 technical / major

  • Message: model or dataset id mismatch.
  • Source: M/F, States/UTs, ibm-granite/granite-speech-3.3-2b, microsoft/VibeVoice-ASR-HF, mistralai/Voxtral-Mini-3B-2507, mistralai/Voxtral-Mini-3B-2507, openai/whisper-large-v3-turbo
  • Target: ibm-granite/granite-speech-3.3-, microsoft/VibeVoice-ASR-, mistralai/Voxtral-Mini-3B-, mistralai/Voxtral-Mini-3B-, openai/whisper-large-v3-
  • Suggested fix: Preserve source model or dataset id exactly.
  • Reason: Review gate exact-match validator failed: missing=['M/F', 'States/UTs', 'ibm-granite/granite-speech-3.3-2b', 'microsoft/VibeVoice-ASR-HF', 'mistralai/Voxtral-Mini-3B-2507', 'mistralai/Voxtral-Mini-3B-2507', 'openai/whisper-large-v3-turbo']; extra=['ibm-granite/granite-speech-3.3-', 'microsoft/VibeVoice-ASR-', 'mistralai/Voxtral-Mini-3B-', 'mistralai/Voxtral-Mini-3B-', 'openai/whisper-large-v3-']

QL-003 technical / major

  • Message: number/unit token mismatch.
  • Target: 1, 10, 10, 100, 12, 14, 15, 2
  • Suggested fix: Preserve source number/unit token exactly.
  • Reason: Review gate exact-match validator failed: extra=['1', '10', '10', '100', '12', '14', '15', '2']

QL-004 accuracy / major

  • Message: QE metric score is low.
  • Target: 0.4775
  • Suggested fix: Review this segment for omission, unrelated translation, or over-compression.
  • Reason: QE score is below threshold 0.55.

QL-005 accuracy / major

  • Message: QE metric score is low.
  • Target: 0.4803
  • Suggested fix: Review this segment for omission, unrelated translation, or over-compression.
  • Reason: QE score is below threshold 0.55.

QL-006 accuracy / major

  • Message: QE metric score is low.
  • Target: 0.5161
  • Suggested fix: Review this segment for omission, unrelated translation, or over-compression.
  • Reason: QE score is below threshold 0.55.

QL-007 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: The Open ASR Leaderboard Adds Its First Global South Language
  • Target: Open ASR Leaderboard에 최초의 글로벌 사우스 언어가 추가되다
  • Suggested fix: Open ASR Leaderboard, 최초의 글로벌 사우스 언어 추가
  • Reason: 의미는 전달되지만 ‘추가되다’는 제목으로 다소 직역투이고 부자연스럽습니다. 제목에서는 핵심 사건을 명확한 명사형이나 능동형으로 표현하는 편이 자연스럽습니다.

QL-008 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Voice Arena and Hugging Face partner to launch open ASR evaluation for Hindi and Indian English
  • Target: Voice Arena와 Hugging Face가 Hindi 및 Indian English를 위한 오픈 ASR 평가를 출시하기 위해 협력합니다
  • Suggested fix: Voice Arena와 Hugging Face, Hindi 및 Indian English용 오픈 ASR 평가 공개
  • Reason: 제목을 문장형으로 직역해 다소 어색하고 길게 들립니다. 특히 ‘평가를 출시하기 위해 협력합니다’는 한국어 제목 표현으로 부자연스럽습니다.

QL-009 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: Held-out private splits.
  • Target: Held-out private splits.
  • Suggested fix: 비공개 홀드아웃 분할 데이터
  • Reason: 일반적인 리스트 항목이 원문 그대로 남아 있어 한국어 독자가 의미를 바로 이해하기 어렵습니다.

QL-010 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: Benchmark-fitting analysis
  • Target: Benchmark-fitting analysis
  • Suggested fix: 벤치마크 적합성 분석(Benchmark-fitting analysis)
  • Reason: 일반 기술 용어인 'Benchmark-fitting analysis'가 문장 끝에 영어로만 남아 있어 한국어 목록 항목의 흐름이 어색하고 독자가 의미를 바로 파악하기 어렵습니다. 핵심 의미인 벤치마크 적합성 분석을 한국어로 제시하고 필요하면 영어를 병기하는 편이 적절합니다.

QL-011 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: Racial disparities in automated speech recognition
  • Target: Racial disparities in automated speech recognition은
  • Suggested fix: 연구 제목 표기를 유지해 Racial disparities in automated speech recognition에서는
  • Reason: 문장 안에서 보고서 또는 연구 제목으로 사용된 표현이 번역되지 않아 한국어 문장 흐름이 끊깁니다. 검색성을 위해 영문 제목을 유지하더라도 제목임을 표시하거나 한국어 설명을 덧붙이는 편이 자연스럽습니다.

QL-012 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: Hindi, **spoken by more than half a billion people**, is the first Indic language on a multilingual tab that currently covers only European languages.
  • Target: **5억 명이 넘는 사람들이 사용하는** Hindi는 현재 유럽 언어만 다루는 다국어 탭에 추가되는 최초의 Indic language입니다.
  • Suggested fix: 5억 명이 넘는 사람들이 사용하는 Hindi는 현재 유럽 언어만 다루는 다국어 탭에 추가되는 최초의 인도계 언어(Indic language)입니다.
  • Reason: 검색성이 필요한 기술·언어 분류 용어인 ‘Indic language’를 한국어로 번역하거나 영문 병기하지 않아, 한국어 독자가 의미를 파악하기 어렵고 원문 용어를 검색하기도 어렵습니다.

QL-013 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: A test set can only expose a failure mode it varies along.
  • Target: 테스트 세트는 변화를 주어 수집한 축을 따라서만 실패 모드를 드러낼 수 있습니다.
  • Suggested fix: 테스트 세트는 변화를 포함하는 축에서 나타나는 실패 모드만 드러낼 수 있습니다.
  • Reason: 원문의 ‘테스트 세트가 변화를 포함하는 축을 따라 나타나는 실패 모드만 드러낼 수 있다’는 의미는 전달되지만, ‘변화를 주어 수집한 축을 따라서만’은 한국어로 다소 어색하고 원문의 논리 관계가 불분명합니다.

QL-014 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: Most benchmarks are built from whatever audio was readily available.
  • Target: 대부분의 벤치마크는 쉽게 구할 수 있었던 오디오로 구성됩니다.
  • Suggested fix: 대부분의 벤치마크는 쉽게 이용할 수 있었던 오디오를 바탕으로 구성됩니다.
  • Reason: ‘was readily available’은 해당 시점에 손쉽게 이용할 수 있었던 오디오라는 뜻인데, ‘구할 수 있었던’으로 옮기면서 수집·획득의 의미가 다소 강해졌습니다.

QL-015 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: Each is a way an aggregate WER can be right on average and wrong for a particular population.
  • Target: 각각은 집계된 WER이 평균적으로는 맞지만 특정 집단에 대해서는 틀릴 수 있는 방식입니다.
  • Suggested fix: 이러한 각 축은 집계된 WER이 평균적으로는 정확해 보여도 특정 집단에서는 부정확할 수 있음을 보여 줍니다.
  • Reason: ‘Each is a way’는 앞서 제시한 각 축이 평균적으로는 정확하지만 특정 집단에서는 부정확할 수 있는 집계 WER의 양상을 보여 주는 기준이라는 뜻입니다. 현재 번역은 ‘각각은 ... 방식’으로 직역되어 의미가 다소 모호합니다.

QL-016 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: The collection method follows from that.
  • Target: 수집 방법은 이러한 설계에서 자연스럽게 도출됩니다.
  • Suggested fix: 수집 방법은 앞서 설명한 내용에서 도출됩니다.
  • Reason: 원문의 'that'은 앞서 언급된 내용을 가리키는 일반적인 지시어인데, 번역에서 이를 '이러한 설계'로 구체화해 원문에 없는 '설계'라는 정보를 추가했습니다. 또한 'naturally'에 해당하는 의미도 원문에는 없습니다.

QL-017 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: comes from recruiting across hundreds of districts
  • Target: 수백 개의 지구에서 모집하여 확보했습니다
  • Suggested fix: 수백 개의 지구에 걸쳐 모집한 데서 비롯됩니다
  • Reason: 원문은 지리적 범위가 수백 개 지역에 걸쳐 모집한 결과임을 설명하지만, 번역은 주어인 ‘지리적 범위’를 직접 확보했다는 의미로 표현해 논리적 주체가 어색하고 의미 관계가 흐려집니다.

QL-018 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: using their own handsets and connections, indoors and out
  • Target: 참여자가 자신의 휴대전화와 네트워크를 사용해 실내와 실외에서 녹음하여 확보했습니다.
  • Suggested fix: 참여자가 자신의 기기와 네트워크를 사용해 실내와 실외에서 직접 수집한 것이며
  • Reason: 원문은 참여자가 자신의 기기와 연결을 사용했다는 사실을 설명하지만, 목표 문장은 '녹음하여 확보했습니다'를 추가해 해당 데이터가 녹음으로 확보되었다고 단정합니다. 문맥상 가능할 수 있으나 이 세그먼트만으로는 명시되지 않은 정보입니다.

QL-019 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: everyday topics that push contributors toward opinion, disagreement, narration and recall, which is where named entities, numbers and unrehearsed phrasing appear.
  • Target: 일상적인 주제가 참여자를 의견 제시, 이견 표현, 서술, 회상으로 이끌도록 하며, 이 과정에서 고유명사, 숫자, 준비되지 않은 표현이 등장합니다.
  • Suggested fix: 일상적인 주제는 참여자가 의견을 말하고, 이견을 표현하고, 이야기를 서술하고, 기억을 떠올리도록 유도합니다. 이 과정에서 고유명사, 숫자, 즉흥적인 표현이 등장합니다.
  • Reason: 현재 번역은 ‘주제가 참여자를 ... 이끌도록 하며’라는 구조가 다소 어색하고, ‘unrehearsed phrasing’를 ‘준비되지 않은 표현’으로 옮겨 한국어 기술 문맥에서 의미가 불분명합니다. 원문은 일상적인 주제가 참여자를 특정 화행으로 유도하고, 그 결과 고유명사·숫자·즉흥적인 표현이 나타난다는 뜻입니다.

QL-020 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Four splits, two languages, collected through one pipeline.
  • Target: 하나의 파이프라인을 통해 수집한 네 개의 스플릿과 두 개의 언어입니다.
  • Suggested fix: 하나의 파이프라인을 통해 네 개의 스플릿과 두 언어를 수집했습니다.
  • Reason: 원문의 압축적인 나열 구조를 그대로 옮겨 한국어 문장이 다소 부자연스럽고, 무엇이 수집되었는지 한 번 더 해석해야 합니다.

QL-021 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: Duration
  • Target: 길이
  • Suggested fix: 지속 시간
  • Reason: Duration은 일반적인 길이가 아니라 녹음 또는 클립의 지속 시간을 뜻하므로, '길이'로 옮기면 Clip length와 의미가 겹치고 구분이 불명확해집니다.

QL-022 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: Devices
  • Target: 기기 수
  • Suggested fix: 기기
  • Reason: 원문 'Devices'는 기기 항목 자체를 가리키며 수량을 명시하지 않습니다. '기기 수'는 개수라는 의미를 추가합니다.

QL-023 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: Lattice (accepted orthographic variants)
  • Target: lattice(허용되는 철자 변형)
  • Suggested fix: Lattice(허용되는 철자 변형)
  • Reason: 표의 다른 수치와 고유명은 보존되었지만, 원문의 기술 용어인 Lattice가 소문자 lattice로 바뀌어 검색성과 명칭 일관성이 떨어집니다. 괄호 설명도 의미는 보존되어 있습니다.

QL-024 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: Lattice (accepted orthographic variants)
  • Target: lattice(허용되는 철자 변형)
  • Suggested fix: Lattice(허용되는 철자 변형)
  • Reason: 표의 마지막 항목에서 고유명사 또는 데이터셋명으로 보이는 Lattice의 대문자가 소문자로 변경되었습니다. 검색성과 명칭 일관성을 위해 원문의 표기를 보존해야 합니다.

QL-025 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: current city and years in the current district
  • Target: 현재 도시, 현재 지구에서 거주한 연수
  • Suggested fix: 현재 도시, 현재 지역에서 거주한 기간(연수)
  • Reason: 원문의 'years in the current district'는 현재 거주 중인 행정 구역에서 보낸 연수를 뜻하지만, '거주한'은 해당 구역을 떠난 경험까지 포함할 수 있어 현재 거주 상태의 의미가 다소 흐려집니다.

QL-026 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: a lattice
  • Target: lattice
  • Suggested fix: 격자(lattice)를 제공합니다
  • Reason: 기술 문맥의 핵심 용어를 한국어로 설명하거나 영문 병기하지 않아 독자가 의미를 바로 파악하기 어렵습니다. 특히 여기서는 각 전사 구간에 허용되는 철자 목록을 나타내는 구조라는 설명이 이어집니다.

QL-027 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Speaker coverage
  • Target: 화자 범위
  • Suggested fix: 화자 커버리지
  • Reason: 문맥상 여러 화자 또는 화자 유형을 얼마나 포함하는지를 뜻하는 제목이라면 ‘화자 범위’는 직역투로 다소 어색하고 의미가 불명확할 수 있습니다.

QL-028 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: **Speaker concentration and diversity beyond the fields above.**
  • Target: **위의 필드로 나타나는 범위를 넘어선 화자 집중도와 다양성.**
  • Suggested fix: 위 필드 외의 화자 집중도와 다양성.
  • Reason: ‘fields’와 ‘beyond’의 관계를 직역해 ‘위의 필드로 나타나는 범위를 넘어선’이 어색하고 의미가 불분명합니다. 위 필드에 포함되지 않는 추가적인 화자 집중도와 다양성을 가리키는 표현으로 다듬는 것이 자연스럽습니다.

QL-029 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: Device manufacturers | 18 | 25 | 23 | 20
  • Target: 기기 제조업체 수 | 18 | 25 | 23 | 20
  • Suggested fix: 기기 제조업체 | 18 | 25 | 23 | 20
  • Reason: 원문은 ‘기기 제조업체’를 나타내지만, ‘수’를 추가해 각 수치가 제조업체의 개수라는 의미로 범위를 좁혔습니다. 표의 문맥상 자연스러울 수 있으나 원문에 없는 의미가 추가되었습니다.

QL-030 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: No region or handset carries it either:
  • Target: 지역이나 휴대전화도 점수를 좌우하지 않습니다:
  • Suggested fix: 지역이나 휴대전화도 이를 설명하지 못합니다:
  • Reason: 원문의 대명사 it은 앞선 맥락에서 특정 특성이나 편향을 가리키는 표현인데, 번역문은 이를 ‘점수를 좌우하지 않습니다’로 구체화했습니다. 원문에 없는 ‘점수’라는 의미를 추가해 주장 범위를 임의로 좁힙니다.

QL-031 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: The Indian English public set
  • Target: Indian English 공개 세트
  • Suggested fix: 인도 영어(Indian English) 공개 데이터셋
  • Reason: Indian English를 영어로만 남겨 한국어 독자가 의미를 즉시 파악하기 어렵고, ‘public set’도 ‘공개 세트’로 직역되어 기술 문맥에서 다소 불명확합니다.

QL-032 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: the Hindi sets, being a Hindi-belt language
  • Target: Hindi 세트는 Hindi belt 언어이므로
  • Suggested fix: 힌디어(Hindi) 데이터셋은 힌디어권(Hindi belt)에 해당하는 언어이므로
  • Reason: Hindi와 Hindi belt가 번역 또는 병기되지 않아 검색성과 가독성이 떨어집니다. 또한 ‘Hindi-belt language’는 언어 자체가 Hindi belt라는 뜻처럼 읽힐 수 있습니다.

QL-033 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: with no single model exceeding 2.1% of segments in any subset.
  • Target: 어떤 하위 집합에서도 단일 모델이 세그먼트의 2.1%를 초과하지 않습니다.
  • Suggested fix: 어떤 하위 집합에서도 단일 모델이 차지하는 세그먼트 비중은 2.1%를 넘지 않습니다.
  • Reason: 의미는 전달되지만 ‘단일 모델이 세그먼트의 2.1%를 초과하지 않습니다’는 주어와 비율의 관계가 다소 어색합니다. 모델별 세그먼트 비중이라는 뜻을 명확히 하는 편이 자연스럽습니다.

QL-034 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: This is English as it is spoken across the country, not the English of one region.
  • Target: 이는 한 지역의 English가 아니라 전국에서 사용되는 English입니다.
  • Suggested fix: 이는 한 지역에서 사용되는 영어가 아니라 전국에서 사용되는 영어입니다.
  • Reason: Indian English를 설명하는 문맥에서 일반 명사인 English를 번역하지 않고 반복해 한국어 문장이 부자연스럽습니다. 원문의 의미는 유지되지만, 독자에게 자연스럽게 전달되도록 ‘영어’로 옮기는 것이 적절합니다.

QL-035 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: where most public ASR test sets ship an identifier, a transcript and a duration.
  • Target: 대부분의 공개 ASR 테스트 세트가 식별자, 전사, 길이만 제공하는 것과 대조적입니다.
  • Suggested fix: 대부분의 공개 ASR 테스트 세트는 식별자, 전사, 길이만 제공하는 반면, Monsoon은 더 많은 메타데이터를 제공합니다.
  • Reason: 원문은 Monsoon의 메타데이터 구성과 대부분의 공개 ASR 테스트 세트의 구성 차이를 설명하지만, where는 단순한 대조가 아니라 앞서 언급한 메타데이터의 구체적 맥락을 연결합니다. 현재 번역의 ‘대조적입니다’는 원문에 없는 명시적 대조 관계를 추가해 의미 흐름을 다소 바꿉니다.

QL-036 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: district and state carry real accent signal
  • Target: 지구와 주는 실제 억양 신호를 포함합니다
  • Suggested fix: district와 주는 실제 억양 신호를 담고 있습니다
  • Reason: 원문의 ‘district and state’는 행정구역인 ‘district와 state’를 가리키며, ‘carry real accent signal’은 억양과 실제로 관련된 신호를 지닌다는 의미입니다. ‘지구’는 한국어에서 지리적 영역이나 지구(planet)로도 읽혀 행정구역 의미가 불명확합니다.

QL-037 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Those runs were on a closed benchmark.
  • Target: 이러한 실행은 비공개 벤치마크에서 수행되었습니다.
  • Suggested fix: 이러한 분석은 비공개 벤치마크에서 진행되었습니다.
  • Reason: ‘runs’는 이 문맥에서 분석 또는 실험 결과를 뜻하므로 ‘실행’은 다소 기계적이며, ‘수행되었습니다’와 함께 번역투가 생깁니다. 분석 결과가 비공개 벤치마크에서 얻어졌다는 의미를 자연스럽게 전달하는 편이 적절합니다.

QL-038 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: Pairs then recorded two-person conversations over a peer-to-peer interface, dual-channel, on assigned everyday topics.
  • Target: 이후 두 명씩 짝을 이루어 지정된 일상 주제에 관해 P2P 인터페이스를 통해 양방향 채널로 대화를 녹음했습니다.
  • Suggested fix: 이후 두 명씩 짝을 이루어 지정된 일상 주제에 관해 피어 투 피어 인터페이스를 통해 2채널로 대화를 녹음했습니다.
  • Reason: dual-channel은 두 채널로 각각 녹음하는 방식을 뜻하는데, ‘양방향 채널’은 채널의 방향성으로 읽혀 기술적 의미가 불명확합니다.

QL-039 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: A per-speaker duration cap, calibrated per language to the population size and geographic distribution of its speakers, prevented a small number of prolific contributors from dominating a language or region
  • Target: 화자별 지속 시간 상한은 각 언어 화자의 인구 규모와 지리적 분포에 맞게 조정하여, 소수의 활발한 참여자가 특정 언어나 지역을 지배하지 못하도록 했습니다.
  • Suggested fix: 언어별 화자 인구 규모와 지리적 분포에 맞게 조정한 화자별 녹음 시간 상한을 적용해, 소수의 녹음을 많이 제출하는 참여자가 한 언어나 지역을 좌우하지 못하도록 했습니다.
  • Reason: 원문은 상한을 설정한 주체가 소수의 기여자가 언어나 지역을 지배하는 것을 막았다고 명시하지만, 번역문은 주어와 목적 관계가 다소 불분명하고 ‘특정 언어나 지역’으로 범위를 좁혀 읽힐 수 있습니다.

QL-040 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: Signal-to-noise ratio estimation removed recordings degraded beyond intelligibility, while natural environmental background noise was deliberately preserved so that the acoustic realism of in-the-wild speech is retained.
  • Target: 신호 대 잡음비 추정으로 명료도를 저해하는 수준으로 품질이 저하된 녹음을 제거했지만, 자연스러운 환경 소음은 의도적으로 보존하여 실제 환경에서 수집된 발화의 음향적 현실성을 유지했습니다.
  • Suggested fix: 신호 대 잡음비 추정으로 알아들을 수 없을 정도로 품질이 저하된 녹음을 제거했지만, 자연스러운 환경 소음은 의도적으로 보존하여 통제되지 않은 실제 환경의 발화가 지닌 음향적 현실성을 유지했습니다.
  • Reason: 원문의 ‘in-the-wild speech’는 통제되지 않은 실제 환경에서 발생하는 발화를 뜻하지만, ‘실제 환경에서 수집된 발화’는 수집 행위와 데이터 출처로 의미가 다소 좁혀집니다. 기술적 맥락을 더 정확히 보존하는 표현이 필요합니다.

QL-041 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: none of which appear on any public leaderboard
  • Target: 이 모델들은 어떤 공개 리더보드에도 등장하지 않으므로
  • Suggested fix: 이 내부 ASR 모델 중 어떤 것도 공개 리더보드에 등장하지 않으므로
  • Reason: 원문의 ‘none of which’는 내부 ASR 모델들이 공개 리더보드에 등장하지 않는다는 뜻이지만, ‘이 모델들은’은 바로 앞의 ‘모델’만 받아 복수 모델 전체를 가리키는 구조로 다소 모호합니다. 또한 원문은 사실을 제시하는 반면 ‘~므로’는 인과 관계를 덧붙여 평가 시스템과 참조 전사의 독립성에 대한 논리적 연결을 과도하게 단정할 수 있습니다.

QL-042 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: It is not the finding the sets exist to deliver
  • Target: 이는 이 세트가 제공하고자 존재하는 핵심 발견이 아니라
  • Suggested fix: 이는 이 세트들이 제공하기 위해 존재하는 발견이 아니라
  • Reason: 원문의 ‘이 세트들이 제공하기 위해 존재하는 발견이 아니다’라는 의미를 ‘제공하고자 존재하는 핵심 발견’으로 옮겨 문법적으로 어색하고 의미 관계가 불명확합니다. 특히 ‘핵심’은 원문에 없는 정보입니다.

QL-043 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: What follows is one example, run on the public Indian English split, to show the kind of evaluation the metadata makes possible.
  • Target: 이하에서는 공개 Indian English 스플릿을 대상으로 수행한 한 가지 예를 통해 메타데이터로 어떤 평가가 가능한지 보여줍니다.
  • Suggested fix: 다음은 공개 Indian English 스플릿에서 실행한 예로, 메타데이터를 통해 어떤 평가가 가능한지 보여줍니다.
  • Reason: ‘이하에서는’과 ‘수행한 한 가지 예를 통해’가 다소 직역투이고 문장이 무겁습니다. 의미는 전달되지만 기술 블로그 문체로는 ‘다음은 ... 실행한 예입니다’처럼 다듬는 편이 자연스럽습니다.

QL-044 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: inside what five hours can resolve
  • Target: 5시간의 데이터가 구분할 수 있는 범위 안에 있습니다
  • Suggested fix: 5시간으로 구분할 수 있는 범위 안에 있습니다
  • Reason: 원문의 ‘five hours’는 측정 또는 평가에 사용된 5시간의 데이터가 아니라, 5시간이라는 시간 규모에서 구분 가능한 해상도/범위를 뜻합니다. ‘5시간의 데이터’로 옮기면 의미가 데이터 양으로 바뀌어 독자가 벤치마크의 한계를 잘못 이해할 수 있습니다.

QL-045 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Ranked on the corpus, they are the same model.
  • Target: 말뭉치 전체에서 순위를 매기면 이들은 동일한 모델입니다.
  • Suggested fix: 말뭉치 전체를 기준으로 순위를 매기면 이 모델들은 사실상 같은 수준입니다.
  • Reason: ‘동일한 모델입니다’는 모델 자체가 같다는 뜻으로 읽힐 수 있어, 원문의 비유적 결론인 성능상 구분되지 않는다는 의미가 불명확합니다.

QL-046 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: mistralai/Voxtral-Mini-3B-2507, fourteen hundredths of a point behind it on the corpus, varies by 1.68
  • Target: 말뭉치 전체에서는 mistralai/Voxtral-Mini-3B-2507보다 14/100포인트 뒤처졌지만, 구역별 차이는 1.68포인트로
  • Suggested fix: 말뭉치 전체에서는 mistralai/Voxtral-Mini-3B-2507보다 0.14포인트 뒤처졌지만, 이 모델의 구역별 차이는 1.68포인트로
  • Reason: 원문의 'fourteen hundredths of a point'는 0.14포인트라는 수치인데, '14/100포인트'로 표기되어 기술 문서의 수치 표현이 부자연스럽고 명확성이 떨어집니다. 또한 원문에서 주어인 Voxtral 모델의 구역별 변동 폭이 1.68이라는 관계는 번역에서도 이해할 수 있지만, 문장 구조상 더 명확히 표현할 수 있습니다.

QL-047 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: The same breakdown runs on age, education, occupation and handset
  • Target: 동일한 분석은 연령, 학력, 직업, 휴대전화에도 적용할 수 있으며
  • Suggested fix: 동일한 분석은 연령, 학력, 직업, 휴대전화에 대해서도 수행되며
  • Reason: 원문의 현재형 진술인 'runs on'이 가능성을 나타내는 '적용할 수 있으며'로 약화되었습니다. 이 문맥에서는 동일한 분석이 해당 속성들에 대해서도 수행된다는 의미를 유지해야 합니다.

QL-048 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Orthographic variation in Hindi
  • Target: Hindi의 표기 변이
  • Suggested fix: 힌디어의 표기 변이
  • Reason: 의미는 전달되지만 영어 고유명사와 한국어 조사·명사를 직접 결합한 표현이 제목에서 다소 어색합니다. 한국어 기술 문서에서는 ‘힌디어의 표기 변이’가 더 자연스럽습니다.

QL-049 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: pull request
  • Target: pull request
  • Suggested fix: 풀 리퀘스트(pull request)를 열면 모델 체크리스트(model checklist)가 표시됩니다.
  • Reason: 일반 기술 용어인 ‘pull request’와 ‘model checklist’가 문장 안에서 그대로 남아 있어 한국어 독자가 의미를 즉시 파악하기 어렵고, 첫 등장 용어의 검색성도 충분히 확보되지 않았습니다.

QL-050 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: We will verify the results on the public sets and compute the metrics on the private ones.
  • Target: 공개 세트에서 결과를 검증하고 비공개 세트의 지표를 계산합니다.
  • Suggested fix: 공개
SEO report — ✅ Pass

SEO Eval Report

Gate: ✅ PASS — deterministic AND rubric

  • File: ../target/_posts/2026-08-28-open-asr-leaderboard-global-south.md
  • Source: —
  • Primary keyword: (none — D5 skipped)
  • Mode: file

Gate

  • Status: PASS
  • Blockers: ✅ pass
  • Deterministic REQUIRED (D1–D7): ✅ pass
  • Rubric (R1–R6): ✅ pass (mean None, min None)

Blockers

✅ body_not_empty: Body is not empty
✅ robots_indexable: Robots is indexable
✅ internal_links_resolve: All internal links resolve
✅ local_images_resolve: All local images resolve

Required checks (gated)

✅ heading_hierarchy: Heading hierarchy: Valid
✅ alt_text_coverage: Alt text coverage: 6/6 images
✅ descriptive_alt_text: Descriptive alt text: 6/6 (≥80% recommend)
✅ image_files_exist: All 0 local image file(s) exist

OpenAI rubric checks

✅ semantic_metadata: PASS (required) — Title, H1, headings, and opening semantically align around adding a Global South language to Open ASR Leaderboard; no semantic conflicts; description field is empty and ignored per instructions.
🟠 alt_semantics: NEEDS_CHANGES (review) — https://storage.googleapis.com/research_team_data/blog_figures/nine-axes-gray.png: Descriptive and relevant; clearly conveys chart content.; https://storage.googleapis.com/research_team_data/blog_figures/flowchart%201.png: Non-descriptive file-name style alt; not reader-friendly; consider a natural-language description of the pipeline steps.; https://storage.googleapis.com/research_team_data/csv_files/corpus_wer_versus_zone_range_coloured_by_zone.png: Describes the comparison but could be improved with capitalization and expansion to 'Word Error Rate (WER) vs. zone range'.

Advisory checks (not gated)

🟠 opening_summary: Opening 3 paragraphs: 38 words (recommend ≥50 for GEO)
✅ h1_count: Markdown H1 count: 1 (review against rendered layout)
✅ citations: Citations/statistics: 40 (recommend ≥1 for GEO)
⚠️ question_headings: Scannable H2/H3 (question or keyword): 0 (0 question, 0 keyword)
⚠️ internal_links: Internal links: 0 (recommend 2-3)
✅ word_count: Body length: 14782 chars (recommend ≥800 for KO)
ℹ️ primary_keyword: No primary_keyword in manifest — keyword check skipped
⚠️ webp_format: WebP format: 0/6 images (≥50% recommend)
⚠️ lazy_loading: Lazy loading: 0 images (optional)

Signals (evidence — not directly gated)

  • Frontmatter: title 42 chars, description 0 chars, author present True
  • Title text: Open ASR Leaderboard에 최초의 글로벌 사우스 언어가 추가되다
  • Description text: —
  • Opening text: * TOC
    {:toc}

이 글은 Hugging Face 블로그의 The Open ASR Leaderboard Adds Its First Global South Language를 한국어로 번역한 글입니다.

  • Opening: first paragraph 191 chars, first 3 paragraphs 278 chars
  • Headings: markdown H1 1, rendered effective H1 2, layout title H1 True
  • Links: total 24, external 24, internal 0, citation signals 40
  • Images: total 6, empty alt 0, filename-like alt 0, missing local files 0

Semantic review packet

  • Title: Open ASR Leaderboard에 최초의 글로벌 사우스 언어가 추가되다
  • Description: —
  • Rendered H1 candidates: Open ASR Leaderboard에 최초의 글로벌 사우스 언어가 추가되다, Open ASR Leaderboard에 최초의 글로벌 사우스 언어가 추가되다
  • Opening: * TOC
    {:toc}

이 글은 Hugging Face 블로그의 The Open ASR Leaderboard Adds Its First Global South Language를 한국어로 번역한 글입니다.

  • Canonical/permalink: —
  • Instruction: Compare title, description, rendered H1, and opening text for meaning consistency. This packet is evidence only; it does not decide pass/fail.

Frontmatter (advisory — written by metadata step, not gated)

✅ title: Title: 42 chars (recommend ≤60)
❌ description: Description is missing
✅ image: OG image: assets/images/blog/posts/2026-08-28-open-asr-leaderboard-global-south/thumbnail.png
✅ categories: Categories: 2 (recommend 2-3)
✅ author: Author: dailybot

SEO metadata suggestion — PARTIAL

This is a suggestion. SEO is applied only when the post frontmatter is updated.
To apply safe fields from a partial suggestion, leave a trusted PR comment: metadata apply.

  • Auto apply: False
  • Requires human: True
  • Mode: frontmatter_only
  • Reason: metadata candidate needs policy decisions or missing title/description

Candidate

  • categories: ['Translation', 'HuggingFace']
  • image: assets/images/blog/posts/2026-08-28-open-asr-leaderboard-global-south/thumbnail.png

Needs policy decision

  • target_url
  • source_url
  • canonical_policy
  • translation_indexing
  • target_locale
  • source_locale

Warnings

  • description is empty in frontmatter
  • source_url is empty in content

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hf-agent:managed Opt PR into HF Agent review automation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant