Skip to content

Translate Hugging Face blog post: Measuring benchmark optimization in speech recognition - #197

Open
Jwaminju wants to merge 1 commit into
mainfrom
translate/asr-benchmark-optimization
Open

Translate Hugging Face blog post: Measuring benchmark optimization in speech recognition#197
Jwaminju wants to merge 1 commit into
mainfrom
translate/asr-benchmark-optimization

Conversation

@Jwaminju

Copy link
Copy Markdown
Collaborator

Source: https://huggingface.co/blog/asr-benchmark-optimization

This PR adds a Korean translation draft for asr-benchmark-optimization.

Downstream handoff:

  • SEO review should use the translation-flow manifest.
  • Quality review should use the translation-flow manifest.

@Jwaminju Jwaminju added the hf-agent:managed Opt PR into HF Agent review automation label Aug 22, 2026
@github-actions

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

🚀 View preview at
https://hugging-face-krew.github.io/pr-preview/pr-197/

Built to branch gh-pages at 2026-08-22 01:50 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@Jwaminju

Jwaminju commented Aug 22, 2026

Copy link
Copy Markdown
Collaborator Author

HF Agent Review

Gate Result
Quality ❌ Fail
SEO ✅ Pass

Head SHA: acbc7298e99282215f3fcbee3586e04287aa90d5

Quality report — ❌ Fail

Quality Report

  • Status: reject
  • Quality Score: 0.0
  • Hard failures: 5
  • Issues: 106
  • Source available: True
  • Source changed: False
  • Source segments: 80
  • Target segments: 80

Scorecard

Dimension Score
adequacy 0.0
technical_accuracy 97.5
completeness 0.0
terminology 0.0
fluency 0.0
publishing_integrity 80.0
style_locale 60.0

Metrics

  • qe_metric: heuristic
  • qe_average: 0.9665
  • qe_min: 0.5509
  • embedding_similarity_average: 0.8944
  • embedding_similarity_min: 0.5733
  • cache_hits: 0
  • cache_misses: 160

MQM Judge

  • Enabled: True
  • Provider: openai
  • Model: gpt-5.6-luna
  • Reasoning effort: none
  • Prompt: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/judges/mqm_prompt.md
  • Prompt hash: 887d2931aa289213f0bdce4a917a0ac8364dad8e470011069b8f8bf758d69a90
  • Style guide hash: 937d8cd893578d30e716a3eb513cdf5f10d6fd3ad8f5e77068b57f96e160de12
  • Requested segments: 80
  • Evaluated segments: 78
  • MQM errors: 58
  • Cache hits: 0
  • Cache misses: 80
  • Severity counts: {'critical': 4, 'major': 13, 'minor': 41}
  • adequacy_average: 0.8831
  • technical_average: 0.9747
  • fluency_average: 0.8765
  • warning: Skipped MQM error for segment t_054: source_span is not verbatim source text.
  • warning: Skipped MQM result for segment t_054: at least one error was invalid.
  • warning: Skipped MQM error for segment p_077: target_span is not verbatim target text.
  • warning: Skipped MQM result for segment p_077: at least one error was invalid.
  • warning: MQM segment coverage is invalid: expected exactly one result for every aligned target segment.

Style Guide

  • Enabled: True
  • Guide: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/style/hf-blog-ko-translation-guide.md
  • Policy: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/configs/style_policy.yml
  • Style score: 60.0
  • Rule hits: {'information_addition': 1, 'link_text_translation': 21, 'modal_strength': 14, 'translationese': 2}

Style Guide Findings

Rule Severity Segment Current Suggested
translationese minor 에 의해 Rewrite the sentence in natural Korean.
translationese minor 을 가지 Rewrite the sentence in natural Korean.
modal_strength major p_002 공개된 음성 인식 벤치마크는 점점 더 모델이 인간 수준의 성능을 보이고 있다는 신호를 시사한다. 그러나 이러한 점수는 모델이 실제 세계에서 어떻게 작동하는지 항상 반영하지 않는다. 공개 벤치마크가 열려 있고 널리 사용되기 때문에, 모델이 테스트 자체에 맞춰 최적화될 수도 있다. 그들의 점수는 벤치마크 특유의 패턴을 학습했기 때문이 아니라 근본적인 과제를 더 잘 수행하게 되었기 때문일 수 있다. Preserve the strength of may using: 수 있습니다, 일 수 있습니다.
modal_strength major p_002 공개된 음성 인식 벤치마크는 점점 더 모델이 인간 수준의 성능을 보이고 있다는 신호를 시사한다. 그러나 이러한 점수는 모델이 실제 세계에서 어떻게 작동하는지 항상 반영하지 않는다. 공개 벤치마크가 열려 있고 널리 사용되기 때문에, 모델이 테스트 자체에 맞춰 최적화될 수도 있다. 그들의 점수는 벤치마크 특유의 패턴을 학습했기 때문이 아니라 근본적인 과제를 더 잘 수행하게 되었기 때문일 수 있다. Preserve the strength of can using: 수 있습니다.
modal_strength major p_009 이를 대규모로 테스트하기 위해, 음소 오류율(PER)이 낮은 독립적 모델들로 구성된 앙상블을 사용한다. PER은 작성된 전사가 오디오의 소리와 얼마나 가깝게 일치하는지 측정하여 모델이 들은 것을 얼마나 충실히 전사하는지에 대한 유용한 대리 지표가 된다. 앙상블의 결과는 모델들이 벤치마크의 참조 전사와 만장일치로 동의하지 않는 사례를 표시하는 데 사용할 수 있다. 그런 표시된 사례들 중 일부를 인간 주석과 비교하여 수정된 전사를 검증한다. Preserve the strength of can using: 수 있습니다.
modal_strength major t_021 nvidia/canary-qwen-2.5b | ❌ Mr President… | ❌ Mr President… | ✅ Thank you Mr. President… Preserve the strength of can using: 수 있습니다.
modal_strength major p_037 합의 불일치 탐침을 바탕으로, 테스트 데이터셋의 오디오 샘플에서 숫자를 의도적으로 침묵시키고 모델이 들리는 것을 전사하도록 한다. 숫자는 오디오에서 문자 그대로 없기 때문에, 모델은 아무 숫자도 출력하지 말아야 하며, 텍스트의 정확한 숫자까지도 출력하지 않아야 한다. Preserve the strength of should using: 좋습니다, 해야 합니다.
modal_strength major t_047 nvidia/canary-qwen-2.5b | Mr President, in the Committee on Budgets we voted on more than one amendments to the 2011 draft budget … voted in the plenary Preserve the strength of can using: 수 있습니다.
modal_strength major p_060 우리의 정자 표기 전환 탐침은 모델이 음향에서 명확하지 않더라도 벤치마크의 참조 전사에서 사용된 정확한 철자를 재현하는지 여부를 테스트한다. 정자 표기 대안은 의미적으로 같고 음성적으로도 동일한 단어들이 서로 다르게 표기되는 경우를 말한다(1 vs one, Mr. vs mister, John vs Jon, Honor vs Honour 등). 이론적으로 모델은 한 철자를 일관되게 선호하거나, 평균적으로 무작위로 번갈아가야 한다. 각 벤치마크의 참조 전사에 맞춰 특정 철자를 사용하도록 체계적으로 바뀌는 경우가 있다면, 테스트가 어떤 철자 표기를 벤치마크가 기대하는지 모델이 파악하고 있는 것을 시사한다. Preserve the strength of can using: 수 있습니다.
modal_strength major p_060 우리의 정자 표기 전환 탐침은 모델이 음향에서 명확하지 않더라도 벤치마크의 참조 전사에서 사용된 정확한 철자를 재현하는지 여부를 테스트한다. 정자 표기 대안은 의미적으로 같고 음성적으로도 동일한 단어들이 서로 다르게 표기되는 경우를 말한다(1 vs one, Mr. vs mister, John vs Jon, Honor vs Honour 등). 이론적으로 모델은 한 철자를 일관되게 선호하거나, 평균적으로 무작위로 번갈아가야 한다. 각 벤치마크의 참조 전사에 맞춰 특정 철자를 사용하도록 체계적으로 바뀌는 경우가 있다면, 테스트가 어떤 철자 표기를 벤치마크가 기대하는지 모델이 파악하고 있는 것을 시사한다. Preserve the strength of should using: 좋습니다, 해야 합니다.

Issues

QL-001 formatting / critical

  • Message: Front matter key authors changed or is missing.
  • Source: user: tzirakis
  • Target: user: tlebryk02
  • Suggested fix: Preserve front matter authors exactly.

QL-002 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Measuring benchmark optimization in speech recognition
  • Target: 음성 인식에서 벤치마크 최적화 측정
  • Suggested fix: 음성 인식 벤치마크 최적화 측정
  • Reason: 의미는 전달되지만 ‘벤치마크 최적화 측정’은 한국어 제목으로 다소 어색하고, 무엇을 측정하는지 관계가 모호합니다.

QL-003 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: Public voice AI benchmarks increasingly suggest that models are performing at human levels.
  • Target: 공개된 음성 인식 벤치마크는 점점 더 모델이 인간 수준의 성능을 보이고 있다는 신호를 시사한다.
  • Suggested fix: 공개 음성 AI 벤치마크는 모델이 점점 인간 수준의 성능을 보이고 있음을 시사합니다.
  • Reason: voice AI는 음성 인식만을 뜻하지 않을 수 있는데, 이를 ‘음성 인식’으로 한정해 원문의 기술 범위를 좁혔습니다. 또한 ‘신호를 시사한다’는 중복되고 부자연스럽습니다.

QL-004 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: Public voice AI benchmarks increasingly suggest that models are performing at human levels. Yet those scores don't always reflect how models work in the real-world.
  • Target: 공개된 음성 인식 벤치마크는 점점 더 모델이 인간 수준의 성능을 보이고 있다는 신호를 시사한다. 그러나 이러한 점수는 모델이 실제 세계에서 어떻게 작동하는지 항상 반영하지 않는다.
  • Suggested fix: 공개 음성 AI 벤치마크는 모델이 점점 인간 수준의 성능을 보이고 있음을 시사합니다. 그러나 이러한 점수가 모델의 실제 환경에서의 작동 방식을 항상 반영하는 것은 아닙니다.
  • Reason: 전체 문장이 공식 기술 블로그에 맞지 않는 평서형으로 번역되었습니다. 또한 “don't always”의 ‘항상 ~하지 않는다’는 한국어에서 모든 경우에 반영하지 않는다는 의미로 읽힐 수 있어, ‘항상 반영하는 것은 아니다’로 제한을 명확히 하는 편이 원문의 의미를 더 정확히 보존합니다.

QL-005 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Their scores may improve because they have learned benchmark-specific patterns and not because they have become better at the underlying task.
  • Target: 그들의 점수는 벤치마크 특유의 패턴을 학습했기 때문이 아니라 근본적인 과제를 더 잘 수행하게 되었기 때문일 수 있다.
  • Suggested fix: 모델의 점수가 오른 것은 기저 과제를 더 잘 수행하게 되었기 때문이 아니라, 벤치마크 특화 패턴을 학습했기 때문일 수 있습니다.
  • Reason: 모델을 가리키는 ‘그들의’는 한국어 기술 문맥에서 부자연스럽고, ‘때문이 아니라 … 때문일 수 있다’ 구조도 원문의 대조를 명확하게 전달하지 못합니다. 원문은 점수 향상의 원인이 벤치마크 특화 패턴 학습일 수 있으며, 기저 과제 수행 능력 향상 때문은 아닐 수 있다는 뜻입니다.

QL-006 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: traditional benchmarks overlook many of the conditions and qualities that make voice systems reliable, natural, contextually appropriate, and effective in practice
  • Target: 전통적인 벤치마크가 음성 시스템을 실세계에서 신뢰할 수 있게 만드는 조건과 특성을 충분히 반영하지 못하기 때문입니다.
  • Suggested fix: 전통적인 벤치마크는 실제 환경에서 음성 시스템을 신뢰할 수 있고, 자연스러우며, 맥락에 적절하고, 효과적으로 작동하게 하는 많은 조건과 특성을 간과합니다.
  • Reason: 원문의 핵심 수식어인 ‘많은(many)’이 빠졌고, reliable, natural, contextually appropriate, and effective라는 네 가지 품질 중 ‘자연스럽고’, ‘맥락에 적절하며’, ‘실제로 효과적인’이라는 내용이 누락되었습니다. 이로 인해 전통적인 벤치마크가 간과하는 평가 요소의 범위와 의미가 축소됩니다.

QL-007 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: That's why we recently introduced held-out sets in Real World VoiceEQ, the Open-ASR Leaderboard, and the Far-field ASR Leaderboard: to measure more of what matters in real-world use.
  • Target: 그래서 우리는 최근 Real World VoiceEQ의 보류 평가셋, Open-ASR Leaderboard 및 Far-field ASR Leaderboard를 도입했습니다: 실제 사용에서 더 중요한 것을 더 많이 측정하기 위함입니다.
  • Suggested fix: 그래서 최근 Real World VoiceEQ, Open-ASR Leaderboard, Far-field ASR Leaderboard에 홀드아웃 세트를 도입했습니다. 실제 환경에서 중요한 요소를 더 많이 측정하기 위해서입니다.
  • Reason: 콜론 뒤의 목적 표현을 ‘~하기 위함입니다’로 직역해 문장이 딱딱하고 부자연스럽습니다. 또한 ‘more of what matters’는 ‘더 중요한 것’이 아니라 ‘중요한 요소를 더 많이’라는 의미입니다.

QL-008 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Our latest research introduces three tests to help quantify it. We evaluated 11 widely used open-source ASR models and found that several of the highest-scoring systems reproduced benchmark transcripts from the VoxPopuli English and LibriSpeech (clean, other) datasets – even when the audio contradicted them, relevant words had been silenced, or the audio equally supported two different written forms.
  • Target: 최신 연구는 이를 정량화하는 데 도움이 되는 세 가지 테스트를 도입했다. 우리는 널리 사용되는 11종의 오픈 소스 ASR 모델을 평가했고, 여러 고득점 시스템이 VoxPopuli 영어 및 LibriSpeech(clean, other) 데이터셋의 벤치마크 트랜스크립트를 재현했다는 것을 발견했다 — 오디오가 이를 반박했고, 관련 단어가 제거되었거나, 오디오가 두 가지 서로 다른 표기를 균등하게 지지했을 때도 말이다.
  • Suggested fix: 최신 연구에서는 이를 정량화하는 데 도움이 되는 세 가지 테스트를 소개합니다. 널리 사용되는 오픈 소스 ASR 모델 11종을 평가한 결과, 가장 높은 점수를 받은 시스템 중 여러 모델이 VoxPopuli 영어 및 LibriSpeech(clean, other) 데이터셋의 벤치마크 전사를 재현하는 것으로 나타났습니다. 오디오가 해당 전사와 일치하지 않거나, 관련 단어가 무음 처리되었거나, 오디오만으로는 서로 다른 두 표기 중 어느 쪽인지 구분하기 어려운 경우에도 마찬가지였습니다.
  • Reason: 전체 의미는 전달되지만 ‘도입했다’, ‘발견했다는 것을’, 대시 뒤의 ‘말이다’가 영어 문장 구조를 강하게 따라가 기술 블로그 문체로는 다소 부자연스럽습니다. 또한 음성이 단어를 무음 처리했다는 의미의 ‘had been silenced’를 ‘제거되었거나’로 옮겨 정보가 약간 달라졌습니다.

QL-009 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: their scores overstated how well they could transcribe speech more generally
  • Target: 점수는 일반적으로 음성을 더 충실히 전사하는 능력에 대해 과대평가될 수 있다
  • Suggested fix: 그 결과 점수는 일반적인 음성 전사 능력을 실제보다 과대평가했다.
  • Reason: 원문의 ‘scores overstated’는 점수가 실제 능력을 과대평가했다는 단정적 서술인데, ‘과대평가될 수 있다’로 옮겨 가능성으로 약화되었습니다. 또한 ‘how well’은 전사 능력의 정도를 뜻하므로 ‘더 충실히’는 원문에 없는 비교 의미를 추가합니다.

QL-010 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: Reference disagreement (VoxPopuli case study)
  • Target: 참조 불일치( VoxPopuli 사례 연구)
  • Suggested fix: 참조 불일치(VoxPopuli 사례 연구)
  • Reason: 괄호 앞에 불필요한 공백이 있어 한국어 제목 표기상 어색합니다. 의미와 기술명은 보존되었습니다.

QL-011 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: VoxPopuli is known to contain a high number of transcription errors (which is why Artificial Analysis released a cleaned version).
  • Target: VoxPopuli는 다수의 전사 오류를 포함하는 것으로 알려져 있다(그래서 Artificial Analysis가 cleaned version를 발표했다).
  • Suggested fix: VoxPopuli에는 전사 오류가 많이 포함된 것으로 알려져 있습니다(그래서 Artificial Analysis가 정제된 버전을 공개했습니다).
  • Reason: 문장 전체가 평서형 반말체인 ‘-다’로 번역되어 주변 기술 블로그의 존댓말 문체와 어긋납니다. 또한 ‘cleaned version’를 영어로 남기면서 한국어 조사 ‘를’이 부자연스럽게 결합되었습니다.

QL-012 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Our consensus disagreement probe tests what happens when leading ASR models encounter these errors: *Do they accurately transcribe what the audio says, or reproduce the benchmark's incorrect reference transcript?*
  • Target: 우리의 합의 불일치 프로브는 선도하는 ASR 모델들이 이러한 오류에 직면했을 때 어떤 일이 벌어지는지 테스트한다: *오디오가 실제로 말하는 내용을 정확히 전사하는가, 아니면 벤치마크의 잘못된 참조 전사를 재생산하는가?*
  • Suggested fix: 저희의 합의 불일치 프로브는 주요 ASR 모델이 이러한 오류를 접했을 때 어떤 결과가 나타나는지 확인합니다. 즉, 오디오에서 실제로 말하는 내용을 정확히 전사하는지, 아니면 벤치마크의 잘못된 참조 전사를 그대로 재현하는지를 살펴봅니다.
  • Reason: ‘선도하는’, ‘테스트한다’, ‘재생산하는가’가 영어식 구조와 직역투를 남기며, 콜론 뒤 질문도 한국어 문장 흐름에 비해 기계적으로 연결되어 있습니다.

QL-013 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: The ensemble results can be used to flag cases in which the models unanimously disagree with the benchmark's reference transcript.
  • Target: 앙상블의 결과는 모델들이 벤치마크의 참조 전사와 만장일치로 동의하지 않는 사례를 표시하는 데 사용할 수 있다.
  • Suggested fix: 앙상블 결과를 사용해 모델들이 벤치마크의 기준 전사와 모두 불일치하는 사례를 표시할 수 있습니다.
  • Reason: 원문의 can은 가능성을 나타내는데, 목표 문장의 ‘사용할 수 있다’는 의미상 대응하지만 문단 전체가 일반적인 합니다체가 아닌 하게체로 번역되어 문체가 일관되지 않습니다. 또한 ‘만장일치로 동의하지 않는’은 이해 가능하지만 기술 문맥에서 모델들이 모두 기준 전사와 불일치하는 사례라는 의미가 더 명확해야 합니다.

QL-014 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: We then compare a sample of those flagged cases against human annotations to validate the corrected transcripts.
  • Target: 그런 표시된 사례들 중 일부를 인간 주석과 비교하여 수정된 전사를 검증한다.
  • Suggested fix: 그다음 식별된 사례 중 일부를 사람의 주석과 비교해 수정된 전사를 검증합니다.
  • Reason: ‘그런 표시된 사례들’과 ‘인간 주석’은 직역투이고, 문단 전체의 종결형과도 어울리지 않습니다. flagged cases는 ‘표시된 사례’보다 ‘플래그된 사례’ 또는 ‘식별된 사례’, human annotations는 ‘사람이 작성한 주석’으로 옮기면 의미와 가독성이 더 분명합니다.

QL-015 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Six of the 11 models we tested reproduced the benchmark's erroneous transcript—giving the "expected" answer even though it contradicted the audio. On the real clip, the formatting follows the same pattern: models that omit "Thank you" also reproduce the benchmark's punctuation style, writing "Mr" without a period, while models that include the audible phrase tend to write "Mr." with the period.
  • Target: 우리가 테스트한 11개 모델 중 여섯 개는 벤치마크의 잘못된 전사를 재현했다 — 오디오와 모순되더라도 "예상된" 정답을 제시했다. 실제 클립에서도 형식은 같은 패턴을 따른다: "Thank you"를 생략한 모델은 또한 벤치마크의 구두점 스타일을 재현하고, "Mr"에 마침표가 없고, 들리는 구절을 포함하는 모델은 보통 "Mr."에 마침표를 찍는 경향이 있다.
  • Suggested fix: 우리가 테스트한 11개 모델 중 6개는 벤치마크의 잘못된 전사를 재현했습니다. 오디오와 모순되더라도 "예상된" 답을 제시한 것입니다. 실제 클립에서도 형식은 같은 패턴을 보입니다. "Thank you"를 생략한 모델은 벤치마크의 구두점 스타일을 따라 마침표 없이 "Mr"라고 쓰는 반면, 들리는 구절을 포함한 모델은 대체로 마침표가 있는 "Mr."라고 씁니다.
  • Reason: 대시는 앞뒤 공백과 함께 사용되었고, 콜론 뒤의 병렬 구조가 한국어 문장으로 자연스럽게 연결되지 않습니다. 특히 "Mr"와 "Mr."의 표기 차이를 설명하는 부분이 문법적으로 어색해 독자가 패턴을 한 번에 파악하기 어렵습니다.

QL-016 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: When we present the same content in newly collected voices from EU parliamentary recordings or generic voices, this behavior often weakens or disappears.
  • Target: 유럽 의회의 녹음이나 일반 음성으로 새로 수집된 음성을 제시하면 이 행동은 종종 약해지거나 사라지는 경우가 많다.
  • Suggested fix: EU 의회 녹음이나 일반 음성에서 새로 수집한 음성으로 동일한 콘텐츠를 제시하면 이 행동은 종종 약해지거나 사라진다.
  • Reason: 원문의 'the same content'가 번역에서 빠져, 무엇을 새로 수집된 음성으로 제시하는지 정보가 불분명해졌습니다. 또한 'often'과 '종종 ... 경우가 많다'가 중복되어 다소 어색합니다.

QL-017 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: for a clone of a new parliamentary recording
  • Target: 새 의회 녹음의 클론에 대해
  • Suggested fix: 새 의회 녹음을 복제한 음성에 대해
  • Reason: 원문의 'clone'은 새 의회 녹음을 복제한 음성/클론을 가리키는데, 현재 번역은 '녹음의 클론'으로 직역되어 음성 합성 문맥에서 대상이 모호합니다.

QL-018 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: The audio in all three clips below actually says the same thing, preceded by an audible "Thank you,"—the clones are text-to-speech renditions of that true sentence, so the courtesy is audible in all three.
  • Target: 이 아래의 세 클립의 오디오도 실제로는 같은 내용을 말하며 앞에 들리는 "Thank you,"—클론은 그 참된 문장의 음성 합성 버전이므로 세 클립 모두 예의 표현이 들린다.
  • Suggested fix: 아래 세 클립의 오디오는 실제로 같은 내용을 말하며, 그 앞에는 들을 수 있는 "Thank you,"가 붙습니다. 클론은 이 실제 문장을 텍스트 음성 변환으로 재현한 것이므로 세 클립 모두에서 이 예의 표현을 들을 수 있습니다.
  • Reason: 원문의 핵심 구조인 ‘세 클립의 오디오는 실제 문장 앞에 들리는 “Thank you,”가 붙어 같은 내용을 말한다’는 관계가 불명확하게 번역되었습니다. 특히 ‘앞에 들리는’의 주체와 연결이 어색하고, 대시 뒤 설명도 문장으로 자연스럽게 이어지지 않아 독자가 오디오 내용과 클론의 관계를 정확히 파악하기 어렵습니다.

QL-019 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: Green highlighting and ✅ mark a transcript that includes the audible "Thank you"; red highlighting and ❌ mark a transcript that reproduces the benchmark's erroneous omission.
  • Target: 초록 하이라이트와 ✅ 표시는 들리는 "Thank you"를 포함하는 전사를; 빨간 하이라이트와 ❌ 표시는 벤치마크의 잘못된 생략을 재현하는 전사를 나타낸다.
  • Suggested fix: 초록색 강조 표시와 ✅는 들리는 "Thank you"를 포함한 전사를 나타냅니다. 빨간색 강조 표시와 ❌는 벤치마크에서 잘못 생략된 부분을 그대로 재현한 전사를 나타냅니다.
  • Reason: 두 절을 세미콜론으로 연결한 구조가 한국어에서 비문에 가깝고, ‘전사를; 빨간’처럼 조사가 문장 연결을 방해합니다. 의미는 대체로 보존되지만 출판용 문장으로는 명확히 다듬어야 합니다.

QL-020 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.6800
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-021 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: Drops the courtesy (❌) out of 11
  • Target: 11개 중 예의를 제외
  • Suggested fix: 11개 중 예의 항목 제외 (❌)
  • Reason: 원문의 ‘courtesy’는 문맥상 ‘예의’라는 추상적 개념을 제외한다는 뜻이 아니라, 11개 항목 중 예의 관련 항목을 제거하거나 제외한다는 의미입니다. 현재 번역은 무엇을 제외하는지 불명확하고, 괄호 안의 ❌ 기호도 누락되어 표의 라벨 정보가 손실되었습니다.

QL-022 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: Parakeet is the only model that flips between reproducing the benchmark on the real clip and getting it right on the same-speaker clone. Phi-4 is the only model still dropping the courtesy on the ep-fresh clone. When we instead resynthesize the sentence in a generic TTS voice unconnected to any parliamentary recording, all eleven models restore the courtesy.
  • Target: Parakeet은 실제 클립에서 벤치마크를 재현하는 것과 같은 화자 클론에서 정확하게 재현하는 것 사이를 오가는 유일한 모델이다. Phi-4는 ep-fresh 클론에서 여전히 예의를 생략하는 유일한 모델이다. 대신 의회 녹음과 연결되지 않은 일반 TTS 보이스로 문장을 재합성하면 모든 열한 모델이 예의를 회복한다.
  • Suggested fix: 문장 종결을 존댓말로 통일합니다. 예: "Parakeet은 실제 클립에서 벤치마크를 재현할 때와 같은 화자 클론에서 이를 정확하게 재현할 때 사이를 오가는 유일한 모델입니다. Phi-4는 ep-fresh 클론에서 여전히 예의를 생략하는 유일한 모델입니다. 대신 의회 녹음과 연결되지 않은 일반 TTS 음성으로 문장을 재합성하면 열한 모델 모두 예의를 회복합니다."
  • Reason: 기술적 의미와 수치는 대체로 보존되었지만 전체 문장이 공식 기술 블로그의 존댓말이 아닌 평서형 반말체로 끝나 문체가 일관되지 않습니다.

QL-023 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: The results suggest that this problem is both widespread and meaningful. Our methodology flagged potential reference errors in 40% of the VoxPopuli test clips we analyzed, affecting roughly 3% of all reference words.
  • Target: 이 결과는 이 문제가 널리 퍼져 있으며 의미가 있음을 시사한다. 우리의 방법론은 분석한 VoxPopuli 테스트 클립의 40%에서 참조 오류 가능성을 표시했고, 전체 참조 단어의 약 3%에 영향을 미쳤다.
  • Suggested fix: 이 결과는 이 문제가 널리 퍼져 있으며 의미가 있음을 시사합니다. 저희 방법론은 분석한 VoxPopuli 테스트 클립의 40%에서 참조 오류 가능성을 표시했고, 전체 참조 단어의 약 3%에 영향을 미쳤습니다.
  • Reason: 의미와 수치, 불확실성은 보존했지만 전체 문장이 해라체로 번역되어 기술 블로그의 기본 존댓말 문체와 맞지 않습니다.

QL-024 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Models exhibiting benchmark-optimized behavior reproduced erroneous reference transcripts 18–30% of the time.
  • Target: 벤치마크 최적화 현상을 보인 모델은 잘못된 참조 전사를 18–30%의 시점에서 재현했다.
  • Suggested fix: 벤치마크에 최적화된 행동을 보이는 모델은 잘못된 참조 전사를 18~30%의 비율로 재현했습니다.
  • Reason: ‘18–30%의 시점에서’는 빈도를 나타내는 영어 표현을 부자연스럽게 옮긴 표현입니다. 또한 문단의 존댓말 문체와 달리 ‘재현했다’가 혼용되어 문체가 일관되지 않습니다.

QL-025 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: The models with the lowest WER—and therefore the strongest reported benchmark performance—are also the most likely to reproduce these errors.
  • Target: 가장 낮은 WER를 보이는 모델들—따라서 가장 강하게 보고되는 벤치마크 성능—도 이러한 오류를 재현할 가능성이 가장 높다.
  • Suggested fix: WER가 가장 낮은 모델, 즉 보고된 벤치마크 성능이 가장 우수한 모델일수록 이러한 오류를 재현할 가능성도 가장 높습니다.
  • Reason: 원문의 ‘strongest reported benchmark performance’는 ‘보고된 벤치마크 성능이 가장 우수한’이라는 뜻인데, ‘가장 강하게 보고되는’으로 옮겨 성능의 우수성이 아니라 보고 방식의 강도로 읽힐 수 있습니다. 문장 종결도 앞 문장과 달리 평서형으로 끝나 문체가 불일치합니다.

QL-026 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: alt="Scatterplot comparing VoxPopuli WER to the rate at which each model reproduces the benchmark's incorrect reference transcript."
  • Target: alt="Scatterplot comparing VoxPopuli WER to the rate at which each model reproduces the benchmark's incorrect reference transcript."
  • Suggested fix: alt="VoxPopuli WER과 각 모델이 벤치마크의 잘못된 참조 전사문을 재현하는 비율을 비교한 산점도"
  • Reason: 이미지 alt text가 영어로 남아 있습니다. 이미지 경로와 HTML 구조는 올바르게 유지했지만, 가이드에 따라 설명용 alt text는 한국어로 번역해야 합니다.

QL-027 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: we deliberately silence numbers in the audio samples of test datasets and ask the models to transcribe what it hears.
  • Target: 테스트 데이터셋의 오디오 샘플에서 숫자를 의도적으로 침묵시키고 모델이 들리는 것을 전사하도록 한다.
  • Suggested fix: 테스트 데이터셋의 오디오 샘플에서 숫자 부분을 의도적으로 음소거한 뒤, 모델에 들리는 내용을 전사하도록 요청합니다.
  • Reason: 공식 기술 블로그 문맥에서 문장 종결이 ‘한다’로 되어 있고, ‘숫자를 침묵시키다’와 ‘모델이 들리는 것’은 영어 구조를 직역한 표현이라 다소 부자연스럽습니다. 의미는 대체로 보존되었습니다.

QL-028 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: much less the exact number in the text.
  • Target: 텍스트의 정확한 숫자까지도 출력하지 않아야 한다.
  • Suggested fix: 하물며 텍스트에 있는 정확한 숫자를 출력해서는 안 됩니다.
  • Reason: ‘much less’는 앞의 주장보다 더 강하게 불가능하다는 강조인데, 현재 번역은 ‘까지도’로 옮겨 핵심 대비는 전달하지만 강조 관계가 충분히 자연스럽게 드러나지 않습니다.

QL-029 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: (although still unlikely for a model to predict)
  • Target: (모델이 예측하기는 어렵다)
  • Suggested fix: (모델이 예측할 가능성은 여전히 낮지만)
  • Reason: 원문의 ‘여전히 가능성이 낮다’는 표현이 ‘예측하기 어렵다’로 옮겨져, 사건의 발생 가능성에 대한 의미가 모델의 난이도 평가로 바뀌었습니다.

QL-030 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: Some of these numbers are semi-predictable (although still unlikely for a model to predict), yet others are quite surprising.
  • Target: 이 숫자들 중 일부는 부분적으로 예측 가능하지만(모델이 예측하기는 어렵다), 다른 숫자는 꽤 놀랍다.
  • Suggested fix: 이 숫자 중 일부는 부분적으로 예측할 수 있지만(모델이 예측할 가능성은 여전히 낮습니다), 다른 숫자는 꽤 놀랍습니다.
  • Reason: 문단 전체가 해라체로 번역되어 기술 블로그의 기본 존댓말과 어조가 맞지 않습니다.

QL-031 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: green highlighting with strikethrough marks reference-transcript words the model correctly did not reproduce (audio-faithful);
  • Target: 초록 하이라이트에 취소선은 모델이 정확히 재현하지 않은 참조 전사를 나타낸다(음향에 충실함);
  • Suggested fix: 초록색 하이라이트와 취소선 표시는 모델이 오디오에 충실하게 정확히 재현하지 않은 참조 전사 단어를 나타냅니다;
  • Reason: 원문은 ‘초록색 하이라이트와 취소선 표시가 ... 단어를 나타낸다’는 뜻인데, 대상 문장에서는 ‘초록 하이라이트에 취소선’이 비문에 가까워 두 시각적 표시의 관계가 불명확합니다. 또한 ‘reference-transcript words’를 ‘참조 전사’로만 옮겨 한국어 기술 문맥에서 단어를 수식하는 구조가 어색하고, ‘audio-faithful’도 ‘음향에 충실함’으로 번역되어 의미가 불분명합니다.

QL-032 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.7000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-033 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: green highlighting with <u>underline</u> marks a correct, audio-faithful insertion in place of the reference's erroneous wording;
  • Target: 초록 하이라이트에 <u>밑줄</u>은 참조의 잘못된 표현을 대체한 올바르고 음향에 충실한 삽입을 표시;
  • Suggested fix: 초록색 하이라이트와 밑줄은 참조의 잘못된 표현을 대신하는, 오디오에 충실한 올바른 삽입을 나타냅니다.
  • Reason: ‘초록 하이라이트에 밑줄은’은 조사와 주어 구조가 어색해 의미를 즉시 파악하기 어렵습니다. 또한 ‘음향에 충실한’은 audio-faithful의 수식 관계를 부자연스럽게 옮겼습니다.

QL-034 accuracy / major

  • Message: MQM judge adequacy score is low.
  • Target: 0.7200
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM adequacy score is below threshold 0.75.

QL-035 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: red highlighting (plain text) reproduces the reference transcript's erroneous, audio-unsupported content: keeping "Mr President", writing "more than 1 amendments" where the audio says "one thousand six hundred", supplying the silenced year "2011", or ending on "plenary".
  • Target: 빨간 하이라이트(일반 텍스트)는 참조 전사의 잘못되고 음향에 의해 뒷받침되지 않는 내용을 재현한다: 예를 들어 "Mr President"를 유지하고, 음성에서 말하는 "one thousand six hundred" 대신 "more than 1 amendments"를 쓰며, 침묵된 연도 "2011"을 제공하거나, "plenary"로 끝내는 것 등.
  • Suggested fix: 빨간색 강조(일반 텍스트)는 참조 전사에 포함된, 오디오로는 뒷받침되지 않는 잘못된 내용을 재현합니다. 예를 들어 "Mr President"를 그대로 유지하거나, 오디오에서 "one thousand six hundred"라고 말한 부분을 "more than 1 amendments"로 적거나, 음성에 나오지 않은 연도 "2011"을 넣거나, "plenary"에서 끝내는 경우입니다.
  • Reason: 의미는 대체로 보존되지만, "음향에 의해 뒷받침되지 않는"과 "제공하거나"가 기술 블로그의 한국어 문맥에서 다소 직역투이고 부자연스럽습니다. 콜론 뒤의 나열도 한국어 문장으로 매끄럽게 연결되지 않습니다.

QL-036 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: `Reference | Mr President, in the Committee on Budgets, we voted on more than 1 amendments to the <span style="background-color:#fee2e
SEO report — ✅ Pass

SEO Eval Report

Gate: ✅ PASS — deterministic AND rubric

  • File: ../target/_posts/2026-08-21-asr-benchmark-optimization.md
  • Source: —
  • Primary keyword: (none — D5 skipped)
  • Mode: file

Gate

  • Status: PASS
  • Blockers: ✅ pass
  • Deterministic REQUIRED (D1–D7): ✅ pass
  • Rubric (R1–R6): ✅ pass (mean None, min None)

Blockers

✅ body_not_empty: Body is not empty
✅ robots_indexable: Robots is indexable
✅ internal_links_resolve: All internal links resolve
✅ local_images_resolve: All local images resolve

Required checks (gated)

✅ heading_hierarchy: Heading hierarchy: Valid
✅ alt_text_coverage: Alt text coverage: 5/5 images
✅ descriptive_alt_text: Descriptive alt text: 5/5 (≥80% recommend)
✅ image_files_exist: All 0 local image file(s) exist

OpenAI rubric checks

✅ semantic_metadata: PASS (required) — Semantic alignment across title, H1, headings, and opening; translation notice confirms origin; description absence ignored.
✅ alt_semantics: PASS (review) — https://huggingface.co/datasets/HumeAI/hf-assets/resolve/main/blog/asr-benchmark-optimization/wer_vs_badref.png: Clear chart description: scatterplot showing WER vs reproduction rate.; https://huggingface.co/datasets/HumeAI/hf-assets/resolve/main/blog/asr-benchmark-optimization/masking_freshpairs.png: Descriptive comparison of recovery rate across benchmarks and held-out audio.; https://huggingface.co/datasets/HumeAI/hf-assets/resolve/main/blog/asr-benchmark-optimization/pair_spacing_sorted.png: Specifies the spacing convention analyzed and that results are sorted by model.

Advisory checks (not gated)

✅ opening_summary: Opening 3 paragraphs: 405 chars (recommend ≥150 for KO/GEO)
✅ h1_count: Markdown H1 count: 1 (review against rendered layout)
✅ citations: Citations/statistics: 42 (recommend ≥1 for GEO)
⚠️ question_headings: Scannable H2/H3 (question or keyword): 0 (0 question, 0 keyword)
⚠️ internal_links: Internal links: 0 (recommend 2-3)
✅ word_count: Word count: 2287 (recommend ≥300)
ℹ️ primary_keyword: No primary_keyword in manifest — keyword check skipped
⚠️ webp_format: WebP format: 0/5 images (≥50% recommend)
⚠️ lazy_loading: Lazy loading: 0 images (optional)

Signals (evidence — not directly gated)

  • Frontmatter: title 19 chars, description 0 chars, author present True
  • Title text: 음성 인식에서 벤치마크 최적화 측정
  • Description text: —
  • Opening text: * TOC
    {:toc}

이 글은 Hugging Face 블로그의 Measuring benchmark optimization in speech recognition를 한국어로 번역한 글입니다.

  • Opening: first paragraph 177 chars, first 3 paragraphs 405 chars
  • Headings: markdown H1 1, rendered effective H1 2, layout title H1 True
  • Links: total 33, external 33, internal 0, citation signals 42
  • Images: total 5, empty alt 0, filename-like alt 0, missing local files 0

Semantic review packet

  • Title: 음성 인식에서 벤치마크 최적화 측정
  • Description: —
  • Rendered H1 candidates: 음성 인식에서 벤치마크 최적화 측정, 음성 인식에서 벤치마크 최적화 측정
  • Opening: * TOC
    {:toc}

이 글은 Hugging Face 블로그의 Measuring benchmark optimization in speech recognition를 한국어로 번역한 글입니다.

  • Canonical/permalink: —
  • Instruction: Compare title, description, rendered H1, and opening text for meaning consistency. This packet is evidence only; it does not decide pass/fail.

Frontmatter (advisory — written by metadata step, not gated)

✅ title: Title: 19 chars (recommend ≤60)
❌ description: Description is missing
✅ image: OG image: assets/images/blog/posts/2026-08-21-asr-benchmark-optimization/thumbnail.png
✅ categories: Categories: 2 (recommend 2-3)
✅ author: Author: dailybot

SEO metadata suggestion — PARTIAL

This is a suggestion. SEO is applied only when the post frontmatter is updated.
To apply safe fields from a partial suggestion, leave a trusted PR comment: metadata apply.

  • Auto apply: False
  • Requires human: True
  • Mode: frontmatter_only
  • Reason: metadata candidate needs policy decisions or missing title/description

Candidate

  • title: 음성 인식에서 벤치마크 최적화 측정
  • categories: ['Translation', 'HuggingFace']
  • image: assets/images/blog/posts/2026-08-21-asr-benchmark-optimization/thumbnail.png

Needs policy decision

  • target_url
  • source_url
  • canonical_policy
  • translation_indexing
  • target_locale
  • source_locale

Warnings

  • description is empty in frontmatter
  • source_url is empty
  • primary_keyword is empty

@Jwaminju Jwaminju added the hf-agent:needs-human HF Agent needs human follow-up label Aug 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hf-agent:managed Opt PR into HF Agent review automation hf-agent:needs-human HF Agent needs human follow-up

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant