Skip to content

Translate Hugging Face blog post: What We Learned by Reproducing 2,200 papers from ICML - #194

Open
Jwaminju wants to merge 1 commit into
mainfrom
translate/icml-2026-open-reproductions
Open

Translate Hugging Face blog post: What We Learned by Reproducing 2,200 papers from ICML#194
Jwaminju wants to merge 1 commit into
mainfrom
translate/icml-2026-open-reproductions

Conversation

@Jwaminju

Copy link
Copy Markdown
Collaborator

Source: https://huggingface.co/blog/icml-2026-open-reproductions

This PR adds a Korean translation draft for icml-2026-open-reproductions.

Downstream handoff:

  • SEO review should use the translation-flow manifest.
  • Quality review should use the translation-flow manifest.

@Jwaminju Jwaminju added the hf-agent:managed Opt PR into HF Agent review automation label Aug 14, 2026
@github-actions

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

🚀 View preview at
https://hugging-face-krew.github.io/pr-preview/pr-194/

Built to branch gh-pages at 2026-08-14 02:44 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@Jwaminju

Jwaminju commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator Author

HF Agent Review

Gate Result
Quality ❌ Fail
SEO ✅ Pass

Head SHA: f40ad5f86a3f2e88d836fe2824e8ea8d6a510ac4

Quality report — ❌ Fail

Quality Report

  • Status: reject
  • Quality Score: 0.0
  • Hard failures: 1
  • Issues: 108
  • Source available: True
  • Source changed: False
  • Source segments: 55
  • Target segments: 55

Scorecard

Dimension Score
adequacy 0.0
technical_accuracy 40.0
completeness 0.0
terminology 0.0
fluency 0.0
publishing_integrity 40.0
style_locale 60.0

Metrics

  • qe_metric: heuristic
  • qe_average: 0.9801
  • qe_min: 0.8833
  • embedding_similarity_average: 0.8281
  • embedding_similarity_min: 0.7352
  • cache_hits: 0
  • cache_misses: 110

MQM Judge

  • Enabled: True
  • Provider: openai
  • Model: gpt-5.6-luna
  • Reasoning effort: none
  • Prompt: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/judges/mqm_prompt.md
  • Prompt hash: 887d2931aa289213f0bdce4a917a0ac8364dad8e470011069b8f8bf758d69a90
  • Style guide hash: 937d8cd893578d30e716a3eb513cdf5f10d6fd3ad8f5e77068b57f96e160de12
  • Requested segments: 55
  • Evaluated segments: 53
  • MQM errors: 68
  • Cache hits: 0
  • Cache misses: 55
  • Severity counts: {'major': 24, 'minor': 44}
  • adequacy_average: 0.8175
  • technical_average: 0.9545
  • fluency_average: 0.7749
  • warning: Skipped MQM error for segment p_030: source_span is not verbatim source text.
  • warning: Skipped MQM result for segment p_030: at least one error was invalid.
  • warning: Skipped MQM error for segment p_039: source_span is not verbatim source text.
  • warning: Skipped MQM result for segment p_039: at least one error was invalid.
  • warning: MQM segment coverage is invalid: expected exactly one result for every aligned target segment.

Style Guide

  • Enabled: True
  • Guide: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/style/hf-blog-ko-translation-guide.md
  • Policy: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/configs/style_policy.yml
  • Style score: 60.0
  • Rule hits: {'alt_text_caption': 3, 'first_mention_bilingual': 1, 'link_text_translation': 12, 'list_consistency': 1, 'modal_strength': 6, 'overstatement': 1, 'translationese': 1}

Style Guide Findings

Rule Severity Segment Current Suggested
translationese minor 에 의해 Rewrite the sentence in natural Korean.
list_consistency minor sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, sentence, phrase Use either sentence-style endings or phrase-style endings consistently within one list.
modal_strength major h_004 더 많은 논문들로 누구도 모두 검토할 수 없었다 Preserve the strength of can using: 수 있습니다.
modal_strength major p_006 심사 역량은 그에 비례해 두 배로 늘어나지 않았다. 대부분의 학회에서 심사자는 자원봉사자이며 논문을 충분히 심사할 시간이나 전문 지식을 갖고 있지 않을 수 있다. 아래는 ICML 2026의 한 주목 논문에 대한 심사의, 심사자의 직접 발언이다: Preserve the strength of may using: 수 있습니다, 일 수 있습니다.
modal_strength major p_010 다만 바뀐 점은 제출의 흐름을 일으키는 동일한 기술이 이를 따라잡는 데에도 도움을 준다는 것이다. Claude Code, Codex, Cursor, Pi 같은 코딩 에이전트가 이제 논문을 읽고, 코드를 작성하고, 실험을 실행하며, 자신이 발견한 것을 보고할 수 있다. 논문을 꼼꼼히 검토하는 데는 리뷰어의 주말이 필요하던 시절이 있었다; 이제 에이전트는 오후 한두 시간에 시도하고, 병렬로 수천 번 실행할 수 있다. Preserve the strength of can using: 수 있습니다.
modal_strength major p_013 우리가 직접 논문을 심사하기보다는, 전체 커뮤니티에 개방했고, 에이전트 프레임워크의 다양성, 컴퓨트 예산, 과학적 취향이 어우러진 다양한 구성원을 포용했습니다. 2026년 7월 15일부터 8월 2일까지 ICML 2026 Open Reproductions challenge은 아래와 같이 작동했습니다: Preserve the strength of up to using: 최대.
modal_strength major p_030 검토된 논문 중 23%(496편)는 하나 이상의 주장을 반증하거나 논쟁의 여지가 있었습니다. 여기에는 모든 주장이 반박되어 확인할 수 없었던 49편이 포함되며, 어쩌면 가장 흥미로운 점은 독립 재현 팀이 같은 주장에 대해 서로 상반된 판단에 도달한 242편이 있었다는 점입니다. 재현성은 이분법이 아니라 대립적(adversarial)입니다. Preserve the strength of may using: 수 있습니다, 일 수 있습니다.
modal_strength major p_052 우리는 인간 심사자로서의 역할이 무엇인지요? 우리의 생각은 지능을 효과적으로 관리하는 일이라고 봅니다. 교수나 주요 연구책임자(PI)가 컴퓨트, 하드웨어, 데이터 접근, 그리고 시점에 맞춘 피드백으로 대학원생들이 좋은 연구를 할 수 있는 환경을 만드는 것처럼, 에이전트에서 최대의 성과를 얻은 참가자들은 올바른 환경을 만들고 적절한 질문을 던지는 뒤 에이전트가 실행하도록 한 사람들이다. Preserve the strength of can using: 수 있습니다.
overstatement major p_013 까지 Use a weaker expression that preserves the source claim strength.
alt_text_caption minor An OpenReview reviewer admitting the proofs were not checked carefully Translate image alt text while preserving the image path.

Issues

QL-001 formatting / critical

  • Message: link target mismatch.
  • Source: https://huggingface.co/spaces/stresearch-dev/63430
  • Suggested fix: Preserve source link target exactly.
  • Reason: Hard gate exact-match validator failed: missing=['https://huggingface.co/spaces/stresearch-dev/63430']

QL-002 technical / major

  • Message: number/unit token mismatch.
  • Target: 0, 1, 3, 3, 7, 7, 7, 8
  • Suggested fix: Preserve source number/unit token exactly.
  • Reason: Review gate exact-match validator failed: extra=['0', '1', '3', '3', '7', '7', '7', '8']

QL-003 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: we ran a hackathon
  • Target: 우리는 hackathon를 운영했고
  • Suggested fix: 해커톤을 개최했고
  • Reason: ‘hackathon’를 외래어 표기 없이 영어 명사로 그대로 사용했고, ‘hackathon를’은 조사 결합도 부자연스럽습니다. 문맥상 해커톤을 개최했다는 의미가 자연스럽습니다.

QL-004 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: 6,816 Trackio logbooks
  • Target: 6,816 Trackio logbooks
  • Suggested fix: Trackio 로그북 6,816개를 게시하며
  • Reason: generic technical term인 ‘logbooks’가 번역되지 않아 한국어 문장 안에서 다소 어색합니다. 제품명 Trackio는 유지하되 일반 명사는 번역하는 편이 명확합니다.

QL-005 formatting / major

  • Message: MQM judge reported formatting issue.
  • Source: In this post, we're sharing what we learned from running this hackathon, and what it suggests about _the role humans will play_ when agents are doing the research experiments.
  • Target: 이 글에서는 이 해커톤을 운영하면서 배운 점과, 에이전트가 연구 실험을 수행할 때 인간이 맡게 될 역할에 대해 시사하는 바를 공유합니다 _인간이 맡게 될 역할_.
  • Suggested fix: 이 글에서는 이 해커톤을 운영하면서 배운 점과 에이전트가 연구 실험을 수행할 때 인간이 맡게 될 역할이 무엇을 시사하는지 공유합니다.
  • Reason: 원문의 강조 구문 _the role humans will play_가 이미 자연스럽게 번역된 뒤 _인간이 맡게 될 역할_로 중복 삽입되어 문장이 비문에 가까워지고 불필요한 강조가 추가되었습니다.

QL-006 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.5500
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-007 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: More papers than anyone can review
  • Target: 더 많은 논문들로 누구도 모두 검토할 수 없었다
  • Suggested fix: 누구도 모두 검토할 수 없을 만큼 많은 논문
  • Reason: 원문은 누구도 검토할 수 있는 양을 넘어서는 논문이 있다는 제목형 표현인데, 번역문은 ‘논문들로’라는 부자연스러운 구조와 과거형 ‘검토할 수 없었다’를 사용해 의미와 제목의 간결한 기능을 훼손합니다.

QL-008 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.2500
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-009 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: continuing an exponential trend that is at least partly driven by AI agents making it faster to run experiments and write them up.
  • Target: 이는 전년 대비 거의 두 배에 달하는 수치이며, AI 에이전트가 실험을 더 빨리 수행하고 글로 옮길 수 있게 해 주는 요인이 부분적으로 작용한 지수적 추세를 이어가고 있다.
  • Suggested fix: 이는 전년 대비 거의 두 배에 달하는 수치입니다. 이러한 증가는 AI 에이전트가 실험을 더 빠르게 수행하고 결과를 논문으로 작성할 수 있게 한 데 부분적으로 영향을 받은 것으로 보이며, 지수적인 증가 추세를 이어가고 있습니다.
  • Reason: 수식 관계가 길고 '요인이 부분적으로 작용한 지수적 추세'가 한국어에서 다소 어색해, 추세의 원인과 ICML의 증가가 한 문장 안에서 모호하게 연결됩니다. 의미는 대체로 보존되었습니다.

QL-010 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Here is a review of one accepted ICML 2026 spotlight paper, in the reviewer's own words:
  • Target: 아래는 ICML 2026의 한 주목 논문에 대한 심사의, 심사자의 직접 발언이다:
  • Suggested fix: 아래는 ICML 2026에서 spotlight 논문으로 선정된 한 논문에 대한 심사자의 평가를, 심사자의 표현 그대로 옮긴 것입니다:
  • Reason: ‘한 주목 논문’과 ‘심사의, 심사자의 직접 발언이다’가 영어 구조를 직역한 형태라 문장이 부자연스럽고 의미 관계가 불명확합니다. ‘spotlight paper’는 논문 유형을 나타내므로 자연스럽게 풀어야 합니다.

QL-011 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: An OpenReview reviewer admitting the proofs were not checked carefully
  • Target: An OpenReview reviewer admitting the proofs were not checked carefully
  • Suggested fix: 증명을 꼼꼼히 확인하지 않았다고 인정하는 OpenReview 리뷰어
  • Reason: 한국어로 번역되어야 할 문단이 원문과 동일한 영어로 남아 있어 독자에게 핵심 내용이 전달되지 않습니다.

QL-012 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Keep it in mind, because we will come back to this exact paper later in the post, and to what happened when we finally did check the proofs carefully.
  • Target: 이를 염두에 두시길. 나중에 이 게시물에서 바로 이 논문에 대해 다시 다루고, 우리가 증명을 꼼꼼히 확인했을 때 일어난 일에 대해서도 다룰 예정입니다.
  • Suggested fix: 이 점을 기억해 두시기 바랍니다. 글의 뒤에서 바로 이 논문을 다시 살펴보고, 나중에 증명을 꼼꼼히 확인한 결과 어떤 일이 있었는지도 다루겠습니다.
  • Reason: ‘이를 염두에 두시길’은 명령형 어미가 다소 부자연스럽고, ‘이 게시물에서’, ‘다룰 예정입니다’가 반복되어 영어 구조를 직역한 느낌이 납니다. 의미는 대체로 보존되지만 기술 블로그 문체로는 더 자연스럽게 연결할 수 있습니다.

QL-013 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: What has changed, though, is that the same technology driving the flood of submissions can also help us keep up with it.
  • Target: 다만 바뀐 점은 제출의 흐름을 일으키는 동일한 기술이 이를 따라잡는 데에도 도움을 준다는 것이다.
  • Suggested fix: 다만 달라진 점은 제출이 급증하게 만든 바로 그 기술이 이제 그 증가세를 따라잡는 데에도 도움이 될 수 있다는 것입니다.
  • Reason: ‘제출의 흐름을 일으키는 동일한 기술’은 영어 구조를 직역한 표현이라 의미가 어색하고, ‘바뀐 점은 ... 것이다’도 기술 블로그 문체에 부자연스럽습니다. 원문의 대조와 기술이 제출 증가에 대응하는 데 도움이 된다는 의미는 유지하면서 자연스럽게 풀어야 합니다.

QL-014 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Checking a paper carefully used to cost a reviewer a weekend; an agent can attempt it in an afternoon, in parallel, thousands of times over.
  • Target: 논문을 꼼꼼히 검토하는 데는 리뷰어의 주말이 필요하던 시절이 있었다; 이제 에이전트는 오후 한두 시간에 시도하고, 병렬로 수천 번 실행할 수 있다.
  • Suggested fix: 논문을 꼼꼼히 검토하려면 예전에는 리뷰어가 주말을 온전히 써야 했습니다. 이제는 에이전트가 오후 한 번의 작업으로 검토를 시도하고, 이를 병렬로 수천 번 반복할 수 있습니다.
  • Reason: 세미콜론을 그대로 유지했고, ‘리뷰어의 주말이 필요하다’와 ‘오후 한두 시간에 시도하다’가 한국어로 어색합니다. 또한 source의 ‘an afternoon’은 특정한 오후가 아니라 오후 시간대라는 뜻이므로 ‘오후 한두 시간’으로 좁히면 안 됩니다.

QL-015 accuracy / major

  • Message: MQM judge adequacy score is low.
  • Target: 0.7200
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM adequacy score is below threshold 0.75.

QL-016 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: with all the diversity of agent frameworks, compute budgets, and scientific taste that brings
  • Target: 에이전트 프레임워크의 다양성, 컴퓨트 예산, 과학적 취향이 어우러진 다양한 구성원을 포용했습니다
  • Suggested fix: 에이전트 프레임워크, 컴퓨트 예산, 과학적 취향이 만들어 내는 모든 다양성과 함께 커뮤니티 전체에 공개했습니다
  • Reason: 원문은 커뮤니티에 개방했을 때 에이전트 프레임워크, 컴퓨트 예산, 과학적 취향의 다양성이 함께 따라온다는 의미입니다. 번역문은 이를 ‘다양한 구성원을 포용했다’는 내용으로 바꾸어 주체와 의미를 추가·변경했습니다.

QL-017 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: ICML 2026 Open Reproductions challenge
  • Target: ICML 2026 Open Reproductions challenge은
  • Suggested fix: ICML 2026 Open Reproductions challenge는
  • Reason: 행사명 뒤에 한국어 조사를 붙일 때 영어 명칭의 문법적 연결이 어색하고, 고유 행사명 표기도 일관되지 않습니다. 행사명은 그대로 보존하고 자연스럽게 문장에 연결하는 것이 좋습니다.

QL-018 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.7000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-019 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: **Pick a paper.** We indexed all 6,341 accepted ICML 2026 papers with their abstracts and extracted the core scientific claims of each one, so an agent could start from a concrete, checkable target rather than a 40-page PDF. Multiple people reproducing the same paper was encouraged.
  • Target: **논문 하나를 고르기.** ICML 2026에서 수락된 6,341편의 논문을 초록과 함께 색인화하고 각 논문의 핵심 과학적 주장을 추출해서 에이전트가 40페이지 분량의 PDF를 처음부터 끝까지 파악하기보다는 구체적이고 검증 가능한 목표에서 시작할 수 있도록 했습니다. 같은 논문의 재현이 여러 사람에 의해 이뤄지는 것을 장려했습니다.
  • Suggested fix: 논문 고르기. ICML 2026에서 승인된 논문 6,341편을 초록과 함께 색인화하고 각 논문의 핵심 과학적 주장을 추출했습니다. 이를 통해 에이전트가 40페이지 분량의 PDF 전체를 처음부터 검토하는 대신, 구체적이고 검증 가능한 목표에서 시작할 수 있도록 했습니다. 같은 논문을 여러 사람이 재현하는 것도 권장했습니다.
  • Reason: 목록 항목 제목이 명사형으로 어색하고, 마지막 문장의 ‘여러 사람에 의해 이뤄지는 것을 장려했습니다’는 번역투가 강해 읽기 어렵습니다. 의미는 대체로 보존되어 있으므로 정확성 문제가 아니라 문체 문제입니다.

QL-020 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: **Bring your own agent.**
  • Target: **당신의 에이전트를 가져오세요.**
  • Suggested fix: 자신의 에이전트를 사용하세요.
  • Reason: 의미는 전달되지만 영어식 소유 대명사와 직역 표현으로 한국어 기술 블로그의 제목형 문장으로는 어색합니다.

QL-021 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: and everything in between
  • Target: 그 사이의 모든 것을 이용했습니다
  • Suggested fix: 그 밖에도 다양한 도구를 사용했습니다
  • Reason: 원문의 ‘everything in between’은 두 예시 사이에 해당하는 다양한 도구를 포괄한다는 뜻인데, 현재 번역은 ‘그 사이’를 물리적 위치처럼 옮겨 의미가 불명확합니다.

QL-022 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: so an agent could pull the paper, its claims, and the challenge instructions with a single command
  • Target: 에이전트가 논문, 주장, 도전 지시를 단일 명령으로 끌어올 수 있도록
  • Suggested fix: 에이전트가 한 번의 명령으로 논문과 논문의 주장, 챌린지 지침을 가져올 수 있도록
  • Reason: ‘pull’은 여기서 필요한 자료를 가져오거나 불러오는 동작을 뜻하며, ‘끌어올 수 있도록’은 기술 문맥에서 부자연스럽고 동작을 오해하게 할 수 있습니다. 또한 ‘challenge instructions’를 ‘도전 지시’로 옮기면 의미가 불명확합니다.

QL-023 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: the write-up
  • Target: 쓰기 구성
  • Suggested fix: 설명 문서
  • Reason: ‘write-up’은 실행 결과를 설명하는 문서나 보고서를 뜻하는데, ‘쓰기 구성’은 의미가 불명확해 로그북에 포함되는 설명 문서를 제대로 전달하지 못합니다.

QL-024 terminology / major

  • Message: MQM judge reported terminology issue.
  • Source: Hugging Face Dataset
  • Target: Hugging Face 데이터셋
  • Suggested fix: Hugging Face Dataset
  • Reason: 제품명인 Hugging Face Dataset을 일반 명사 ‘데이터셋’으로 번역해 제품 식별성과 검색성이 떨어집니다.

QL-025 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.7200
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-026 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: **toy** (evidence at reduced scale)
  • Target: **소품(toy)**(감축된 데이터에서의 증거)
  • Suggested fix: toy(축소된 규모에서의 증거)
  • Reason: 여기서 toy는 장난감이나 소품이 아니라 축소된 규모에서 얻은 증거를 나타내는 판정 레이블입니다. ‘소품’으로 번역하면 판정 기준의 의미를 오해하게 됩니다.

QL-027 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: An automated Logbook Judge (running an open-weights model, GLM-5.2)
  • Target: 자동화된 로그북 심사관(Logbook Judge, 오픈 가중치 모델 GLM-5.2를 실행)
  • Suggested fix: 자동화된 로그북 심사관(Logbook Judge, 오픈 웨이트 모델 GLM-5.2 사용)
  • Reason: 괄호 안의 ‘모델을 실행’은 영어식으로 어색하고, open-weights는 모델의 특성을 설명하는 표현이므로 ‘오픈 웨이트 모델을 사용하는’처럼 옮기는 편이 자연스럽고 의미가 분명합니다.

QL-028 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: evidence at reduced scale
  • Target: 감축된 데이터에서의 증거
  • Suggested fix: 축소된 규모에서의 증거
  • Reason: reduced scale은 데이터가 감축되었다는 뜻이 아니라 실험이나 평가를 축소된 규모로 수행했다는 의미입니다.

QL-029 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: $20 in Hugging Face compute credits
  • Target: Hugging Face 컴퓨트 크레딧 20
  • Suggested fix: Hugging Face 컴퓨트 크레딧 $20
  • Reason: 금액의 달러 단위인 $20에서 $가 누락되어 크레딧 금액이 불완전하게 전달됩니다.

QL-030 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: Anatomy of a reproduction logbook: logbook pages, agent traces, and artifacts
  • Target: Anatomy of a reproduction logbook: logbook pages, agent traces, and artifacts
  • Suggested fix: 재현 로그북의 구성: 로그북 페이지, 에이전트 추적 기록, 산출물
  • Reason: 제목 전체가 번역되지 않아 한국어 독자에게 제목의 핵심 내용이 전달되지 않습니다. 기술 용어와 구조는 유지하되 제목을 한국어로 옮겨야 합니다.

QL-031 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: organization
  • Target: organization
  • Suggested fix: 조직
  • Reason: 일반 명사인 organization이 번역되지 않아 한국어 목록 항목 안에서 문장이 어색하고 의미 전달이 덜 자연스럽습니다.

QL-032 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.6200
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-033 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: **6,816** reproduction logbooks published
  • Target: **6,816** 재현 로그북이 게시되었다
  • Suggested fix: 6,816개의 재현 로그북 게시
  • Reason: 목록 항목의 서술이 과거형 평서문으로 끝나 어색하며, 공식 기술 블로그의 존댓말 문체와도 맞지 않습니다. 의미와 수치는 유지되지만 목록 항목으로는 명사구 또는 자연스러운 합니다체가 적절합니다.

QL-034 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: papers attempted
  • Target: 논문이 시도되었고
  • Suggested fix: 논문 2,226편에서 시도가 이루어졌고
  • Reason: ‘papers attempted’는 논문을 시도했다는 뜻이 아니라 해당 논문들에서 시도가 이루어졌거나 논문이 시도 대상으로 포함되었다는 의미입니다. 현재 번역은 논문 자체가 시도된 것처럼 표현해 의미가 부자연스럽고 오해를 일으킵니다.

QL-035 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: many by several independent teams
  • Target: 다수는 여러 독립 팀에 의해 수행되었다
  • Suggested fix: 그중 다수는 여러 독립 팀이 각각 수행했습니다
  • Reason: ‘다수는’의 지시 대상이 불분명하고 ‘여러 독립 팀에 의해 수행되었다’는 표현은 번역투입니다. 여러 독립 팀이 각각 수행했다는 관계가 드러나도록 주어와 수식을 자연스럽게 정리해야 합니다.

QL-036 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: **35,908** claims judged, with all verdicts frozen in a public dataset at challenge close
  • Target: **35,908** 주장들이 평가되었고, 모든 판정은 도전 종료 시 공개 데이터세트에 고정되었다
  • Suggested fix: 35,908개의 주장을 평가했으며, 모든 판정은 챌린지 종료 시 공개 데이터세트에 기록되었습니다
  • Reason: 의미는 대체로 전달되지만 ‘주장들이’는 한국어에서 부자연스럽고, ‘판정을 데이터세트에 고정했다’도 직역투입니다. 또한 리스트 항목의 종결 표현이 문어체로 다소 어색합니다.

QL-037 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: HF Jobs launched
  • Target: HF 작업이 시작되었고
  • Suggested fix: 2,962개의 HF Jobs가 시작됨
  • Reason: HF Jobs는 Hugging Face의 제품·기능명으로 보존해야 하며, ‘HF 작업’으로 번역하면 제품명이 일반 작업으로 오인될 수 있습니다. 또한 목록 항목의 간결한 명사구 구조와 달리 문장형으로 끝납니다.

QL-038 terminology / minor

  • Message: MQM judge reported terminology issue.
  • Source: full agent-trace datasets
  • Target: 전체 에이전트 추적 데이터세트
  • Suggested fix: 전체 agent-trace 데이터셋
  • Reason: ‘full agent-trace’의 핵심 검색어인 agent-trace가 ‘에이전트 추적’으로만 번역되어 원문의 기술 용어 검색성이 떨어집니다. ‘데이터세트’ 자체는 이해 가능하지만, 기술 블로그에서는 dataset을 데이터셋으로 통일하는 편이 자연스럽습니다.

QL-039 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: had at least one claim independently verified
  • Target: 최소 하나의 주장을 독립적으로 검증했습니다
  • Suggested fix: 최소 하나의 주장이 독립적으로 검증되었습니다
  • Reason: 원문은 논문에 포함된 주장 중 하나 이상이 독립적으로 검증되었다는 수동적 사실을 말하지만, 번역은 논문이 주장을 직접 검증했다는 능동적 의미로 바뀌었습니다.

QL-040 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: with every extracted claim verified
  • Target: 모든 추출된 주장을 검증한 채로
  • Suggested fix: 추출된 모든 주장이 검증된 상태로
  • Reason: ‘검증한 채로’는 논문 재현 결과를 설명하는 문맥에서 어색하고 영어식 구조가 남아 있습니다.

QL-041 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: with nothing falsified
  • Target: 아무 것도 부정되지는 않았습니다
  • Suggested fix: 반증된 주장은 없었습니다
  • Reason: 원문의 ‘falsified’는 주장이 거짓으로 판명되거나 반증되었다는 뜻입니다. ‘부정되다’는 이 기술적 의미를 충분히 명확하게 전달하지 못합니다.

QL-042 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: The remainder sat in the middle:
  • Target: 나머지는 중간에 위치했다:
  • Suggested fix: 나머지는 중간 범주에 해당했습니다.
  • Reason: ‘중간에 위치했다’는 이 맥락에서 다소 직역투이며, 콜론을 그대로 유지해 문장이 부자연스럽습니다. 나머지 논문이 중간 범주에 해당한다는 의미로 자연스럽게 풀어 쓰는 것이 좋습니다.

QL-043 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: 280 where nothing could be established either way
  • Target: 280편은 어느 쪽으로도 확립될 수 없었다
  • Suggested fix: 280편은 어느 쪽인지 판단할 수 없었습니다
  • Reason: ‘nothing could be established either way’는 양쪽 어느 쪽인지 판단할 수 없었다는 뜻인데, 현재 번역은 무엇도 확립될 수 없었다는 어색하고 불명확한 표현으로 의미를 충분히 전달하지 못합니다.

QL-044 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Reproductions done well
  • Target: 잘 수행된 재현
  • Suggested fix: 재현을 잘하는 방법
  • Reason: 의미는 전달되지만 ‘수행된 재현’은 영어 구조를 직역한 표현처럼 느껴져 한국어 제목으로 다소 어색합니다.

QL-045 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Some papers came through the gauntlet looking great
  • Target: 일부 논문은 도전 과정을 거쳐 훌륭하게 보였고
  • Suggested fix: 일부 논문은 까다로운 검증 과정을 거친 뒤에도 훌륭한 결과를 보여 주었고
  • Reason: ‘came through the gauntlet looking great’는 어려운 검증이나 선별 과정을 통과한 뒤에도 좋은 평가를 받았다는 뜻에 가깝습니다. 현재 번역은 ‘훌륭하게 보였다’가 어색하고 원문의 관용적 의미가 충분히 전달되지 않습니다.

QL-046 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: The one linked included and published the full agent trace.
  • Target: 연결된 버전은 전체 에이전트 추적을 포함해 공개했습니다.
  • Suggested fix: 링크된 팀은 전체 에이전트 추적을 포함해 공개했습니다.
  • Reason: 원문의 주어인 'The one linked'는 링크된 팀 또는 재현 결과를 가리키지만, 현재 번역의 '연결된 버전'은 링크된 버전이라는 별도 대상을 만들어 의미가 불분명해졌습니다. 또한 원문은 전체 에이전트 추적을 포함하고 공개했다는 사실을 말하므로, 링크된 주체를 명확히 드러내야 합니다.

QL-047 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: A paper about unreliable LLM judges holding up under scrutiny by LLM agents :)
  • Target: LLM 심판의 신뢰성에 대한 논문이 LLM 에이전트의 검증 아래 견뎌냈다는 내용 :)
  • Suggested fix: 신뢰할 수 없는 LLM 심판을 다룬 논문이 LLM 에이전트의 검토에도 견뎌냈다는 내용 :)
  • Reason: 원문은 ‘신뢰할 수 없는 LLM 심판’을 다룬 논문이 LLM 에이전트의 검토를 받아도 견고함을 보였다는 유머러스한 요약입니다. 번역문은 ‘LLM 심판의 신뢰성에 대한 논문’으로 바꾸어 논문의 주제를 왜곡하고, ‘논문이 ... 견뎌냈다’는 구조도 불명확하게 만들었습니다.

QL-048 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.5800
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-049 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: formally claimed they had falsified something
  • Target: 형식적으로 무언가를 반증했다고 주장했다
  • Suggested fix: 무언가를 조작했다고 공식적으로 주장했다
  • Reason: 여기서 falsified는 어떤 주장이나 결과를 거짓으로 입증했다는 뜻이 아니라, 위조·조작을 했다고 인정했다는 의미입니다. '반증했다'로 옮기면 과학적 반증 행위로 의미가 바뀌어 참가자들의 진술 내용을 오해하게 합니다.

QL-050 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: We adversarially re-verified every claimed falsification:
  • Target: 우리는 그 주장된 반증을 적대적으로 재확인했다:
  • Suggested fix: 우리는 조작했다는 각각의 주장을 반대 입장에서 엄격하게 다시 검증했습니다.
  • Reason: '적대적으로 재확인했다'와 '주장된 반증'은 영어 구조를 직역한 표현으로 한국어 기술 문맥에서 부자연스럽고 의미도 불명확합니다. adversarially re-verified는 각 진술을 반대 입장에서 엄격하게 다시 검증했다는 뜻입니다.

QL-051 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: re-reading the paper, re-reading the logbook, and re-deriving the math or re-implementing the experiment from the paper's own text.
  • Target: 논문을 다시 읽고, 로그북을 다시 읽고, 논문 자체의 텍스트에서 수학을 재도출하거나 실험을 재구현했다.
  • Suggested fix: 논문과 실험 기록을 다시 검토하고, 논문에 적힌 내용만으로 수식을 다시 유도하거나 실험을 재현했습니다.
  • Reason: '논문 자체의 텍스트에서 수학을 재도출하거나'는 직역투이고, 문장 끝이 앞 문장의 과거 시제와 어울리지 않습니다. 논문 내용을 바탕으로 수학을 다시 유도하거나 실험을 재현했다는 의미를 자연스럽게 연결해야 합니다.

QL-052 technical / major

  • Message: MQM judge technical score is low.
  • Target: 0.7200
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM technical score is below threshold 0.75.

QL-053 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: A few of the confirmed falsifications, linking to the logbook that found it:
  • Target: 확인된 반증 중 일부를 아래에 제시하며, 이를 발견한 로그북으로 연결합니다:
  • Suggested fix: 확인된 반증 중 일부를, 이를 확인한 로그북 링크와 함께 소개합니다:
  • Reason: 원문의 'linking to the logbook that found it'은 반증을 확인한 근거 기록으로 연결한다는 의미인데, '이를 발견한 로그북'은 로그북 자체가 반증을 발견한 것처럼 읽혀 주체와 의미 관계가 어색하게 바뀝니다.

QL-054 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.6800
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-055 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: A theorem that falls after step 224.
  • Target: **224단계 이후에 떨어지는 정리.**
  • Suggested fix: 224단계 이후에 무너지는 정리.
  • Reason: “falls”를 “떨어지는”으로 직역해 한국어 제목이 부자연스럽고, 정리가 반박되거나 무너진다는 의미가 충분히 전달되지 않습니다.

QL-056 terminology / major

  • Message: MQM judge reported terminology issue.
  • Source: convex hull
  • Target: 볼록 궤
  • Suggested fix: 볼록 껍질
  • Reason: convex hull은 기술 용어인 “볼록 껍질”을 의미하는데 “볼록 궤”로 번역되어 수학적 개념이 잘못 전달됩니다.

QL-057 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: whenever the origin starts inside their convex hull
  • Target: 원점이 그들의 볼록 궤 안에 들어오면
  • Suggested fix: 원점이 토큰 입자들의 볼록 껍질 안에서 시작할 때마다
  • Reason: “starts inside”의 초기 조건과 “whenever”의 일반적 조건이 어색하고 부정확하게 옮겨졌습니다. 또한 “their”는 토큰 입자들의 볼록 껍질을 가리킵니다.

QL-058 terminology / minor

  • Message: MQM judge reported termino
SEO report — ✅ Pass

SEO Eval Report

Gate: ✅ PASS — deterministic AND rubric

  • File: ../target/_posts/2026-08-13-icml-2026-open-reproductions.md
  • Source: —
  • Primary keyword: (none — D5 skipped)
  • Mode: file

Gate

  • Status: PASS
  • Blockers: ✅ pass
  • Deterministic REQUIRED (D1–D7): ✅ pass
  • Rubric (R1–R6): ✅ pass (mean None, min None)

Blockers

✅ body_not_empty: Body is not empty
✅ robots_indexable: Robots is indexable
✅ internal_links_resolve: All internal links resolve
✅ local_images_resolve: All local images resolve

Required checks (gated)

✅ heading_hierarchy: Heading hierarchy: Valid
✅ alt_text_coverage: Alt text coverage: 3/3 images
✅ descriptive_alt_text: Descriptive alt text: 3/3 (≥80% recommend)
✅ image_files_exist: All 0 local image file(s) exist

OpenAI rubric checks

✅ semantic_metadata: PASS (required) — Semantic meaning is consistent across title, H1, headings, and opening text; no mismatch detected.
✅ alt_semantics: PASS (review) — —

Advisory checks (not gated)

✅ opening_summary: Opening 3 paragraphs: 59 words (recommend ≥50 for GEO)
✅ h1_count: Markdown H1 count: 1 (review against rendered layout)
✅ citations: Citations/statistics: 21 (recommend ≥1 for GEO)
⚠️ question_headings: Scannable H2/H3 (question or keyword): 0 (0 question, 0 keyword)
⚠️ internal_links: Internal links: 0 (recommend 2-3)
✅ word_count: Body length: 8700 chars (recommend ≥800 for KO)
ℹ️ primary_keyword: No primary_keyword in manifest — keyword check skipped
⚠️ webp_format: WebP format: 0/3 images (≥50% recommend)
⚠️ lazy_loading: Lazy loading: 0 images (optional)

Signals (evidence — not directly gated)

  • Frontmatter: title 27 chars, description 0 chars, author present True
  • Title text: ICML에서 2,200편의 논문 재현으로 배운 것
  • Description text: —
  • Opening text: * TOC
    {:toc}

이 글은 Hugging Face 블로그의 What We Learned by Reproducing 2,200 papers from ICML를 한국어로 번역한 글입니다.

  • Opening: first paragraph 178 chars, first 3 paragraphs 491 chars
  • Headings: markdown H1 1, rendered effective H1 2, layout title H1 True
  • Links: total 14, external 14, internal 0, citation signals 21
  • Images: total 3, empty alt 0, filename-like alt 0, missing local files 0

Semantic review packet

  • Title: ICML에서 2,200편의 논문 재현으로 배운 것
  • Description: —
  • Rendered H1 candidates: ICML에서 2,200편의 논문 재현으로 배운 것, ICML에서 2,200편의 논문 재현으로 배운 것
  • Opening: * TOC
    {:toc}

이 글은 Hugging Face 블로그의 What We Learned by Reproducing 2,200 papers from ICML를 한국어로 번역한 글입니다.

  • Canonical/permalink: —
  • Instruction: Compare title, description, rendered H1, and opening text for meaning consistency. This packet is evidence only; it does not decide pass/fail.

Frontmatter (advisory — written by metadata step, not gated)

✅ title: Title: 27 chars (recommend ≤60)
❌ description: Description is missing
✅ image: OG image: assets/images/blog/posts/2026-08-13-icml-2026-open-reproductions/thumbnail.png
✅ categories: Categories: 2 (recommend 2-3)
✅ author: Author: dailybot

SEO metadata suggestion — PARTIAL

This is a suggestion. SEO is applied only when the post frontmatter is updated.
To apply safe fields from a partial suggestion, leave a trusted PR comment: metadata apply.

  • Auto apply: False
  • Requires human: True
  • Mode: frontmatter_only
  • Reason: metadata candidate needs policy decisions or missing title/description

Candidate

  • categories: ['Translation', 'HuggingFace']
  • image: assets/images/blog/posts/2026-08-13-icml-2026-open-reproductions/thumbnail.png

Needs policy decision

  • target_url
  • source_url
  • canonical_policy
  • translation_indexing
  • target_locale
  • source_locale

Warnings

  • None

@Jwaminju Jwaminju added the hf-agent:needs-human HF Agent needs human follow-up label Aug 14, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hf-agent:managed Opt PR into HF Agent review automation hf-agent:needs-human HF Agent needs human follow-up

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant