Skip to content

Translate Hugging Face blog post: State of Open Models: Summer 2026 Observations - #195

Open
Jwaminju wants to merge 1 commit into
mainfrom
translate/state-of-open-models-summer-2026
Open

Translate Hugging Face blog post: State of Open Models: Summer 2026 Observations#195
Jwaminju wants to merge 1 commit into
mainfrom
translate/state-of-open-models-summer-2026

Conversation

@Jwaminju

Copy link
Copy Markdown
Collaborator

Source: https://huggingface.co/blog/state-of-open-models-summer-2026

This PR adds a Korean translation draft for state-of-open-models-summer-2026.

Downstream handoff:

  • SEO review should use the translation-flow manifest.
  • Quality review should use the translation-flow manifest.

@Jwaminju Jwaminju added the hf-agent:managed Opt PR into HF Agent review automation label Aug 15, 2026
@github-actions

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1

🚀 View preview at
https://hugging-face-krew.github.io/pr-preview/pr-195/

Built to branch gh-pages at 2026-08-15 01:49 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@Jwaminju

Jwaminju commented Aug 15, 2026

Copy link
Copy Markdown
Collaborator Author

HF Agent Review

Gate Result
Quality ❌ Fail
SEO ❌ Fail

Head SHA: e20140bfea99c66e294656863d20e9f99416263e

Quality report — ❌ Fail

Quality Report

  • Status: reject
  • Quality Score: 0.0
  • Hard failures: 25
  • Issues: 168
  • Source available: True
  • Source changed: False
  • Source segments: 76
  • Target segments: 76

Scorecard

Dimension Score
adequacy 0.0
technical_accuracy 0.0
completeness 0.0
terminology 0.0
fluency 0.0
publishing_integrity 0.0
style_locale 60.0

Metrics

  • qe_metric: heuristic
  • qe_average: 0.9609
  • qe_min: 0.4875
  • embedding_similarity_average: 0.9901
  • embedding_similarity_min: 0.8033
  • cache_hits: 0
  • cache_misses: 152

MQM Judge

  • Enabled: True
  • Provider: openai
  • Model: gpt-5.6-luna
  • Reasoning effort: none
  • Prompt: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/judges/mqm_prompt.md
  • Prompt hash: 887d2931aa289213f0bdce4a917a0ac8364dad8e470011069b8f8bf758d69a90
  • Style guide hash: 937d8cd893578d30e716a3eb513cdf5f10d6fd3ad8f5e77068b57f96e160de12
  • Requested segments: 76
  • Evaluated segments: 75
  • MQM errors: 53
  • Cache hits: 0
  • Cache misses: 76
  • Severity counts: {'critical': 25, 'major': 18, 'minor': 10}
  • adequacy_average: 0.5409
  • technical_average: 0.7624
  • fluency_average: 0.6556
  • warning: Skipped MQM error for segment p_015: source_span is not verbatim source text.
  • warning: Skipped MQM result for segment p_015: at least one error was invalid.
  • warning: MQM segment coverage is invalid: expected exactly one result for every aligned target segment.

Style Guide

  • Enabled: True
  • Guide: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/style/hf-blog-ko-translation-guide.md
  • Policy: /home/runner/work/hugging-face-krew.github.io/hugging-face-krew.github.io/workflow/skills/quality/configs/style_policy.yml
  • Style score: 60.0
  • Rule hits: {'alt_text_caption': 13, 'first_mention_bilingual': 2, 'link_text_translation': 22, 'modal_strength': 18}

Style Guide Findings

Rule Severity Segment Current Suggested
modal_strength major p_008 In almost every month of 2026, the largest and most performant open model from a Chinese lab was larger than anything an American lab released of its own. China's monthly ceiling ran between 754B and 2.78 trillion parameters; America's own ceiling stayed under 130B in five of seven months, the exception being NVIDIA's Nemotron 3 Ultra at 561B in May and June, and Inkling from Thinking Machines Lab. Preserve the strength of may using: 수 있습니다, 일 수 있습니다.
modal_strength major p_008 In almost every month of 2026, the largest and most performant open model from a Chinese lab was larger than anything an American lab released of its own. China's monthly ceiling ran between 754B and 2.78 trillion parameters; America's own ceiling stayed under 130B in five of seven months, the exception being NVIDIA's Nemotron 3 Ultra at 561B in May and June, and Inkling from Thinking Machines Lab. Preserve the strength of can using: 수 있습니다.
modal_strength major p_018 At the frontier scale, the picture is very different. Most U.S. releases above 100B parameters this year are not new models, but built on top of Chinese models. Only a few major original American models appear at this scale: Thinking Machines’ Inkling (952B), NVIDIA’s Nemotron 3 Ultra (561B), Nemotron 3 Super (124B), and Arcee AI’s Trinity-Large (399B). Preserve the strength of can using: 수 있습니다.
modal_strength major p_019 AMD contributed many conversions but no original model at this scale. This work is still important: it enables trillion-parameter Chinese models to run efficiently on American hardware. But it represents a distribution and optimization layer rather than model creation. Preserve the strength of can using: 수 있습니다.
modal_strength major p_027 Chinese frontier labs are the only accounts on the Hub where the heavy band carries the volume. Effectively all of MiniMax's 2026 downloads are of models above 70B, along with 88% of Moonshot's, 55% of DeepSeek's and 39% of Z.ai's. No large American account looks like this: Google, Microsoft and IBM Granite record essentially none of their 2026 downloads above 70B, and NVIDIA and Meta only 14% and 9%. Preserve the strength of can using: 수 있습니다.
modal_strength major p_035 DeepSeek and Z.ai ship models between 700 billion and 1.65 trillion parameters under plain MIT. Chinese labs license their largest models about as permissively as their smallest, and more permissively than American labs license theirs: on the American side of the same size band, 29% is Apache or MIT, 41% sits under custom terms and 30% declares nothing at all. Preserve the strength of can using: 수 있습니다.
modal_strength major p_057 Repositories declaring the gguf library rose 464%, lerobot 194% and Apple's mlx148%, against 16% for transformers and peft and 21% for diffusers. The modelling core is growing at roughly the platform average. The layer that decides where a model can physically run local inference formats, Apple silicon, robot control stacks, is growing three to seven times faster than that. Preserve the strength of can using: 수 있습니다.
modal_strength major p_060 We could not have written this section in March, because the instrument did not exist. The agent-usage dataset, published in July, records the agent/<name> token that coding agents send when they call the Hub through huggingface\_hub or the hf CLI — searching for models, pushing datasets, running Jobs, creating Spaces. For the first time we can see how much agent traffic the Hub receives and which harnesses it comes from. Preserve the strength of can using: 수 있습니다.
modal_strength major p_061 Agents calling the Hugging Face Hub Claude Code led July with 44.4%, but a single month conceals the real finding: it held 67.8% in April and 6.4% in May, while Codex climbed steadily from 10.4% to 20.8%. This is a market with no incumbent, where one release or one changed default can move half the traffic in a month. Preserve the strength of may using: 수 있습니다, 일 수 있습니다.
modal_strength major p_061 Agents calling the Hugging Face Hub Claude Code led July with 44.4%, but a single month conceals the real finding: it held 67.8% in April and 6.4% in May, while Codex climbed steadily from 10.4% to 20.8%. This is a market with no incumbent, where one release or one changed default can move half the traffic in a month. Preserve the strength of can using: 수 있습니다.

Issues

QL-001 technical / major

  • Message: number/unit token mismatch.
  • Target: 1, 2, 8
  • Suggested fix: Preserve source number/unit token exactly.
  • Reason: Review gate exact-match validator failed: extra=['1', '2', '8']

QL-002 accuracy / major

  • Message: QE metric score is low.
  • Target: 0.4875
  • Suggested fix: Review this segment for omission, unrelated translation, or over-compression.
  • Reason: QE score is below threshold 0.55.

QL-003 accuracy / major

  • Message: QE metric score is low.
  • Target: 0.4875
  • Suggested fix: Review this segment for omission, unrelated translation, or over-compression.
  • Reason: QE score is below threshold 0.55.

QL-004 accuracy / major

  • Message: QE metric score is low.
  • Target: 0.5458
  • Suggested fix: Review this segment for omission, unrelated translation, or over-compression.
  • Reason: QE score is below threshold 0.55.

QL-005 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: State of Open Models: Summer 2026 Observations
  • Target: 오픈 모델 현황: 2026년 여름 관찰
  • Suggested fix: 오픈 모델 현황: 2026년 여름 동향
  • Reason: ‘Observations’를 ‘관찰’로 옮기면 한국어 제목에서 다소 직역투이고 부자연스럽습니다. 제목의 의미는 유지되지만 ‘살펴보기’ 또는 ‘동향’이 더 자연스럽습니다.

QL-006 fluency / major

  • Message: MQM judge reported fluency issue.
  • Source: In the AI world, time feels compressed. A few months after our spring report in our biannual analysis worked through the ecosystem, there are quite a few findings that we have observed until this summer.
  • Target: AI 세계에서 시간은 압축된 느낌이다. 우리의 spring report이 생태계를 살펴본 연 2회 분석 이후 불과 몇 달이 지났고, 이번 여름까지 관찰한 여러 발견들이 있다.
  • Suggested fix: AI 분야에서는 시간이 빠르게 흘러가는 듯합니다. 생태계 전반을 살펴본 연 2회 분석의 봄 보고서를 발표한 지 불과 몇 달이 지났지만, 이번 여름까지 관찰한 결과에서도 여러 가지 새로운 사실을 확인했습니다.
  • Reason: 문장이 영어 구조를 기계적으로 따라가 한국어로 부자연스럽고 의미 관계가 불분명합니다. 특히 'spring report이', '연 2회 분석 이후', '관찰한 여러 발견들이 있다'가 자연스럽지 않으며, 공식 블로그 문체에 필요한 존댓말도 지켜지지 않았습니다.

QL-007 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: spring report
  • Target: spring report
  • Suggested fix: 봄 보고서
  • Reason: 일반 명사인 'spring report'를 문장 안에서 영어로만 남겨 한국어 문장의 자연스러움과 의미 전달을 떨어뜨립니다. 제품명이나 고유명사가 아니므로 계절 보고서로 옮기는 것이 적절합니다.

QL-008 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: This report lays out these observations from January to August 2026 and presents the data behind each one.
  • Target: 이 보고서는 2026년 1월부터 8월까지의 관찰을 제시하고 각 관찰 뒤의 데이터를 제시한다.
  • Suggested fix: 이 보고서는 2026년 1월부터 8월까지의 관찰을 정리하고, 각각을 뒷받침하는 데이터를 제시합니다.
  • Reason: '제시하고 ... 제시한다'가 반복되고 '각 관찰 뒤의 데이터'는 한국어에서 어색한 직역입니다. 원문의 의미인 각 관찰을 뒷받침하는 데이터를 자연스럽게 표현해야 합니다.

QL-009 accuracy / major

  • Message: MQM judge adequacy score is low.
  • Target: 0.6200
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM adequacy score is below threshold 0.75.

QL-010 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: Public model repositories grew from 2.43 to 2.96 million over the period, datasets from 711,000 to 1 million, Spaces from 1.00 to 1.44 million.
  • Target: 공개 모델 저장소는 기간 동안 2.43 to 2.96 million으로 증가했고, 데이터셋은 711,000에서 1 million으로, Spaces는 1.00에서 1.44 million으로 증가했다.
  • Suggested fix: 공개 모델 저장소는 해당 기간에 243만 개에서 296만 개로 증가했고, 데이터셋은 71만 1천 개에서 100만 개로, Spaces는 100만 개에서 144만 개로 증가했다.
  • Reason: 첫 번째 증가 범위가 한국어로 번역되지 않고 'to'가 그대로 남아 있으며, million도 일관되게 처리되지 않아 수치 정보의 가독성과 정확성이 손상되었습니다.

QL-011 technical / major

  • Message: MQM judge reported technical issue.
  • Source: fewer than 200 lifetime downloads
  • Target: lifetimesDownloads가 200 미만이고
  • Suggested fix: 전체 누적 다운로드 수가 200회 미만이며
  • Reason: 원문의 'lifetime downloads'가 존재하지 않는 형태의 'lifetimesDownloads'로 잘못 변환되어, 모델의 누적 다운로드 수라는 의미가 불명확해졌습니다.

QL-012 fluency / minor

  • Message: MQM judge reported fluency issue.
  • Source: Everything below happens inside that shape.
  • Target: 아래의 모든 내용은 그 형태 안에서 일어난다.
  • Suggested fix: 이하의 모든 내용은 이러한 분포 안에서 일어난다.
  • Reason: '그 형태 안에서 일어난다'는 원문의 분포 구조를 가리키는 표현을 기계적으로 옮겨 기술 블로그 문맥에서 어색합니다.

QL-013 style_locale / minor

  • Message: MQM judge reported style_locale issue.
  • Source: Models and datasets on HF hub are growing on a daily basis.
  • Target: Hugging Face Hub의 모델과 데이터셋은 매일 증가하고 있다.
  • Suggested fix: Hugging Face Hub의 모델과 데이터셋은 매일 증가하고 있습니다.
  • Reason: 전체 문장이 해라체로 번역되어 기술 블로그의 기본 존댓말 문체와 일치하지 않습니다.

QL-014 formatting / minor

  • Message: MQM judge reported formatting issue.
  • Source: **1\. The frontier is moving fast**
  • Target: **1\\. The frontier is moving fast**
  • Suggested fix: 1. The frontier is moving fast
  • Reason: 대상 텍스트에서 마침표 앞의 백슬래시가 원문보다 하나 더 추가되어 Markdown 이스케이프 구조가 변경되었습니다. 렌더링 결과나 후속 처리에 영향을 줄 수 있으므로 원문의 이스케이프를 그대로 유지해야 합니다.

QL-015 accuracy / minor

  • Message: MQM judge reported accuracy issue.
  • Source: smaller models
  • Target: 작고 가벼운 모델
  • Suggested fix: 작은 모델
  • Reason: 원문은 모델의 규모가 더 작다는 의미만 전달하지만, ‘가벼운’은 모델의 경량성이나 자원 요구량이 낮다는 추가 의미를 포함할 수 있습니다.

QL-016 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: Largest open-model releases from Chinese and US labs by month in 2026
  • Target: Largest open-model releases from Chinese and US labs by month in 2026
  • Suggested fix: 2026년 중국 및 미국 연구소의 월별 주요 오픈 모델 공개
  • Reason: 영문 캡션이 번역되지 않아 한국어 독자가 핵심 정보를 이해할 수 없습니다.

QL-017 technical / major

  • Message: MQM judge technical score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM technical score is below threshold 0.75.

QL-018 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-019 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: **In almost every month of 2026, the largest and most performant open model from a Chinese lab was larger than anything an American lab released of its own.** China's monthly ceiling ran between 754B and 2.78 trillion parameters; America's own ceiling stayed under 130B in five of seven months, the exception being NVIDIA's Nemotron 3 Ultra at 561B in May and June, and Inkling from Thinking Machines Lab.
  • Target: **In almost every month of 2026, the largest and most performant open model from a Chinese lab was larger than anything an American lab released of its own.** China's monthly ceiling ran between 754B and 2.78 trillion parameters; America's own ceiling stayed under 130B in five of seven months, the exception being NVIDIA's Nemotron 3 Ultra at 561B in May and June, and Inkling from Thinking Machines Lab.
  • Suggested fix: 문단 전체를 한국어로 번역하되 2026년, 754B, 2.78 trillion parameters, 130B, 561B, 모델명과 연구소명은 정확히 유지합니다.
  • Reason: 전체 문단이 한국어로 번역되지 않고 원문 영어가 그대로 반복되어, 한국어 독자에게 핵심 비교와 수치 정보가 전달되지 않습니다.

QL-020 technical / major

  • Message: MQM judge technical score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM technical score is below threshold 0.75.

QL-021 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-022 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: Every lab has a different size strategy
  • Target: Every lab has a different size strategy
  • Suggested fix: 각 연구소마다 규모를 정하는 전략이 다릅니다
  • Reason: 번역 대상 문장이 원문과 동일하게 남아 있어 한국어 독자에게 의미가 전달되지 않습니다.

QL-023 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: The chart splits the labs into two camps. Moonshot, MiniMax, Xiaomi and Z.ai publish almost nothing below 70B, so a developer's first encounter with them is a model too large to run on anything they own. Tencent and Alibaba Qwen cover the whole range instead, from under 1B upward.
  • Target: The chart splits the labs into two camps. Moonshot, MiniMax, Xiaomi and Z.ai publish almost nothing below 70B, so a developer's first encounter with them is a model too large to run on anything they own. Tencent and Alibaba Qwen cover the whole range instead, from under 1B upward.
  • Suggested fix: 차트는 연구소들을 두 진영으로 나눕니다. Moonshot, MiniMax, Xiaomi, Z.ai는 70B 미만 모델을 거의 공개하지 않으므로, 개발자가 이들의 모델을 처음 접할 때는 자신이 보유한 어떤 장비에서도 실행하기 어려울 만큼 큰 모델을 만나게 됩니다. 반면 Tencent와 Alibaba Qwen은 1B 미만부터 그 이상까지 전체 범위의 모델을 제공합니다.
  • Reason: 전체 문단이 한국어로 번역되지 않고 영어 원문 그대로 남아 있어 한국어 블로그에 게시할 수 없습니다.

QL-024 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.7200
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-025 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: Two things made the first camp possible. Building large stopped being a differentiator. Xiaomi and Meituan both cleared a trillion parameters this year, and neither was a household name in open weights twelve months ago. And a lab no longer has to ship a small model to be reachable, because the community's quantization layer will make a large one runnable within days, a dependency we return to below.
  • Target: Two things made the first camp possible. Building large stopped being a differentiator. Xiaomi and Meituan both cleared a trillion parameters this year, and neither was a household name in open weights twelve months ago. And a lab no longer has to ship a small model to be reachable, because the community's quantization layer will make a large one runnable within days, a dependency we return to below.
  • Suggested fix: 문단 전체를 한국어로 번역합니다. 예: "첫 번째 진영이 가능해진 데에는 두 가지 요인이 있었습니다. 대규모 모델을 만드는 것만으로는 더 이상 차별화가 되지 않았습니다. Xiaomi와 Meituan은 올해 모두 1조 개의 파라미터를 넘어섰으며, 12개월 전만 해도 두 회사 모두 오픈 웨이트 분야에서 대중적으로 알려진 이름은 아니었습니다. 또한 커뮤니티의 양자화 계층 덕분에 대규모 모델도 며칠 안에 실행할 수 있게 되므로, 연구소는 접근성을 확보하기 위해 더 이상 소형 모델을 출시할 필요가 없습니다. 이 의존성에 대해서는 아래에서 다시 살펴보겠습니다."
  • Reason: 전체 문단이 한국어로 번역되지 않고 영어 원문이 그대로 반복되어, 독자에게 필요한 번역 내용이 누락되었습니다.

QL-026 technical / major

  • Message: MQM judge technical score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM technical score is below threshold 0.75.

QL-027 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.0500
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-028 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: That leaves the size profile as a statement of intent rather than of capability. A frontier only portfolio stakes everything on benchmark position and API demand. A full spectrum portfolio is a bid to be the family developers standardise on. Both are rational, they are playing for different prizes.
  • Target: That leaves the size profile as a statement of intent rather than of capability. A frontier only portfolio stakes everything on benchmark position and API demand. A full spectrum portfolio is a bid to be the family developers standardise on. Both are rational, they are playing for different prizes.
  • Suggested fix: 크기 구성은 역량보다는 의도를 보여주는 표현이 됩니다. 최첨단 모델에만 집중하는 포트폴리오는 벤치마크 순위와 API 수요에 모든 것을 걸고 있습니다. 전체 스펙트럼을 아우르는 포트폴리오는 개발자들이 표준으로 삼을 모델군이 되겠다는 시도입니다. 두 전략 모두 합리적이지만, 서로 다른 목표를 추구합니다.
  • Reason: 전체 문단이 한국어로 번역되지 않고 영어 원문이 그대로 반복되어 한국어 독자에게 핵심 내용이 전달되지 않습니다.

QL-029 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.0500
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-030 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: The United States, meanwhile, is not absent from open source.
  • Target: The United States, meanwhile, is not absent from open source.
  • Suggested fix: 한편 미국도 오픈 소스 분야에서 빠지지 않습니다.
  • Reason: 번역문이 원문과 동일한 영어 문장으로 남아 있어 한국어 독자에게 핵심 의미가 전달되지 않습니다.

QL-031 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-032 accuracy / major

  • Message: MQM judge reported accuracy issue.
  • Source: New homegrown models
  • Target: New homegrown models
  • Suggested fix: 새로운 자체 개발 모델
  • Reason: 제목성 문구가 번역되지 않아 한국어 독자가 의미를 바로 이해하기 어렵습니다. ‘homegrown’은 자체 개발 또는 자국산이라는 의미로 문맥에 맞게 옮겨야 합니다.

QL-033 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: At the frontier scale, the picture is very different. Most U.S. releases above 100B parameters this year are not new models, but built on top of Chinese models. Only a few major original American models appear at this scale: Thinking Machines’ Inkling (952B), NVIDIA’s Nemotron 3 Ultra (561B), Nemotron 3 Super (124B), and Arcee AI’s Trinity-Large (399B).
  • Target: At the frontier scale, the picture is very different. Most U.S. releases above 100B parameters this year are not new models, but built on top of Chinese models. Only a few major original American models appear at this scale: Thinking Machines’ Inkling (952B), NVIDIA’s Nemotron 3 Ultra (561B), Nemotron 3 Super (124B), and Arcee AI’s Trinity-Large (399B).
  • Suggested fix: 최전선 규모에서는 상황이 매우 다릅니다. 올해 미국에서 출시된 1,000억 개 초과 파라미터 모델 대부분은 새로운 모델이 아니라 중국 모델을 기반으로 구축되었습니다. 이 규모에서 등장한 주요 오리지널 미국 모델은 Thinking Machines의 Inkling(952B), NVIDIA의 Nemotron 3 Ultra(561B), Nemotron 3 Super(124B), Arcee AI의 Trinity-Large(399B) 등 몇 개뿐입니다.
  • Reason: 전체 문단이 번역되지 않고 영어 원문 그대로 남아 있어 한국어 독자에게 제공할 핵심 내용이 누락되었습니다.

QL-034 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: AMD contributed many conversions but no original model at this scale. This work is still important: it enables trillion-parameter Chinese models to run efficiently on American hardware. But it represents a **distribution and optimization layer rather than model creation**.
  • Target: AMD contributed many conversions but no original model at this scale. This work is still important: it enables trillion-parameter Chinese models to run efficiently on American hardware. But it represents a **distribution and optimization layer rather than model creation**.
  • Suggested fix: AMD는 여러 변환 작업에 기여했지만, 이 규모의 독창적인 모델을 직접 개발한 것은 아닙니다. 그럼에도 이 작업은 중요합니다. 미국 하드웨어에서 수조 개 파라미터를 가진 중국어 모델을 효율적으로 실행할 수 있게 해 주기 때문입니다. 하지만 이는 모델 생성이 아니라 배포 및 최적화 계층에 해당합니다.
  • Reason: 전체 문단이 영어 원문 그대로 남아 있어 한국어 독자에게 번역 내용이 제공되지 않습니다. 이는 핵심 정보가 누락된 게시 차단 수준의 문제입니다.

QL-035 technical / major

  • Message: MQM judge technical score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM technical score is below threshold 0.75.

QL-036 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-037 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: Meanwhile, Chinese open models are increasingly optimized for domestic chips, the same competition in reverse, where models are designed around specific hardware ecosystems.
  • Target: Meanwhile, Chinese open models are increasingly optimized for domestic chips, the same competition in reverse, where models are designed around specific hardware ecosystems.
  • Suggested fix: 한편 중국 오픈 모델은 자국 칩에 맞춰 점점 더 최적화되고 있습니다. 이는 모델이 특정 하드웨어 생태계를 중심으로 설계되는, 반대 방향으로 진행되는 동일한 경쟁입니다.
  • Reason: 타깃 문장이 원문과 동일한 영어로 제공되어 한국어 번역이 누락되었습니다. 독자가 핵심 주장과 모델·하드웨어 생태계 간 관계를 한국어로 이해할 수 없습니다.

QL-038 technical / major

  • Message: MQM judge technical score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM technical score is below threshold 0.75.

QL-039 fluency / major

  • Message: MQM judge fluency score is low.
  • Target: 0.0000
  • Suggested fix: Review this segment against the MQM judge feedback.
  • Reason: MQM fluency score is below threshold 0.75.

QL-040 formatting / minor

  • Message: MQM judge reported formatting issue.
  • Source: **2\. Attention ≠ Adoption**
  • Target: **2\\. Attention ≠ Adoption**
  • Suggested fix: 2. Attention ≠ Adoption
  • Reason: 원문의 Markdown 이스케이프가 한 단계 더 추가되어 백슬래시가 중복되었습니다. 렌더링 결과가 달라지거나 제목 번호의 마크다운 처리가 어긋날 수 있습니다.

QL-041 fluency / major

  • Message: MQM judge reported fluency issue.
  • Source: We took the top 25 model repositories by downloads accumulated this year and the top 25 by likes. **Exactly one repository appears in both lists.**
  • Target: We took the top 25 model repositories by downloads accumulated this year and the top 25 by likes. **Exactly one repository appears in both lists.**
  • Suggested fix: 올해 누적 다운로드 수 기준 상위 25개 모델 리포지토리와 좋아요 수 기준 상위 25개 모델 리포지토리를 선정했습니다. 두 목록에 모두 포함된 리포지토리는 정확히 하나입니다.
  • Reason: 전체 문단이 번역되지 않고 영어 원문 그대로 남아 있어 한국어 블로그 독자에게 번역문으로 제공할 수 없습니다.

QL-042 accuracy / critical

  • Message: MQM judge reported accuracy issue.
  • Source: `We counted downloads inside the window rather than lifetime, so nothing is credited for merely having existed longer, and controlling for age makes the split sharper. Not one model published in 2026 reaches the download top 25, while thirteen of the twenty-five date from 2022. all-MiniLM-L6-v2 was pulled 1.55 billion times in seven months against 5,156 likes; Kimi-K3 was pulled about
SEO report — ❌ Fail

SEO Eval Report

Gate: 🟡 NEEDS_CHANGES — deterministic AND rubric

  • File: ../target/_posts/2026-08-14-state-of-open-models-summer-2026.md
  • Source: —
  • Primary keyword: (none — D5 skipped)
  • Mode: file

Gate

  • Status: NEEDS_CHANGES
  • Blockers: ✅ pass
  • Deterministic REQUIRED (D1–D7): ✅ pass
  • Rubric (R1–R6): ❌ fail (mean None, min None)

Blockers

✅ body_not_empty: Body is not empty
✅ robots_indexable: Robots is indexable
✅ internal_links_resolve: All internal links resolve
✅ local_images_resolve: All local images resolve

Required checks (gated)

✅ heading_hierarchy: Heading hierarchy: Valid
✅ alt_text_coverage: Alt text coverage: 13/13 images
✅ descriptive_alt_text: Descriptive alt text: 13/13 (≥80% recommend)
✅ image_files_exist: All 0 local image file(s) exist

OpenAI rubric checks

❌ semantic_metadata: NEEDS_CHANGES (required) — frontmatter.description is empty (missing meta description).
✅ alt_semantics: PASS (review) — https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/blog/state-of-open-models-summer-2026/dataset-growth.png: Describes the chart content and the key metric (cumulative growth to one million).; https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/blog/state-of-open-models-summer-2026/frontier-ceiling-by-country.png: Specifies the scoped comparison and timeframe shown in the image.; https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/blog/state-of-open-models-summer-2026/lab-size-strategy.png: Conveys the general takeaway about varying lab strategies.

Advisory checks (not gated)

✅ opening_summary: Opening 3 paragraphs: 56 words (recommend ≥50 for GEO)
✅ h1_count: Markdown H1 count: 1 (review against rendered layout)
✅ citations: Citations/statistics: 55 (recommend ≥1 for GEO)
⚠️ question_headings: Scannable H2/H3 (question or keyword): 0 (0 question, 0 keyword)
⚠️ internal_links: Internal links: 0 (recommend 2-3)
✅ word_count: Word count: 2773 (recommend ≥300)
ℹ️ primary_keyword: No primary_keyword in manifest — keyword check skipped
⚠️ webp_format: WebP format: 0/13 images (≥50% recommend)
⚠️ lazy_loading: Lazy loading: 0 images (optional)

Signals (evidence — not directly gated)

  • Frontmatter: title 21 chars, description 0 chars, author present True
  • Title text: 오픈 모델 현황: 2026년 여름 관찰
  • Description text: —
  • Opening text: * TOC
    {:toc}

이 글은 Hugging Face 블로그의 State of Open Models: Summer 2026 Observations를 한국어로 번역한 글입니다.

  • Opening: first paragraph 175 chars, first 3 paragraphs 401 chars
  • Headings: markdown H1 1, rendered effective H1 2, layout title H1 True
  • Links: total 26, external 26, internal 0, citation signals 55
  • Images: total 13, empty alt 0, filename-like alt 0, missing local files 0

Semantic review packet

  • Title: 오픈 모델 현황: 2026년 여름 관찰
  • Description: —
  • Rendered H1 candidates: 오픈 모델 현황: 2026년 여름 관찰, 오픈 모델 현황: 2026년 여름 관찰
  • Opening: * TOC
    {:toc}

이 글은 Hugging Face 블로그의 State of Open Models: Summer 2026 Observations를 한국어로 번역한 글입니다.

  • Canonical/permalink: —
  • Instruction: Compare title, description, rendered H1, and opening text for meaning consistency. This packet is evidence only; it does not decide pass/fail.

Frontmatter (advisory — written by metadata step, not gated)

✅ title: Title: 21 chars (recommend ≤60)
❌ description: Description is missing
✅ image: OG image: assets/images/blog/posts/2026-08-14-state-of-open-models-summer-2026/thumbnail.png
✅ categories: Categories: 2 (recommend 2-3)
✅ author: Author: dailybot

SEO metadata suggestion — SKIPPED

This is a suggestion. SEO is applied only when the post frontmatter is updated.
To apply safe fields from a partial suggestion, leave a trusted PR comment: metadata apply.

  • Auto apply: False
  • Requires human: True
  • Mode: frontmatter_only
  • Reason: metadata suggestion skipped because SEO gate did not pass

Candidate

  • No metadata candidate fields were produced.

Needs policy decision

  • None

Warnings

  • seo_gate_did_not_pass

@Jwaminju Jwaminju added the hf-agent:needs-human HF Agent needs human follow-up label Aug 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hf-agent:managed Opt PR into HF Agent review automation hf-agent:needs-human HF Agent needs human follow-up

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant