Skip to content

[serge] integration failure triage - 2026-09-11 #48724

Description

@github-actions

Automated integration-failure triage for the daily CI window 2026-09-05 → 2026-09-11.

This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.

Serge dispatched one task per failure group below — each opens or updates its own PR on a serge/fix/itf-<fingerprint> branch. This table is refreshed in place as Serge runs: a group links its #<pr> when opened, shows 🚫 no fix when Serge found no safe change, ⚠️ task failed on error, or (pending) while still running (a late PR links on the next nightly run).

Dispatched failure groups

Model Error Occurrences PR
inkling (regressed by PR #47827) mixed — tensor values differ (4) 4 #48667
generation cuda_runtime — other (12) 12 🚫 no fix
blt import_or_config — other (4) 4 ⚠️ task failed
depth_anything output_mismatch — tensor values differ (2) 2 #48725
kimi_k25 other — other (2) 2 🚫 no fix

Not dispatched — environment / dependency

These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).

2 models ran out of device memory (7 failures) — needs runner capacity, not a patch, so none of these were dispatched: bamba (4), llama4 (3).

18 group(s) skipped — stopped failing before this run (4 integration tests regressed by commit ce5c8f5e4352 (PR #47988), musicgen_melody, glm_ocr, hunyuan_vl +14 more). Not dispatched and not awaiting a human.

Outcome recap

Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.

Group Reason LLM Tokens (in / out) Failing tests
generation [not_reproduced] GPU reproduce: the targeted tests did NOT fail at the base commit (not_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR. moonshotai/Kimi-K2.7-Code — / — test_validate_assistant · test_matches_immediate_check_at_max_length · test_matches_immediate_check_on_a_batch · +3 more
blt LLM returned unparseable output (finish_reason=stop, 49 LLM turns · 49 tool calls · 74.8s · 2150112 in / 11702 out tokens, 1 salvage re-ask(s) did not recover it) moonshotai/Kimi-K2.7-Code 2,150,112 / 11,702 test_model_logits · test_model_logits_bf16
kimi_k25 [not_fixed] GPU verification did not confirm the fix (not_fixed): the patch did not turn the targeted tests green. moonshotai/Kimi-K2.7-Code 2,083,714 / 19,205 test_model_logits · test_model_logits_batched

Generated 2026-09-12T00:16:31+00:00.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions