Automated integration-failure triage for the daily CI window 2026-09-05 → 2026-09-11.
This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.
Serge dispatched one task per failure group below — each opens or updates its own PR on a serge/fix/itf-<fingerprint> branch. This table is refreshed in place as Serge runs: a group links its #<pr> when opened, shows 🚫 no fix when Serge found no safe change, ⚠️ task failed on error, or (pending) while still running (a late PR links on the next nightly run).
Dispatched failure groups
| Model |
Error |
Occurrences |
PR |
inkling (regressed by PR #47827) |
mixed — tensor values differ (4) |
4 |
#48667 |
generation |
cuda_runtime — other (12) |
12 |
🚫 no fix |
blt |
import_or_config — other (4) |
4 |
⚠️ task failed |
depth_anything |
output_mismatch — tensor values differ (2) |
2 |
#48725 |
kimi_k25 |
other — other (2) |
2 |
🚫 no fix |
Not dispatched — environment / dependency
These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).
2 models ran out of device memory (7 failures) — needs runner capacity, not a patch, so none of these were dispatched: bamba (4), llama4 (3).
18 group(s) skipped — stopped failing before this run (4 integration tests regressed by commit ce5c8f5e4352 (PR #47988), musicgen_melody, glm_ocr, hunyuan_vl +14 more). Not dispatched and not awaiting a human.
Outcome recap
Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.
| Group |
Reason |
LLM |
Tokens (in / out) |
Failing tests |
generation |
[not_reproduced] GPU reproduce: the targeted tests did NOT fail at the base commit (not_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR. |
moonshotai/Kimi-K2.7-Code |
— / — |
test_validate_assistant · test_matches_immediate_check_at_max_length · test_matches_immediate_check_on_a_batch · +3 more |
blt |
LLM returned unparseable output (finish_reason=stop, 49 LLM turns · 49 tool calls · 74.8s · 2150112 in / 11702 out tokens, 1 salvage re-ask(s) did not recover it) |
moonshotai/Kimi-K2.7-Code |
2,150,112 / 11,702 |
test_model_logits · test_model_logits_bf16 |
kimi_k25 |
[not_fixed] GPU verification did not confirm the fix (not_fixed): the patch did not turn the targeted tests green. |
moonshotai/Kimi-K2.7-Code |
2,083,714 / 19,205 |
test_model_logits · test_model_logits_batched |
Generated 2026-09-12T00:16:31+00:00.
Automated integration-failure triage for the daily CI window
2026-09-05 → 2026-09-11.This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.
Serge dispatched one task per failure group below — each opens or updates its own PR on a
serge/fix/itf-<fingerprint>branch. This table is refreshed in place as Serge runs: a group links its#<pr>when opened, shows🚫 no fixwhen Serge found no safe change,⚠️ task failedon error, or(pending)while still running (a late PR links on the next nightly run).Dispatched failure groups
inkling(regressed by PR #47827)generationbltdepth_anythingkimi_k25Not dispatched — environment / dependency
These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).
2 models ran out of device memory (7 failures) — needs runner capacity, not a patch, so none of these were dispatched:
bamba(4),llama4(3).18 group(s) skipped — stopped failing before this run (
4 integration tests regressed by commit ce5c8f5e4352 (PR #47988),musicgen_melody,glm_ocr,hunyuan_vl+14 more). Not dispatched and not awaiting a human.Outcome recap
Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.
generationnot_reproduced) — the failure is stale, already fixed, flaky, or environment-specific. Skipping this group; no investigation, no PR.moonshotai/Kimi-K2.7-Codebltmoonshotai/Kimi-K2.7-Codekimi_k25not_fixed): the patch did not turn the targeted tests green.moonshotai/Kimi-K2.7-CodeGenerated 2026-09-12T00:16:31+00:00.