Skip to content

[serge] integration failure triage - 2026-08-18 #48050

Description

@github-actions

Automated integration-failure triage for the daily CI window 2026-08-12 → 2026-08-18.

This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.

Serge dispatched one task per failure group below — each opens or updates its own PR on a serge/fix/itf-<fingerprint> branch. This table is refreshed in place as Serge runs: a group links its #<pr> when opened, shows 🚫 no fix when Serge found no safe change, ⚠️ task failed on error, or (pending) while still running (a late PR links on the next nightly run).

Dispatched failure groups

Model Error Occurrences PR
kimi_k25 (regressed by PR #47573) mixed — other (2), tensor values differ (2) 4 🚫 no fix
generation output_mismatch — list output differs (4) 4 #48070
qwen3_moe other — other (4) 4 🚫 no fix
exaone4 OOM — other (1) 1 🚫 no fix
gpt_oss import_or_config — other (1) 1 🚫 no fix
cosmos3_omni (regressed by PR #47096) mixed — list output differs (2) 2 #48060
deepseek_vl other — other (3) 3 ⚠️ task failed
qwen3_vl_moe OOM — other (1) 1 🚫 no fix
voxtral_realtime (regressed by PR #47622) mixed — other (2) 2 ⚠️ task failed
phimoe other — other (2) 2 🚫 no fix

Not dispatched — environment / dependency

These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).

5 models ran out of device memory (13 failures) — needs runner capacity, not a patch, so none of these were dispatched: bamba (4), llama4 (3), cohere2_vision (2), gemma4 (2), pi0 (2).

Outcome recap

Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.

Group Reason LLM Tokens (in / out) Failing tests
kimi_k25 (regressed by PR #47573) [not_fixed] GPU verification did not confirm the fix (not_fixed): the patch did not turn the targeted tests green. moonshotai/Kimi-K2.7-Code 2,067,305 / 3,048 test_model_logits · test_model_logits_batched
qwen3_moe [reproduced] GPU reproduce + classify: this is an ENVIRONMENT issue (The traceback shows a 4-bit quantized 15B model being dispatched to CPU/disk because the two-GPU CI runner lacks enough GPU RAM, which is a hardware/environment capacity problem, not a library bug or stale test expectation.) — no source patch can fix it, so no investigation and no PR. moonshotai/Kimi-K2.7-Code — / — test_model_15b_a2b_generation · test_model_15b_a2b_logits · test_model_15b_a2b_long_prompt_sdpa · +1 more
exaone4 [reproduced] GPU reproduce + classify: this is an ENVIRONMENT issue (The failure is a CUDA out-of-memory error during a large-model generation integration test, which is a property of the GPU memory capacity and allocation state rather than a bug in the model or test logic.) — no source patch can fix it, so no investigation and no PR. moonshotai/Kimi-K2.7-Code — / — test_model_generation_beyond_sliding_window
gpt_oss [reproduced] GPU reproduce + classify: this is an ENVIRONMENT issue (The traceback is a clean ImportError stating the kernels package is missing or has an incompatible version, which is an environment/dependency problem, not a bug in the library code or test expectations.) — no source patch can fix it, so no investigation and no PR. moonshotai/Kimi-K2.7-Code — / — test_model_outputs_02

Generated 2026-08-18T23:02:10+00:00.

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions