Automated integration-failure triage for the daily CI window 2026-08-12 → 2026-08-18.
This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.
Serge dispatched one task per failure group below — each opens or updates its own PR on a serge/fix/itf-<fingerprint> branch. This table is refreshed in place as Serge runs: a group links its #<pr> when opened, shows 🚫 no fix when Serge found no safe change, ⚠️ task failed on error, or (pending) while still running (a late PR links on the next nightly run).
Dispatched failure groups
| Model |
Error |
Occurrences |
PR |
kimi_k25 (regressed by PR #47573) |
mixed — other (2), tensor values differ (2) |
4 |
🚫 no fix |
generation |
output_mismatch — list output differs (4) |
4 |
#48070 |
qwen3_moe |
other — other (4) |
4 |
🚫 no fix |
exaone4 |
OOM — other (1) |
1 |
🚫 no fix |
gpt_oss |
import_or_config — other (1) |
1 |
🚫 no fix |
cosmos3_omni (regressed by PR #47096) |
mixed — list output differs (2) |
2 |
#48060 |
deepseek_vl |
other — other (3) |
3 |
⚠️ task failed |
qwen3_vl_moe |
OOM — other (1) |
1 |
🚫 no fix |
voxtral_realtime (regressed by PR #47622) |
mixed — other (2) |
2 |
⚠️ task failed |
phimoe |
other — other (2) |
2 |
🚫 no fix |
Not dispatched — environment / dependency
These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).
5 models ran out of device memory (13 failures) — needs runner capacity, not a patch, so none of these were dispatched: bamba (4), llama4 (3), cohere2_vision (2), gemma4 (2), pi0 (2).
Outcome recap
Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.
| Group |
Reason |
LLM |
Tokens (in / out) |
Failing tests |
kimi_k25 (regressed by PR #47573) |
[not_fixed] GPU verification did not confirm the fix (not_fixed): the patch did not turn the targeted tests green. |
moonshotai/Kimi-K2.7-Code |
2,067,305 / 3,048 |
test_model_logits · test_model_logits_batched |
qwen3_moe |
[reproduced] GPU reproduce + classify: this is an ENVIRONMENT issue (The traceback shows a 4-bit quantized 15B model being dispatched to CPU/disk because the two-GPU CI runner lacks enough GPU RAM, which is a hardware/environment capacity problem, not a library bug or stale test expectation.) — no source patch can fix it, so no investigation and no PR. |
moonshotai/Kimi-K2.7-Code |
— / — |
test_model_15b_a2b_generation · test_model_15b_a2b_logits · test_model_15b_a2b_long_prompt_sdpa · +1 more |
exaone4 |
[reproduced] GPU reproduce + classify: this is an ENVIRONMENT issue (The failure is a CUDA out-of-memory error during a large-model generation integration test, which is a property of the GPU memory capacity and allocation state rather than a bug in the model or test logic.) — no source patch can fix it, so no investigation and no PR. |
moonshotai/Kimi-K2.7-Code |
— / — |
test_model_generation_beyond_sliding_window |
gpt_oss |
[reproduced] GPU reproduce + classify: this is an ENVIRONMENT issue (The traceback is a clean ImportError stating the kernels package is missing or has an incompatible version, which is an environment/dependency problem, not a bug in the library code or test expectations.) — no source patch can fix it, so no investigation and no PR. |
moonshotai/Kimi-K2.7-Code |
— / — |
test_model_outputs_02 |
Generated 2026-08-18T23:02:10+00:00.
Automated integration-failure triage for the daily CI window
2026-08-12 → 2026-08-18.This issue was generated by AI-assisted automation. The grouping, summaries, and recommended follow-up can be incomplete or misleading; verify the failures before acting.
Serge dispatched one task per failure group below — each opens or updates its own PR on a
serge/fix/itf-<fingerprint>branch. This table is refreshed in place as Serge runs: a group links its#<pr>when opened, shows🚫 no fixwhen Serge found no safe change,⚠️ task failedon error, or(pending)while still running (a late PR links on the next nightly run).Dispatched failure groups
kimi_k25(regressed by PR #47573)generationqwen3_moeexaone4gpt_osscosmos3_omni(regressed by PR #47096)deepseek_vlqwen3_vl_moevoxtral_realtime(regressed by PR #47622)phimoeNot dispatched — environment / dependency
These groups were triaged but NOT handed to Serge: their failure mode is a property of the runner or the environment, so no minimal source patch can fix them. They need a human (runner capacity, a dependency pin).
5 models ran out of device memory (13 failures) — needs runner capacity, not a patch, so none of these were dispatched:
bamba(4),llama4(3),cohere2_vision(2),gemma4(2),pi0(2).Outcome recap
Why each group that opened no PR ended without one, and what it cost — surfaced here from the Serge dashboard. Failing tests links straight to each test's dashboard view, since these are the groups a human has to pick up.
kimi_k25(regressed by PR #47573)not_fixed): the patch did not turn the targeted tests green.moonshotai/Kimi-K2.7-Codeqwen3_moemoonshotai/Kimi-K2.7-Codeexaone4moonshotai/Kimi-K2.7-Codegpt_osskernelspackage is missing or has an incompatible version, which is an environment/dependency problem, not a bug in the library code or test expectations.) — no source patch can fix it, so no investigation and no PR.moonshotai/Kimi-K2.7-CodeGenerated 2026-08-18T23:02:10+00:00.