Skip to content

[CI Monitor] Daily Report - 2026-07-25 #146

Description

@amd-bot

Daily Cross-Workflow Summary — 2026-07-25

Snapshot: 2026-07-25 23:45 UTC · Only completed runs/jobs counted in trends · Auto-updated every 30 min

TL;DR

🟡 YELLOW · 2 🆕 today (R290 mi355x accuracy gate, R291 GHA platform error) · carry-over nightly families (R257 17d, R270 10d, R273, R262, R289) · R288 resolved (DCP hang did not recur) · in-flight fixes: R270→#31741 (blocked), R257→#31068/#32003, R290→#31727 (candidate)
👉 Today's ask: rerun the mi355x dsv4flash-fp8 leg 2–3× to separate R290 gate-flake from a real 3-day accuracy slide (0.917→0.914→0.905); rerun pr-test-amd to clear the transient R291 GHA platform error (it skipped all 22 AMD test jobs → no test signal on the latest run). Keep chasing #31741 (R270, 10th day) and #31068/#32003 (R257, 17th day). No fresh code regression.

Caution

pr-test-amd latest completed run 30157771919 (Jul-25 12:19) has NO test signal — a GitHub Actions internal platform error hit check-changes+gate jobs (R291), so all 22 downstream AMD test jobs were skipped, not tested. The real test-family failures come from the prior run 30136388246 (Jul-25 00:29). Rerun/update the branch before trusting a "green".

Caution

Jul-25 nightly-test-amd 30168349328 and nightly-test-amd-rocm720 30168262867 are IN-FLIGHT (queued, head 69a3c54c) → excluded from trends. Latest completed nightlies are the Jul-24 runs (30115195247 / 30115085856). pr-test-amd-rocm720 30168318210 also IN-FLIGHT; latest completed = Jul-24 30115153295.

Note

release-docker-amd-rocm720-nightly 30157811305 = startup_failure today (infra — workflow never provisioned, 0 jobs; was green Jul-18→24). release-docker-amd-nightly 30157965906 ✅ green. amd-aiter-scout: no run (cron Mon/Thu; today Sat).

Workflow status

Workflow Latest completed run ❌ jobs 7d trend (completed only, old→new) Δ vs yesterday
nightly-test-amd Jul-24 30115195247 (failure) 8 ❌·❌·⊘·❌·❌·❌·❌ ~same (R289 confirmed intermittent)
nightly-test-amd-rocm720 Jul-24 30115085856 (failure) 10 ❌·❌·⊘·❌·❌·❌·❌ R288 gone (DCP hang not reproduced)
nightly-amd-mi355x-disagg Jul-25 30143908399 (failure) 3 ✅·✅·✅·✅·✅·✅·❌ 🔴 went RED (R290 🆕)
release-docker-amd-nightly Jul-25 30157965906 0 ✅·✅·✅·✅·⚠️·✅·✅ ✅ still green
release-docker-amd-rocm720-nightly Jul-25 30157811305 ⚠️ startup_failure 0 (infra) ✅·✅·✅·✅·✅·✅·⚠️ ⚠️ startup_failure (infra)
nightly-amd-mi355x-disagg (see above)
amd-aiter-scout (no run today) n/a (cron Mon/Thu)
pr-test-amd Jul-25 30157771919 (failure — R291 infra) 4 (gate/infra) ❌·❌·❌·❌·❌·❌·❌ +R291 🆕 (GHA platform error)
pr-test-amd-rocm720 Jul-24 30115153295 (failure) 8 ❌·❌·❌·❌·❌·❌·❌ all known families

Failure clusters (deduplicated across all workflows)

R290 · 🆕 · GSM8K accuracy gate miss — DeepSeek-V4-Flash-FP8 1P1D (MI355X disagg) — mi355x-disagg · ⚠️ candidate #31727

  • Status: first RED for this workflow since Jul-18; config passed the prior 6 runs (Jul-19→24). This exact gate signature is intermittent: 3/13 completed runs (Jul-12, Jul-18, Jul-25). Marginal — 0.905 vs 0.91 gate.
  • Top hypothesis: [HIGH] the 0.91 GSM8K gate is too tight for this DSV4-Flash-FP8 DP8/EP8 config (real accuracy 0.905–0.914, ±0.005 sampling variance); runs straddle the bar. Disconfirming: the 3-day trend is monotonic (0.917→0.914→0.905), weakly more consistent with a small real drift than pure noise — a rerun is needed to decide.
  • In-flight fix: ❌ no threshold-retune PR; ⚠️ candidate #31727 (DSV4-Flash FP8 fused-RMS scale metadata on gfx950) could nudge numerics. Suggested triage: rerun the leg 2–3× at HEAD ebcb74ab; if it oscillates ≥0.91 it's gate flake, else bisect 99b29bf1..ebcb74ab (48 commits, FP8/MoE/DSV4 only).
Workflow Job (shard) Test File Test Function Error Log
nightly-amd-mi355x-disagg dsv4flash-fp8-1k1k-1p1d-dp8ep8 GSM8K gate (few_shot_gsm8k.py, 1319q) N/A (accuracy gate) accuracy=0.905 threshold=0.91 → failing before sweep Log
nightly-amd-mi355x-disagg collect-results rollup N/A fails because the leg above failed Log
nightly-amd-mi355x-disagg dsv4pro-fp4-1k1k-1p1d N/A (leg cancelled during setup) N/A cancelled ~7s (runner mia1-p01-g20 infra) — separate, benchmark never ran Log

R291 · 🆕 · GitHub Actions internal platform error (gate jobs, 0 steps) — pr-test-amd · ❌ transient infra

  • Status: isolated to the single run 30157771919 (Jul-25 12:22 UTC, 44 s window); prior run 12h earlier passed these jobs normally. Not a code regression.
  • Top hypothesis: [FACT] GitHub Actions platform internal error failed check-changes+gate/finish jobs before any step ran (0 steps, no runner, identical annotation on all 4 jobs). Because check-changes never emitted its change-filter outputs, all 22 downstream AMD test jobs were skipped → this run has zero test signal. Disconfirming: (none — directly observable).
  • In-flight fix: ❌ none (nothing to fix; platform-side). Suggested triage: simply rerun the run / update the branch; if it recurs, escalate to GitHub Actions infra.
Workflow Job (shard) Test File Test Function Error Log
pr-test-amd call-gate / pr-gate N/A (non-test) N/A GitHub Actions has encountered an internal error… Log
pr-test-amd check-changes N/A (non-test) N/A same internal error → 22 test jobs skipped Log

R257 · MI35x (gfx950) GPU-hang / HW-exception family — hipErrorCapturedEvent · GPU mem-fault · runner-loss — nightly-amd + nightly-rocm720 · ⚠️ #31068/#32003

  • Status: carry-over, 17th day (first Jul-09/10), never-passed on MI35x. Largest cluster: ~9 job-legs across both nightlies on the Jul-24 completed runs. Model-/workflow-independent; alternates on identical SHAs → environment/HW over a bisectable code regression.
  • Top hypothesis: [MEDIUM] a TP/MoE collective (aiter custom_all_reduce / IPC-meta gather) is captured into the decode HIP graph; the RCCL ProcessGroupNCCL watchdog's isCompleted()hipEventQuery on that event is illegal mid-capture → hipErrorCapturedEvent → uncaught → SIGABRT (exit -6); compounded by MI35x node/driver hangs & runner-loss. Disconfirming: alternates pass/fail on identical SHAs; Jul-13 showed a distinct GPU Hang exit 134 → not a clean code regression.
  • In-flight fix: ⚠️ #31068 (hung-GPU pre-flight recovery), #32003 (route scheduled jobs to MI300). Suggested triage: make MoE/TP all-reduce capture-safe on ROCm; chase #31068/#32003; validate MI35x pool health.
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd nightly-8-gpu-mi35x-qwen3-235b-mxfp4 test_qwen3_instruct_mxfp4.py TestQwen3Instruct2507MXFP4.setUpClass hipErrorCapturedEvent during decode graph capture → -6 Log
nightly-test-amd nightly-2-gpu-mi35x-glm51-mxfp4 test_qwen3_instruct_mxfp4.py setUpClass hipErrorCapturedEvent → -6 Log
nightly-test-amd-rocm720 nightly-8-gpu-mi35x-qwen3-235b-mxfp4-rocm720 test_qwen3_instruct_mxfp4.py setUpClass hipErrorCapturedEvent (aiter custom_all_reduce) → -6 Log
nightly-test-amd-rocm720 nightly-accuracy-8-gpu-mi35x-rocm720 test_qwen3_moe_eval_mi35x.py test_qwen3_moe_accuracy GPU memory access fault on Qwen3-30B-A3B prefill → abort Log
nightly-test-amd-rocm720 nightly-2-gpu-mi35x-glm51-mxfp4-rocm720 test_glm51_mxfp4_tp2_gsm8k_mi35x.py setUpClass runner lost communication (log blob missing) Log

R270 · Mamba JIT transfer kernel NameError on ROCm (CUDA-only import gate) — nightly-amd + nightly-rocm720 kernel · ✅ #31741 (blocked)

  • Status: carry-over, 10th day (first Jul-16), never-passed on AMD. 12 NameError subtests per kernel leg; identical on both Jul-24 completed runs.
  • Top hypothesis: [FACT] transfer_kv_mamba_* imports gated behind if _is_cuda: while the call site is reachable on HIP → NameError at runtime. Disconfirming: (none — directly observable).
  • In-flight fix: ✅ #31741 (open since Jul-20, mergeable_state: blocked). Suggested triage: rebase/merge #31741, then verify the JIT kernel actually compiles under ROCm, not just that the import resolves.
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd nightly-test-1-gpu-kernel test/registered/kernels/ops/mamba/test_transfer_mamba.py 12 subtests (roundtrip/single_item/full_indices) NameError: transfer_kv_mamba_lf_pf → exit 255 Log
nightly-test-amd-rocm720 nightly-test-1-gpu-kernel-rocm720 test_transfer_mamba.py same 12 subtests same Log

R289 · DeepSeek-V3.2 MTP GPU hang → scheduler watchdog timeout (MI35x, DSA backend) — nightly-amd · ⚠️ intermittent, no fix

  • Status: carry-over from Jul-24 (now 2nd data point), confirmed intermittent — passes ~4 of 6 recent nights; the identical signature already occurred Jul-22 on an earlier SHA, so no Jul23→24 commit introduced it. Not a fresh regression.
  • Top hypothesis: [LOW] intermittent MI35x (gfx950) GPU hang in the MTP draft-extend decode path — CPU parked on a HIP D2H sync (dsa_backend.py:793 .item()); classic GPU-side stall. Disconfirming: no window commit touches dsa_backend.py/eagle_worker; the .item() is where CPU waits on the GPU, not the root cause.
  • In-flight fix: ❌ none (#32209 is a different model/config). Suggested triage: rerun 2–3× same SHA to confirm flakiness before any bisect; sync with the MI35x debug-branch owners.
Workflow Job (shard) Test File Test Function Error Log
nightly-test-amd nightly-accuracy-8-gpu-mi35x-deepseek-v32-mtp test/registered/amd/accuracy/mi35x/test_deepseek_v32_mtp_eval_mi35x.py TestDeepseekV32TPMTP.test_a_gsm8k GSM8K stalls at 128/200 → watchdog → KeyError:'answer' → exit 255 Log
Known stable / recurring clusters (carrying over, no new action today) · click to expand
Cluster Workflow · Rep. job Test / fn Status In-flight fix
R273 accuracy near-misses: Qwen3-4B missing-threshold config gap, Mixtral/Mistral gsm8k answer-extraction, VLM MMMU shortfall, MiniMax-M2.5 (~0.929), DeepSeek-R1-MXFP4 MTP (~0.85), DeepSeek-V3.2 tilelang (~0.86–0.90) nightly-amd acc-2gpu · nightly-rocm720 dsv32 test_gsm8k_eval_amd.py / test_*mxfp4*.py / test_vlms_mmmu_eval_amd.py marginal/boundary gates, recurring ⚠️ recalib precedent #31702; Qwen3-4B gap #26730
R262 InternVL2_5-2B CLIPImageProcessor has no .tokenizer — server crash nightly-rocm720 vlm test_vlms_mmmu_eval_amd.py (InternVL2_5-2B) recurring since Jul-14
DeepSeek-V4-Pro NextN/MTP server-launch timeout (slow CPU FP8 wo_a dequant) nightly-rocm720 dsv4-pro-mtp test_deepseek_v4_pro_*_mtp.py recurring (triton-only)
Diffusion (multimodal-gen) family: 1-/2-GPU runner-loss, 60/90-min step timeout, FLUX.2 FP8 HIPBLAS_STATUS_NOT_SUPPORTED, update_weights_from_disk 400 shape-mismatch pr-test-amd mm-gen 1gpu (0) · pr-rocm720 mm-gen 1gpu (0) ×many test_server_1_gpu.py / test_server_2_gpu.py never-passed/recurring ⚠️ skip #31925
Disaggregation family: Mooncake KV-transfer timeout, PD+PP event-loop hang (KeyError:'answer'), Kimi-K2.5-MXFP4 load monitored_barrier SIGKILL pr-test-amd disagg mi35x · stage-c kimi test_disaggregation_pp.py / test_kimi_k25_mxfp4_bcg_mi35x.py recurring/flaky ⚠️ #31800
ModuleNotFoundError: msgpack on ROCm 7.2 image (diffusion extra not installed) pr-rocm720 mm-gen-unit test_msgpack_roundtrip_* never-passed since Jul-11 #31899
MXFP4 GSM8K just below strict gate (Qwen3-30B-A3B online FP8→MXFP4 requant) pr-rocm720 stage-b-small-mi35x test_*_mxfp4_*.py borderline eval flake
Triton offline throughput below AMD 3500 tok/s gate pr-rocm720 stage-b-1gpu-large test_bench_serving_* never-passed
GPU-Hang HW-exception / HSA-OOM (infra): sgl-kernel build exit-134, unified-radix 2-GPU watchdog, DeepSeek-V3.2 TP8 OUT_OF_RESOURCES pr-test-amd stage-b-1gpu-small (8) · pr-rocm720 stage-c-8gpu infra/timeout recurring ⚠️ #31068, #32003
Workflow drill-down · click to expand

nightly-test-amd · latest completed 30115195247 (Jul-24) · 8 failures: R257 (qwen3-235b-mxfp4 ×2 legs, hipErrorCapturedEvent) · R270 (kernel) · R289 (dsv32-mtp watchdog hang) · R273 (acc-2gpu Mixtral+Qwen3-4B config, acc-2gpu-vlm MMMU, minimax-m25, nightly-4-gpu GLM-4.1V MMMU timeout). Jul-25 30168349328 IN-FLIGHT.

nightly-test-amd-rocm720 · latest completed 30115085856 (Jul-24) · 10 failures: R270 (kernel) · R257 (qwen3-235b-mxfp4, acc-8gpu mem-fault, glm51 runner-loss) · R262 (InternVL2_5) · R273 (dsv32, dsr1-mxfp4 MTP, minimax-m25, Mixtral/Qwen3-4B) · dsv4-pro-mtp launch timeout · dsv4-flash. R288 (DCP hang) did NOT recurqwen35-triton-dcp absent from today's failures. Jul-25 30168262867 IN-FLIGHT.

nightly-amd-mi355x-disagg · 30143908399 (Jul-25) · 3 failures = R290 🆕 (dsv4flash-fp8 gate 0.905<0.91) + collect-results rollup + 1 unrelated leg cancellation (runner mia1-p01-g20 infra). Was green Jul-19→24.

pr-test-amd · latest completed 30157771919 (Jul-25 12:19) · R291 🆕 GHA platform error → 4 gate/finish jobs failed, 22 test jobs skipped (no signal). Prior run 30136388246 (Jul-25 00:29): diffusion + disagg families only.

pr-test-amd-rocm720 · latest completed 30115153295 (Jul-24) · 8 failures: diffusion family (FLUX.2 FP8/timeout/runner-loss) · msgpack ModuleNotFound · triton throughput gate · MXFP4 borderline · stage-c-8gpu HSA-OOM (DeepSeek-V3.2). Jul-25 30168318210 IN-FLIGHT.

How this report is generated

  • Only status == "completed" runs/jobs counted in trends. Jul-25 nightly-amd/rocm720 and pr-rocm720 runs are queued (IN-FLIGHT) → excluded; latest completed nightlies are the Jul-24 runs. mi355x-disagg Jul-25 and pr-test-amd Jul-25 12:19 completed and are used directly.
  • NEW today: R290 mi355x GSM8K gate miss (workflow flipped green→red, marginal/flaky 3/13, candidate #31727); R291 GitHub Actions platform error (pr-test-amd, isolated single run, 22 test jobs skipped).
  • Resolved/not-reproduced: R288 DCP all-reduce hang — qwen35-triton-dcp-rocm720 passed in the Jul-24 completed run.
  • Carry-over: R257 MI35x hang family (17th day), R270 mamba NameError (10th day, #31741 blocked), R273 accuracy near-misses, R262 InternVL, R289 DSV3.2-MTP (confirmed intermittent).
  • Infra: release-docker-amd-rocm720-nightly startup_failure (Jul-25) — platform-side, 0 jobs, no cluster.
  • Confidence: FACT/HIGH/MEDIUM/LOW/SPECULATION. Bot does NOT assign Priority — engineers decide from cluster size + persistence + fix availability.

Generated by amd-bot · last updated 2026-07-25 23:45 UTC


Generated by amd-bot using Claude Code CLI (last updated: 2026-07-25 23:45 UTC)


CI Monitor — 2026-07-25

Repo: sgl-project/sglang

Monitored Workflows:

  • nightly-test-amd.yml
  • nightly-test-amd-rocm720.yml
  • release-docker-amd-nightly.yml
  • release-docker-amd-rocm720-nightly.yml
  • nightly-amd-mi355x-disagg.yml
  • amd-aiter-scout.yml
  • pr-test-amd.yml
  • pr-test-amd-rocm720.yml

Per-workflow failure reports are appended as comments below; the Daily Cross-Workflow Summary is rendered above this section.

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions