Skip to content

ci: preserve sanitizer diagnostics and failed suite logs - #2487

Open
DeusData wants to merge 3 commits into
mainfrom
ci/failure-summary-sanitizer-reports
Open

DeusData wants to merge 3 commits into
mainfrom
ci/failure-summary-sanitizer-reports

Conversation

@DeusData

@DeusData DeusData commented Oct 2, 2026

Copy link
Copy Markdown
Owner

When a suite aborted under a sanitizer, the job summary could show only the shadow-byte legend because it searched for FAIL. The harness now includes the running test and the useful beginning of sanitizer or crash reports, and retains failed suite logs plus results.txt for diagnosis.

Two commits separate the harness behavior from the workflow addition. The four parallel-test jobs upload those failed logs after a job failure, using the existing pinned upload action and seven-day retention.

Validation:

  • All 63 failure-report contract checks passed on macOS, Linux, and real Windows; removing the behavior previously failed 30 checks.
  • Parallel-harness, executable-bit, shell line-ending, and venue-parity contracts passed.
  • A current macOS ASan/UBSan build passed the bounded suite run (63 tests) and real parallel-harness shard (30 tests).
  • Existing assertion-only failure output remains byte-identical, and log collection preserves exit status.

The full three-OS batch gate and the first remote artifact upload remain to be verified.

A Windows sanitizer job went red on one suite that aborted under
AddressSanitizer, and the job log showed an empty "every failure site"
section followed by the shadow-byte legend. The summary in
scripts/run-tests-parallel.sh searched the suite log for "FAIL" only. A
sanitizer report never contains that word, and its useful part (error
kind, access size, top frames) is at the start of the report, far above
the last 15 lines.

The summary now lives in scripts/suite-failure-report.sh, which the
harness sources. Between the unchanged failure-site and last-15-lines
sections it prints, for a failing suite, the test that was running and
the first sanitizer report. The framework flushes each test name before
the test body runs, so the last test line ahead of a report is the test
that was running; a report that comes after the last test finished is
labelled as such instead of blaming a test. The report is shown from its
header (AddressSanitizer, LeakSanitizer, ThreadSanitizer,
MemorySanitizer, UndefinedBehaviorSanitizer) to its own SUMMARY line, at
most 40 lines, without the shadow dump and the legend; a recoverable
"runtime error:" line is shown with its stack only. Further sanitizer
headers, "runtime error:" lines and SUMMARY lines follow (at most 10),
then signal and timeout lines (at most 5).

A suite that dies with no report and no completion summary gets the
running test named. A trap-mode UBSan build stops on an illegal
instruction and prints nothing, so that name is the only evidence its
log holds. A log with an assertion failure and no report prints exactly
what it printed before.

The same file keeps the logs a red run needs. An EXIT trap in the
harness copies the log of every suite that did not record rc=0, plus
results.txt, into test-logs/failed/. The trap covers the scheduler
giving up and the union guard as well as the normal end, and it never
changes the exit status.

scripts/run-test-wave.py reported pass=0 for a suite that died without
its summary line. In that case, and only then, it now counts the PASS
markers in the log. A suite with a summary line is never recounted, so
the totals of a green run are unchanged.

tests/test_failure_report_contract.sh drives the real functions with
fabricated suite logs and runs as Step 0i2 of scripts/test.sh. It fails
on the previous behaviour in every sanitizer, running-test and
result-line check and passes on it in the unchanged-output checks. The
awk code uses no interval expressions and no character classes; the test
passes with BSD awk and grep on macOS and with mawk and GNU grep on
Ubuntu. scripts/README.md names the new internal file.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
The job log of a red test job carries a summary of each failing suite,
not its log, and the only artifact was shard-manifest.txt. A failure
whose summary was not enough could only be understood by reproducing it.

The four jobs in .github/workflows/_test.yml that run the parallel
harness (test-unix, test-windows, test-diag, test-lsan-macos) get one
more step, "Failing suite logs". It runs only when the job failed and
uploads build/c/test-logs/failed/: the logs of the suites that did not
end green, plus results.txt, as the harness leaves them there. One red
suite is typically a few kilobytes; a whole run's logs are about 25 MB.
Retention is 7 days. The artifact names follow the shard-manifest ones
(suite-logs-<os>-<compiler or msystem>-<job index>, suite-logs-diag,
suite-logs-lsan-macos), and the step uses the upload-artifact reference
the file already pins.

A job that failed outside the suite waves has no such directory; the
step then warns and uploads nothing. The step cannot change a job's
result: it is skipped on a green job, and a failed job stays failed
whatever the upload does. Triggers, permissions, job dependencies,
timeouts, the matrix and the shard-completeness download pattern are
unchanged.

Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
Signed-off-by: Martin Vogel <martin.vogel.tech@gmail.com>
@DeusData

DeusData commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

Checkpoint and handover (2026-10-03 UTC)

Published head: 743c6dc01beed5398929b35b9288d9384d37bb22. The remote head was verified.

The CI summary implementation and harness are published. This checkpoint also commits the two previously staged executable-bit changes for the failure-report helper and contract script; both source blobs are byte-identical. Historical coverage includes 63 contracts across the recorded platform runs and macOS sanitizer evidence; no fresh runtime checks were run for this mode-only follow-up. VM lifecycle/ownership and mixed-version cleanup compatibility questions remain open; passing checks alone do not settle those policy choices.

Hosted snapshot at 2026-10-03 22:02:11 UTC: 3 queued. Confirm the required checks on this exact head before treating it as ready.

No additional local build, test, lint, sanitizer, benchmark or CI runs were performed at this checkpoint, as requested. Earlier executed evidence remains historical; prepared tests and the newer source-reviewed changes must still be validated by the hosted gate.

The campaign is paused at the maintainer’s request. Local monitoring has stopped; hosted jobs remain running. No merge was performed. Thanks for reviewing this change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant