Skip to content

Report perfcollect initialization failures in controller output - #901

Open
LoopedBard3 wants to merge 4 commits into
dotnet:mainfrom
LoopedBard3:loopedbard3-linux-trace-collection
Open

LoopedBard3 wants to merge 4 commits into
dotnet:mainfrom
LoopedBard3:loopedbard3-linux-trace-collection

Conversation

@LoopedBard3

@LoopedBard3 LoopedBard3 commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

Summary

Legacy Linux trace collection can exit during initialization (for example, LTTng not installed) while the benchmark continues. Crank currently reports Trace collected, then shows only an HTTP 404 when downloading the missing trace.

  • Retain the last 16,384 characters of perfcollect stdout/stderr and record launch/nonzero-exit failures in the existing Job.Error field.
  • Keep each legacy collector process and completion task in its job's JobContext, including native/Docker startup, finalization, abort, and deletion cleanup.
  • Await collector completion and drained output before disposal. Report success only when the collector exits successfully and the trace exists and is nonempty; a nonempty partial trace does not turn collector failure into success.
  • Reject missing or zero-length files at the trace download endpoint, without deleting diagnostic artifacts.
  • Refresh job details after a trace-download 404 with a five-second cancellation deadline. Use only a successfully refreshed error; preserve the original HTTP diagnostic on failure, cancellation, an unsuccessful response, or a missing error rather than displaying a stale job error.
  • Document the LTTng-UST 2.12/2.13 userspace ABI incompatibility and legacy installer's limitations on newer Ubuntu releases. Recommend collect-linux for supported .NET 10+ Linux hosts, including kernel, root, user-events, and tracefs prerequisites; do not recommend disabling LTTng.

No new collection options or serialized job states. The benchmark is not aborted when the collector fails; existing controller handling of Job.Error can return a nonzero exit code even when the benchmark itself succeeds.

Documentation references: dotnet/runtime#57784, perfcollect installer, and collect-linux prerequisites.

Validation

Passed 104 related tracing unit tests on Windows, including all four Copilot review regressions:

dotnet test .\test\Microsoft.Crank.UnitTests\Microsoft.Crank.UnitTests.csproj --no-restore --filter "FullyQualifiedName~PerfcollectTests|FullyQualifiedName~TraceDownloadTests|FullyQualifiedName~JobsControllerTests.TraceRejectsMissingOrEmptyFilesAndPreservesNonEmptyDownloads|FullyQualifiedName~DotNetTrace|FullyQualifiedName~TraceExtensionsTests|FullyQualifiedName~JobMixedVersionCompatTests" --verbosity quiet

Coverage includes launch/exit failures, bounded diagnostics, successful collection, partial/missing/empty traces, stopping one job while another collector remains running, endpoint rejection of empty traces, and stale-error/refresh-timeout fallback. The timeout regression uses an infinite HttpClient timeout and a handler that waits until the request's cancellation token is canceled.

A quick local review by GPT-5.6 Sol Fast identified a refresh-cancellation fallback issue, which was fixed; the subsequent Copilot PR review findings are addressed in 5a9fdc9.

The process tests use fake collectors, and the controller tests use an in-memory HTTP handler. A real Linux perfcollect capture was not run.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@LoopedBard3
LoopedBard3 marked this pull request as draft October 5, 2026 21:36
@LoopedBard3
LoopedBard3 marked this pull request as draft October 5, 2026 21:36
LoopedBard3 and others added 2 commits October 5, 2026 15:39
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Collector state is not job-scoped, and several failure paths can still report or download invalid traces incorrectly.

Review effort: Balanced
Findings: 2 High severity · 2 Medium severity

Open (4)
What changed in this PR

Improves legacy Linux trace failure reporting and documents modern tracing alternatives.

Changes:

  • Captures bounded perfcollect diagnostics and validates trace output.
  • Refreshes agent errors after failed trace downloads.
  • Adds regression tests and Linux tracing guidance.
File Description
Startup.cs Tracks perfcollect completion, errors, and trace validity.
JobConnection.cs Retrieves collector errors after trace 404s.
PerfcollectTests.cs Tests collector failures and trace validation.
TraceDownloadTests.cs Tests controller error reporting and fallback.
setup_linux.md Documents LTTng compatibility and collect-linux.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/Microsoft.Crank.Agent/Startup.cs Outdated
Comment thread src/Microsoft.Crank.Agent/Startup.cs Outdated
Comment thread src/Microsoft.Crank.Agent/Startup.cs
Comment thread src/Microsoft.Crank.Controller/JobConnection.cs Outdated
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@LoopedBard3 LoopedBard3 self-assigned this Oct 6, 2026
@LoopedBard3
LoopedBard3 marked this pull request as ready for review October 6, 2026 21:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants