Skip to content

Avoid reporting expected BuildHost shutdowns - #85853

Draft
mwiemer-microsoft wants to merge 3 commits into
mainfrom
mwiemer-fix-buildhost-shutdown-race
Draft

mwiemer-microsoft wants to merge 3 commits into
mainfrom
mwiemer-fix-buildhost-shutdown-race

Conversation

@mwiemer-microsoft

@mwiemer-microsoft mwiemer-microsoft commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Ref roslyn-CI failure (link will expire around 2026-10-09):

[xUnit.net 00:03:15.48]     Microsoft.CodeAnalysis.MSBuild.UnitTests.NetCoreTests.TestOpenProject_FileBasedApp_RefDirective_Self [FAIL]
Command: exec "/opt/hostedtoolcache/dotnet/sdk/11.0.100-rc.1.26425.128/vstest.console.dll" @"/mnt/vss/_work/1/s/artifacts/tmp/Debug/vstest-rsp/vstest_58.rsp"
xUnit output log: /mnt/vss/_work/1/s/artifacts/log/Debug/xUnitFailure-Microsoft.CodeAnalysis.Workspaces.MSBuild.UnitTests_58.log
Unhandled exception. System.IO.DirectoryNotFoundException: Could not find a part of the path '/mnt/vss/_work/1/s/artifacts/log/Debug/xUnitFailure-Microsoft.CodeAnalysis.Workspaces.MSBuild.UnitTests_58.log'.

Full persistent log: 2026-09-30.23-roslyn-CI-log-224-flaky-failure-pr-85853.txt


AI-generated:

Problem

When BuildHostProcessManager intentionally shuts down a BuildHost, the RPC stream closes before the OS necessarily reports the child process as exited. The disconnect handler can observe HasExited == false during this window and log The BuildHost process is not responding, even though the BuildHost is exiting normally. That failure diagnostic caused NetCoreTests.TestOpenProject_FileBasedApp_RefDirective_Self to fail its Assert.Empty(workspace.Diagnostics) assertion on loaded Linux CI. The CI output ended with the BuildHost's normal RPC channel closed; process exiting. message. RunTests then failed writing its xUnitFailure log because artifacts/log/Debug was missing, obscuring the assertion in the job log; AzDO's structured test result retained it.

Pre-fix reproduction command

This repeats the exact test that failed in Linux CI; on an unpatched checkout, Linux or a heavily loaded machine may expose the timing race:

for i in $(seq 1 20); do
  dotnet test src/Workspaces/MSBuild/Test/Microsoft.CodeAnalysis.Workspaces.MSBuild.UnitTests.csproj \
    -c Debug -f net10.0 --no-build --no-restore \
    --filter 'FullyQualifiedName~NetCoreTests.TestOpenProject_FileBasedApp_RefDirective_Self' \
    --verbosity quiet || break
done

Reproduction limitation: this race is timing-dependent, not deterministic. The available machine is Windows, not Linux. On the unpatched code it passed 20 sequential runs and 40 runs with four parallel testhosts; the original disposal test also passed 80/80 pre-fix runs. The failure itself was observed in the loaded Linux CI run. Thus this command represents the failing test and may reproduce under suitable Linux load, but we cannot claim it reliably fails locally.

Fix

Manager disposal records its intent before waiting on the process-map gate. Disconnect callbacks check that intent both when claiming a process and just before logging, including the case where a callback already claimed it. Detaching handlers during disposal avoids subsequent callbacks; unexpected disconnects still report failure. The revised regression test uses a barrier after the disconnect callback claims a running BuildHost but before logging; it starts manager disposal and then releases and awaits the callback. This controlled test failed against the previous callback behavior and passes with the new guard. A companion test verifies genuine unexpected disconnects still report failure.

Validation

  • Analyzer-enabled Debug/net10.0 build and focused BuildHostProcessManagerTests plus the original file-based-app test: 27 passed. Debug/net472 focused tests: 2 passed.
  • The controlled regression test failed pre-guard and passed post-guard.
  • The first review revision's CI failed because the new test used a non-generic TaskCompletionSource, unavailable on net472. Commit 056fdef2460 corrected the cross-target test; both net472 and net10.0 builds and tests passed locally.
  • CI on the current head 056fdef2460 passed, including Linux Debug tests, Windows Helix, analyzer/correctness, and integration checks.

Detach disconnect handlers before intentionally disposing managed BuildHost processes so graceful RPC shutdown cannot be mistaken for a process failure. Add regression coverage for disposal diagnostics.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 2 pipeline(s).
There may be pipelines that require an authorized user to comment /azp run to run.

@jasonmalinowski

Copy link
Copy Markdown
Member

I'm wondering if this is related to #84412 and may actually explain a more root cause that we were looking for there.

@mwiemer-microsoft
mwiemer-microsoft marked this pull request as ready for review September 30, 2026 23:56
@mwiemer-microsoft
mwiemer-microsoft requested a review from a team as a code owner September 30, 2026 23:56
Copilot AI balanced review requested due to automatic review settings September 30, 2026 23:56
@mwiemer-microsoft
mwiemer-microsoft enabled auto-merge (squash) September 30, 2026 23:57

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The implementation does not address the demonstrated interleaving, and its test cannot detect the regression.

Review effort: Balanced
Findings: 1 Medium severity · 1 Low severity

Open (2)
What changed in this PR

Attempts to suppress false BuildHost shutdown diagnostics by detaching disconnect handlers.

Changes:

  • Detaches handlers before disposing BuildHost processes.
  • Adds a shutdown diagnostic test.

Holistic Assessment: The reported issue is valid, but the new unsubscription does not alter the relevant locking outcomes, and the test passes without the fix.

File Description
BuildHostProcessManager.cs Detaches handlers during disposal.
BuildHostProcessManagerTests.cs Tests disposal diagnostics.

Comment on lines +243 to +244
foreach (var process in processesToDispose)
process.Disconnected -= BuildHostProcess_Disconnected;

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI-generated:

Agreed—the original unsubscription does not cover a callback that has already claimed the process. Commit 62907f419e2 now records disposal intent before waiting on the gate and checks it both when the callback takes the gate and immediately before logging after it has released the gate. A callback that already claimed the process remains responsible for disposing it, but does not report an intentional shutdown as a failure. The controlled test failed on the old callback behavior and passes with this guard. I’ve left the thread open for review.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

resolving ai-to-ai thread

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looking at this carefully, it seems that the original Copilot review in this thread was correct: because _processes was cleared in the Dispose, this race it was imagining cannot occur. The test the coding agent then wrote doesn't really make sense: it's saying if we dispose the build host process manager while we're processing a disconnect (but one that was started after the disconnect was observed) shouldn't trigger any diagnostics. But that's wrong, because we only got into this situation because indeed the process went away first!

Whatever the bug is, this isn't the right fix. I'd propose closing this PR.

await manager.GetBuildHostAsync(BuildHostProcessKind.NetCore, CancellationToken.None);
await manager.DisposeAsync();

await Task.Delay(500);

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI-generated:

Agreed. Commit 62907f419e2 replaces the delay with a TestAccessor barrier after the disconnect callback claims the live BuildHost and before it logs. The test starts manager disposal at that barrier, then releases and awaits the callback. It failed against the pre-fix callback logic and passed with the disposal guard; a separate test confirms an unexpected disconnect still reports a failure. The existing Windows test suite and original file-based-app test pass. I’ve left the thread open for review.

@mwiemer-microsoft

Copy link
Copy Markdown
Member Author

AI-generated:

I looked at #84412 and the underlying #84411 scenario. They involve related disconnect ordering, but I cannot establish that this is the root cause of that crash. This PR addresses a narrower case: during intentional BuildHostProcessManager disposal, the RPC pipe can disconnect before the child has exited, and our disconnect callback can emit a false nonresponsive-process diagnostic. #84412 addresses the BuildHost server side: writing a response after its client disconnects can throw an unhandled IOException, regardless of why the pipe closed.

The #84411 report describes sustained LSP activity, not a known manager teardown; requests on an already-cached BuildHost do not themselves trigger the disposal path. Concurrent workspace teardown could connect the two, but we have no evidence it occurred. I think the fixes are complementary. I would retain the server-side protection in #84412 and this client-side intentional-shutdown fix. Logs showing manager teardown during the active session would be needed to confirm a shared trigger.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: e22abb5d-c5f6-4804-852a-4608fb19758f
Copilot AI balanced review requested due to automatic review settings October 1, 2026 01:35

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The disposal guard suppresses genuine unexpected disconnects that race with later disposal.

Review effort: Balanced
Findings: 1 Medium severity · 1 Low severity

Open (2)

Use generic TaskCompletionSource instances so the new shutdown-race test compiles on both net10.0 and .NET Framework.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 1, 2026 02:18

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The race is correctly guarded with deterministic positive and negative coverage.

Review effort: Balanced
Findings: 1 Medium severity · 1 Low severity

Open (2)

@mwiemer-microsoft

Copy link
Copy Markdown
Member Author

@RikkiGibson requesting review :)

@RikkiGibson

Copy link
Copy Markdown
Member

I didn't debug the tests, but read the description and tried to interpret how the problem was originally occurring:

  • BuildHostProcessManager.DisposeAsync() is called.
  • This calls BuildHostProcess.DisposeAsync().
  • This calls _rpcClient.Shutdown().
  • This calls event handler BuildHostProcess_Disconnected.
  • We extract processToDispose from _processes.
  • We call processToDispose.LogProcessFailure(). Process has not exited yet so we log an error.

Is that an accurate sequence of events?

One thing doesn't track for me with this--we shouldn't have been able to extract a processToDispose in this case. The list should have already been cleared.

@mwiemer-microsoft
mwiemer-microsoft marked this pull request as draft October 2, 2026 18:33
auto-merge was automatically disabled October 2, 2026 18:33

Pull request was converted to draft

@mwiemer-microsoft

Copy link
Copy Markdown
Member Author

(marking draft: learning about BuildHost as someone new to the domain, want to get this right instead of fixing symptoms)

@jasonmalinowski

Copy link
Copy Markdown
Member

I didn't debug the tests, but read the description and tried to interpret how the problem was originally occurring:

  • BuildHostProcessManager.DisposeAsync() is called.
  • This calls BuildHostProcess.DisposeAsync().
  • This calls _rpcClient.Shutdown().
  • This calls event handler BuildHostProcess_Disconnected.
  • We extract processToDispose from _processes.
  • We call processToDispose.LogProcessFailure(). Process has not exited yet so we log an error.

Is that an accurate sequence of events?

One thing doesn't track for me with this--we shouldn't have been able to extract a processToDispose in this case. The list should have already been cleared.

@RikkiGibson Your analysis was correct here (or at least matches my own analysis!) and that matched the Copilot code review analysis too. @mwiemer-microsoft's agent then doubled down on the analysis, by introducing a new test that triggered Dispose() on the BuildHostProcessManager while it's in the middle of processing the early disposal, claiming "that shouldn't log a diagnostic. But of course that should -- the process still failed first!

@mwiemer-microsoft I'd recommend closing this PR since it seems your agent is going down a bad path (or at least, it's missing the real root cause and then making things more confusing).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants