wait for service registration cleanup until allocs marked lost #26424

tgross · 2025-08-04T19:39:39Z

When a node misses a heartbeat and is marked down, Nomad deletes service registration instances for that node. But if the node then successfully heartbeats before its allocations are marked lost, the services are never restored. The node is unaware that it has missed a heartbeat and there's no anti-entropy on the node in any case.

We already delete services when the plan applier marks allocations as stopped, so deleting the services when the node goes down is only an optimization to more quickly divert service traffic. But because the state after a plan apply is the "canonical" view of allocation health, this breaks correctness.

Remove the code path that deletes services from nodes when nodes go down. Retain the state store code that deletes services when allocs are marked terminal by the plan applier. Also add a path in the state store to delete services when allocs are marked terminal by the client. This gets back some of the optimization but avoids the correctness bug because marking the allocation client-terminal is a one way operation.

Fixes: #16983
Ref: https://hashicorp.atlassian.net/browse/NMD-516

Contributor Checklist

Changelog Entry If this PR changes user-facing behavior, please generate and add a
changelog entry using the make cl command.
Testing Please add tests to cover any new functionality or to demonstrate bug fixes and
ensure regressions will be caught.
- See also my testing notes in Nomad Service Discovery drops services #16983 (comment)
Documentation If the change impacts user-facing functionality such as the CLI, API, UI,
and job configuration, please update the Nomad website documentation to reflect this. Refer to
the website README for docs guidelines. Please also consider whether the
change requires notes within the upgrade guide.

Reviewer Checklist

Backport Labels Please add the correct backport labels as described by the internal
backporting document.
Commit Type Ensure the correct merge method is selected which should be "squash and merge"
in the majority of situations. The main exceptions are long-lived feature branches or merges where
history should be preserved.
Enterprise PRs If this is an enterprise only PR, please add any required changelog entry
within the public repository.

jrasell

LGTM!

When a node misses a heartbeat and is marked down, Nomad deletes service registration instances for that node. But if the node then successfully heartbeats before its allocations are marked lost, the services are never restored. The node is unaware that it has missed a heartbeat and there's no anti-entropy on the node in any case. We already delete services when the plan applier marks allocations as stopped, so deleting the services when the node goes down is only an optimization to more quickly divert service traffic. But because the state after a plan apply is the "canonical" view of allocation health, this breaks correctness. Remove the code path that deletes services from nodes when nodes go down. Retain the state store code that deletes services when allocs are marked terminal by the plan applier. Also add a path in the state store to delete services when allocs are marked terminal by the client. This gets back some of the optimization but avoids the correctness bug because marking the allocation client-terminal is a one way operation. Fixes: #16983

…d lost (#26424) (#26448) When a node misses a heartbeat and is marked down, Nomad deletes service registration instances for that node. But if the node then successfully heartbeats before its allocations are marked lost, the services are never restored. The node is unaware that it has missed a heartbeat and there's no anti-entropy on the node in any case. We already delete services when the plan applier marks allocations as stopped, so deleting the services when the node goes down is only an optimization to more quickly divert service traffic. But because the state after a plan apply is the "canonical" view of allocation health, this breaks correctness. Remove the code path that deletes services from nodes when nodes go down. Retain the state store code that deletes services when allocs are marked terminal by the plan applier. Also add a path in the state store to delete services when allocs are marked terminal by the client. This gets back some of the optimization but avoids the correctness bug because marking the allocation client-terminal is a one way operation. Fixes: #16983 Co-authored-by: Tim Gross <[email protected]>

vercel bot deployed to Preview – nomad-ui August 4, 2025 19:40 View deployment

tgross mentioned this pull request Aug 4, 2025

Nomad Service Discovery drops services #16983

Closed

tgross added backport/ent/1.8.x+ent Changes are backported to 1.8.x+ent backport/ent/1.9.x+ent Changes are backported to 1.9.x+ent backport/1.10.x backport to 1.10.x release line type/bug theme/service-discovery theme/service-discovery/nomad labels Aug 4, 2025

tgross force-pushed the b-service-deregistration-on-node-down branch from 5f7a8d8 to b42d77a Compare August 4, 2025 19:46

vercel bot deployed to Preview – nomad-ui August 4, 2025 19:47 View deployment

tgross added this to the 1.11.0 milestone Aug 4, 2025

tgross marked this pull request as ready for review August 4, 2025 20:05

tgross requested review from a team as code owners August 4, 2025 20:05

tgross requested review from jrasell and pkazmierczak August 4, 2025 20:05

jrasell previously approved these changes Aug 5, 2025

View reviewed changes

tgross dismissed jrasell’s stale review via 75071f7 August 5, 2025 13:58

tgross force-pushed the b-service-deregistration-on-node-down branch from b42d77a to 75071f7 Compare August 5, 2025 13:58

vercel bot deployed to Preview – nomad-ui August 5, 2025 13:59 View deployment

tgross requested a review from jrasell August 5, 2025 14:09

jrasell approved these changes Aug 6, 2025

View reviewed changes

tgross merged commit 6563d0e into main Aug 6, 2025
37 checks passed

tgross deleted the b-service-deregistration-on-node-down branch August 6, 2025 17:40

hc-github-team-nomad-core mentioned this pull request Aug 6, 2025

Backport of wait for service registration cleanup until allocs marked lost into release/1.10.x #26448

Merged

6 tasks

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

wait for service registration cleanup until allocs marked lost #26424

wait for service registration cleanup until allocs marked lost #26424

Uh oh!

tgross commented Aug 4, 2025 •

edited

Loading

Uh oh!

jrasell left a comment

Uh oh!

Uh oh!

Uh oh!

wait for service registration cleanup until allocs marked lost #26424

wait for service registration cleanup until allocs marked lost #26424

Uh oh!

Conversation

tgross commented Aug 4, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Contributor Checklist

Reviewer Checklist

Uh oh!

jrasell left a comment

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

tgross commented Aug 4, 2025 •

edited

Loading