Skip to content

cuda.core: add a bounded asyncio wait for NVML device and system events - #2924

Open
0z5a wants to merge 13 commits into
NVIDIA:mainfrom
0z5a:cp1458-async-nvml-events
Open

0z5a wants to merge 13 commits into
NVIDIA:mainfrom
0z5a:cp1458-async-nvml-events

Conversation

@0z5a

@0z5a 0z5a commented Sep 19, 2026 •

Copy link
Copy Markdown

Incremental review: Files changed against main.

Scope

Related to #1458. Adds DeviceEvents.wait_async(timeout_ms=0) and RegisteredSystemEvents.wait_async(timeout_ms=0, buffer_size=1) so an asyncio control loop can await NVML events while continuing other work.

Async iterators, a public executor setting, background listener management and system hot-plug handling remain outside this proposal.

Behavior

  • Waits on one event set are serialized by a thread lock across synchronous callers, asyncio tasks and separate event loops. A synchronous caller can block behind another consumer; asynchronous callers yield while waiting for the lock, and that time counts against their timeout budget.
  • Async native waits use positive timeout slices of at most 100 ms against a monotonic deadline. timeout_ms=0 means indefinite waiting for the async APIs and skip waiting for synchronous wait().
  • Cancellation drains the outstanding slice before propagating CancelledError, including repeated cancellation. A native error encountered during that drain is saved for the next consumer.
  • A consumed event or system batch is retained for the next wait after cancellation. Both sync and async system waits deliver parked batches in buffer_size pieces, and failed conversion preserves the payload for a retry. Invalid buffer sizes are rejected before consuming an event.
  • Native waits run on private, lazily created, reusable non-daemon workers, with at most eight workers. Importing cuda.core.system starts no threads. Closing a coroutine or its loop keeps the event-set lock held until its outstanding native call finishes and saves its outcome.

Review follow-up and current validation

Review fixes in 30feadc address the batch-loss, cancellation-error and unlocked-lease findings, replace concurrent-wait exceptions with serialization, correct the fake early-timeout behavior, and add closed-loop and conversion-failure controls. Release notes cover the async APIs, wait locking overhead and invalid-buffer behavior; generated stubs match the updated docstrings.

The final branch at 5199f5e incorporates main at 608a37ddc71a37235f2476e57e15a2b3cd7259a9, retaining both sets of release notes. CPU controls, mypy and full Cython code generation were rechecked after this integration. GitHub confirms there are no merge conflicts.

  • CPU controls: 43 passed, 4 native GPU cases deselected. The same controls on the previous implementation reproduce 22 failures and 21 passes.
  • The CPU controls compile the affected Cython event classes and the NVML batch class/dtype excerpt with Cython 3.2.9. NVML entry points, initialization and binding enums are modeled. They exercise one-, two- and three-event cancellation handoffs through both public wait paths, buffer limits, conversion retry, sync/async/thread/loop serialization, repeated cancellation, asyncio.timeout, TaskGroup, early timeouts and closed-loop cleanup.
  • Ruff 0.15.9 checks/format, mypy 1.20.0 across 81 source files, cython-lint 0.19.0, full _device.pyx and _system_events.pyx Cython code generation, stub generation, SPDX, mempool hygiene, release-note presence and whitespace checks passed locally.
  • Native NVML/GPU tests and GPU benchmarks were not rerun for this follow-up on the macOS host. The GPU measurements below belong to earlier revisions.

Historical GPU measurements

The following measurements were previously reported for the pre-review implementation. They provide context for the proposal; the current review fixes need a fresh native run before these measurements can qualify this revision.

Measured on one host, 2 × NVIDIA L20 (SM89, 46 GB), driver 595.91.07 / NVML 13.595.91.07, cuda-bindings 12.9.8, Python 3.10.12, CUDA toolkit 12.8.61.

  • Base SHA: 8b6e9f52db1074a580fc1678c71e7c75c42db2ea
  • Measured revision: d6bff844d, before the review follow-up below. These tables were measured after wait_async became a real async def; they have not been re-measured for the current locking and error-handoff changes.
  • Loaded extensions and their SHA256: evidence/env-patch.json; independent clean rebuild: evidence/env-clean.json
  • Deterministic contract tests: 27 passed — cuda_core/tests/system/test_system_events_async.py drives every race with a fake native wait (slice budget, early timeout, single consumer, cancel during the in-flight slice, repeated cancel, hand-off of an event consumed by a cancelled slice, batch remainder, exactly-once free of a failed registration, saturated scheduling, reuse across event loops, argument bounds) and exercises real event sets on both L20s (timeout budget, cancel-to-drain, two independent event sets, sync/async lease conflict)
  • Regression: cuda_core/tests/system + test_stream.py: 150 passed, 131 skipped, 8 xfailed, 984 subtests passed
  • Second host (8 x RTX 5090, Python 3.12.13, cuda-bindings 13.4.2, CUDA 13.4 headers), same commit and tree: 27 passed, and cuda_core/tests in full — 4317 passed, 265 skipped, 17 xfailed, 872 subtests passed, 0 failed (4463 collected). 8-GPU fan-in with quiet event sets: stop latency +80.3 % (2 GPUs), +86.8 % (4), +88.6 % (8), with 1+N worker threads and the bound reached exactly at 8; the blocking baseline grows as ~(N-1) x timeout because this host's driver serialises concurrent blocking waits. Single-arm replication of the headline rows: heartbeat p50 300.90 ms -> 1.218 ms, stop p50 1701.7 ms -> 101.0 ms, exit 10015.8 ms -> 164.3 ms (details in the comment below)
  • Independent clean rebuild of this same commit (fresh worktree, fresh venv, pip install --no-build-isolation -e ., recorded in evidence/env-clean.json with the SHA256 of all 47 loaded extensions): 27 passed, and the paired results below reproduce

Every row is a paired four-arm quartet (A→P→P→A / P→A→A→P, a fresh process per arm), 3–5 independent quartets per row, session-level bootstrap intervals, with an A/A placebo measured at −0.0 % / −0.1 %.

Metric Baseline wait_async Improvement 95 % interval Quartets
Control-loop heartbeat p50 — 1 ms heartbeat while a control loop cycles 100 ms waits 300.9 ms 1.112 ms +99.6 % [+99.6 %, +99.6 %] 4
Control-loop heartbeat p99, same run 301.0 ms 1.255 ms +99.6 % [+99.6 %, +99.6 %] 4
Cycle throughput in that same run 9.971 /s 9.95 /s −0.2 % [−0.2 %, −0.2 %] 4
Stop latency p50 — sync wait(timeout_ms=1000) baseline 1702 ms 150.6 ms +91.6 % [+89.7 %, +93.2 %] 5
Stop latency p95 — sync wait(timeout_ms=1000) baseline 2173 ms 200.4 ms +89.7 % [+88.2 %, +92.1 %] 5
Stop latency p50 — sync wait(timeout_ms=100) baseline 200.2 ms 140.7 ms +33.7 % [+12.8 %, +49.5 %] 5
Stop latency under a real CLOCK/PSTATE load (2 GPUs) 1288 ms 374.7 ms +79.1 % [−38.9 %, +92.8 %] 4
Process exit with a wait in flight (timeout_ms=0, 10 s watchdog) 10020 ms (killed) 113.7 ms +98.9 % [+98.9 %, +98.9 %] 5
Process exit with a wait in flight (timeout_ms=1000) 10010 ms (killed) 118.7 ms +98.8 % [+98.7 %, +98.9 %] 5
Idle CPU over 6 s with timeout_ms=0 (the documented "wait indefinitely") 8873 ms 6014 ms +32.2 % [+32.1 %, +32.3 %] 4
Idle CPU over 6 s with timeout_ms=1000 6221 ms 6025 ms +3.0 % [−0.3 %, +6.1 %] 5
1 ms idle-churn throughput — blocking baseline 913.1 /s 818.5 /s −10.3 % [−12.0 %, −7.6 %] 5
1 ms idle-churn throughput — hand-rolled non-blocking baseline 849.0 /s 821.4 /s −3.3 % [−4.2 %, −2.5 %] 5
100 ms cycle throughput — hand-rolled non-blocking baseline 9.959 /s 9.947 /s −0.1 % [−0.1 %, −0.2 %] 5
Heartbeat p50 with the hand-rolled non-blocking baseline 1.114 ms 1.110 ms +0.3 % [+0.1 %, +0.7 %] 3

Independent clean rebuild at the same revision (4 fresh quartets per configuration): stop latency p50 +93.5 %, p95 +88.2 %, real-event stop +88.0 %, exit +98.8 %, 1 ms churn −11.8 %, idle CPU +4.3 %, real-event count +0.3 %. Repeating the real-event count over runs gives −23.3 % and +0.3 % with intervals that cross zero: it is noise around parity, and no run lost an event.

Notes that keep the numbers honest:

  • The heartbeat row is the point of the API. A blocking wait() inside an asyncio control loop starves that loop: with the sync API the same 1 ms heartbeat degrades to ~301 ms between beats, because it only gets to run between blocking waits. wait_async keeps it at 1.11 ms (p99 1.27 ms) with the same cycle throughput (−0.2 %), and a hand-rolled worker-thread + queue wrapper around the existing sync API reaches the same cadence (−0.1 %, i.e. parity), which is what makes this a fair like-for-like comparison.
  • The 1 ms idle-churn row is the only negative one, and it is the price of not blocking. A blocking wait(1 ms) occupies the calling thread for the whole wait; any non-blocking wait pays a wakeup plus an event-loop round trip (a hand-rolled worker-thread + queue wrapper measures −8 % against the same blocking baseline before this API exists, and bare asyncio.sleep(1 ms) on this host costs 1121 µs). This API lands at −3.3 % against that hand-rolled non-blocking wrapper — the lease/deadline/pending/drain machinery plus the async def frame — and −0.1 % at the 100 ms slice scale the design is built around. Earlier revisions of this branch measured −17.8 % here; the commits contain the fixes for that.
  • timeout_ms=0 does not mean "wait indefinitely" on this stack, contrary to the docstring at that time and the NVML C documentation: it returns immediately with TimeoutError (295k calls/s busy loop, 2 s of CPU for 2 s of wall). The async path therefore never passes 0 to the driver — it slices a positive timeout and implements the unbounded budget itself. The synchronous docstrings are corrected in this review follow-up to say that 0 skips waiting.

Limits

  • Earlier GPU runs used two L20s and, separately, eight RTX 5090s. The current review follow-up has been checked locally with modeled NVML entry points; native GPU tests and the performance rows still need to be rerun for this revision.
  • Real XID / ECC / GPU-lost events were not injected — driver unbind, GPU reset and fault injection are not acceptable on a shared host — so those paths are covered by the fake backend plus the native error-propagation test.
  • Real system-event hot-plug (bind/unbind) was not constructed; batch hand-over and remainder slicing are covered by a native test with a synthetic batch.
  • The host was shared (load average in the hundreds): absolute latencies include scheduler noise, so results are reported as paired differences with intervals rather than absolutes, and the same quartet schedule was re-run on an independent clean build.

@copy-pr-bot

copy-pr-bot Bot commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@github-actions github-actions Bot added the cuda.core Everything related to the cuda.core module label Sep 19, 2026
@0z5a

0z5a commented Sep 20, 2026

Copy link
Copy Markdown
Author

Marking this ready for review — it is the concrete proposal for #1458, and I am keeping it as a draft only until the API shape is settled.

Naming / shape questions (why it is still a draft)

  1. Is wait_async the name you want on DeviceEvents / RegisteredSystemEvents? The alternatives I weighed were await_event(...), wait_for_event(...), and overloading wait(..., async_=...). wait_async leaves the existing sync wait untouched and reads as its sibling, but I have no attachment to it — say the word and I will rename.
  2. Both methods return a coroutine rather than being declared async def; await events.wait_async(...) is unchanged. That saves one coroutine frame per wait, which measures ~14 µs/cycle at a 1 ms budget. If you would rather have real async def methods, I will switch it back and take that cost.
  3. For system events, buffer_size sits on wait_async next to the sync wait. Happy to move it if you prefer a different shape.

What is in the branch (904e0fea8, base 8b6e9f52d)

  • DeviceEvents.wait_async() and RegisteredSystemEvents.wait_async(), slicing the native wait at ≤100 ms against a monotonic deadline, a single-consumer lease per event set, drain-before-cancel-return, and a pending slot so an event consumed by a cancelled slice is handed to the next consumer instead of dropped.
  • 22 tests, most of them driving the races deterministically through a fake native wait, plus real event sets on two L20s. Importing cuda.core.system still starts no threads and needs no CUDA context.
  • Paired four-arm measurements on 2× L20 (fresh process per arm, A/A placebo at −0.0 %): stop latency 1701 → 142 ms, process exit with a wait in flight 10 s watchdog kill → 114 ms, idle CPU with timeout_ms=0 9244 → 6259 ms/6 s, and a 1 ms heartbeat on the same loop 301 → 1.11 ms while cycle throughput is unchanged. Full tables and raw logs are in the branch description.

One thing I found that may want its own decision: on this stack wait(timeout_ms=0) returns immediately with TimeoutError (295k calls/s busy loop), contrary to the docstring and the NVML C docs. The async path never passes 0 to the driver, so it is not a problem here, but the sync API's documented "wait indefinitely" does not hold — happy to file that separately or fold a fix in here, whichever you prefer.

Requesting review from @mdboom and @leofang. If it is easier to review in pieces, I can split device events and system events, or drop the system-event half entirely for a first PR.

@0z5a

0z5a commented Sep 20, 2026

Copy link
Copy Markdown
Author

Follow-ups pushed after the review request, all CI-visible:

  • 5709e8ba1 restores wait_async as a real async def (stubgen-pyx cannot annotate a coroutine returned from a plain def, and iscoroutinefunction() should report what callers expect). Cython's annotation typing turns the cdef-class return into a C type the coroutine wrapper cannot return, so the two annotations are marked as hints with @cython.annotation_typing(False); the generated stubs keep EventData / SystemEvents. Cost: one more coroutine frame per wait (~14 us on a 1 ms cycle).
  • Two annotation commits on top for mypy-cuda-core (pre-commit.ci is green now).
  • Also fixed what the first pre-commit run reported: import order, a test closure that a later del made look undefined, and a subprocess call that now says why its argv is fixed.

The measurements in the description are from 904e0fea8; the same quartet runs are being repeated on the final revision and I will replace the tables when they land.

@0z5a
0z5a force-pushed the cp1458-async-nvml-events branch 2 times, most recently from e5507ea to d6bff84 Compare September 20, 2026 05:46
@0z5a

0z5a commented Sep 20, 2026

Copy link
Copy Markdown
Author

Re-measurement and clean rebuild are done

Follow-up to the review request above: the tables are final now, and this branch is the tree they were measured on.

Branch state. The branch was rebuilt through the Git Data API (git transport to github.com is not available from the test host) and now carries the work-tree commit chain exactly, 9 commits, d6bff844d at the head — the same SHA the numbers below were taken from, so the tree, the tests and the measurements are the same object. 7 files, +1073/−2.

What changed since the earlier tables. wait_async went back to being a real async def so that iscoroutinefunction() and the generated stubs report what callers expect (one extra coroutine frame per wait, ~1 %, visible only in the 1 ms idle-churn row), plus four cases the contract implied but the suite did not exercise: exactly-once free of a failed registration, serialised slicing under a saturated worker bound, reuse of one event set across event loops (with a second loop rejected by the lease), and argument bounds on both the sync and async paths. 27 passed, regression cuda_core/tests/system + test_stream.py 150 passed / 131 skipped / 8 xfailed / 984 subtests, pre-commit.ci green.

Independent clean rebuild of this same commit (fresh worktree, fresh venv, pip install --no-build-isolation -e .; evidence/env-clean.json records the SHA and the SHA256 of all 47 loaded extensions), 4 fresh quartets per configuration:

Clean rebuild (4 quartets each) Baseline wait_async Improvement 95 % interval
Stop latency p50 (T=1000 ms) 1701 ms 113.4 ms +93.5 % [+92.3 %, +94.1 %]
Stop latency p95 (T=1000 ms) 1702 ms 200.4 ms +88.2 % [+88.2 %, +88.2 %]
Stop latency under a real CLOCK/PSTATE load (2 GPUs) 1061 ms 100.6 ms +88.0 % [+73.7 %, +92.9 %]
Process exit with a wait in flight (T=0, 10 s watchdog) 10010 ms (killed) 126.2 ms +98.8 % [+98.5 %, +98.9 %]
1 ms idle-churn throughput 924.8 /s 816 /s −11.8 % [−12.4 %, −11.1 %]
Idle CPU over 6 s (2 GPUs) 6723 ms 6427 ms +4.3 % [−7.5 %, +14.2 %]
Real events processed (6 s window) 11.75 12.5 +0.3 % [−9.4 %, +9.5 %]

The headline row is unchanged from the earlier run — a 1 ms heartbeat in the control loop goes from ~301 ms between beats with the blocking API to 1.11 ms with wait_async, at the same cycle throughput — and the clean rebuild reproduces it in direction and magnitude, including the one negative row (1 ms idle churn, decomposed in the report: ~8 % is the fixed price of not blocking the caller, measured against a hand-rolled worker-thread + queue wrapper, −0.1 % at the 100 ms slice scale the design targets). Real-event counts land within noise of parity in both sets, with no event lost and no cross-device mix-up.

Two things still need a maintainer's call, both asked above: the wait_async name (and whether these should return a coroutine rather than be async def), and what to do about timeout_ms=0 — measured here it returns immediately with TimeoutError (≈295k calls/s, i.e. a busy loop) rather than waiting indefinitely as the docstring and the NVML C documentation say. This PR documents that mismatch and deliberately does not change the synchronous semantics.

@0z5a

0z5a commented Sep 20, 2026

Copy link
Copy Markdown
Author

Second-host verification (8 × RTX 5090)

Same commit — this branch head d6bff844d, same tree 92a2bbe95f — rebuilt and re-run on a second machine with a different stack: 8 × RTX 5090 (SM120), driver 580.82.07, Python 3.12.13, cuda-bindings 13.4.2, CUDA 13.4 headers. Nothing in the change set differs; this is independent evidence that the behaviour and the numbers are not artefacts of the first host.

Full test suite. cuda_core/tests completes for the first time here: 4317 passed, 265 skipped, 17 xfailed, 872 subtests passed, 0 failed / 0 errors in 337 s (4463 tests collected). Previously only tests/system + test_stream.py had been run end to end.

Contract tests. The 27 async-event cases pass unchanged (27 passed in 3.54s), and the graph host-callback contract tests from #2925 pass on this host as well (3 passed).

Stop latency at scale. The first host only had two free GPUs, so the dispatcher's worker bound was covered by unit tests alone. On eight GPUs, quiet event sets, timeout_ms=1000, three repetitions per configuration:

Devices Blocking baseline wait_async Improvement Worker threads
2 1001.4–1002.5 ms 196.2–197.1 ms +80.3…80.4 % 3
4 3004.6–3004.9 ms 396.2–396.3 ms +86.8 % 5
8 7007.9–7009.0 ms 796.9–798.5 ms +88.6 % 9

Two things worth noting from this:

  • The blocking baseline grows as ≈ (N−1) × timeout on this host: the driver serialises concurrent blocking waits inside one process, so stopping N blocking waiters costs N−1 full timeouts. The async path is serialised by the same driver, but at the 100 ms slice granularity (≈ N × 100 ms), which is why the improvement grows with device count instead of staying flat.
  • threading.active_count() shows 1 + N worker threads at every configuration, i.e. all N slices are in flight simultaneously — at N = 8 the bound (_MAX_WORKERS = 8) is exactly reached, with no queueing and no errors. Each device's event counter stayed at 0, so no waiter received another device's event.

Caveat: the (N−1) × timeout serialisation is this host's driver behaviour and is not extrapolated to the first host, where only 2 GPUs were available.

Direction replication (single-arm, not paired statistics). Re-running the same harness on this host reproduces the headline rows:

Metric First host (paired, this PR's tables) Second host (single-arm check)
Heartbeat p50 in a 100 ms wait cycle 300.9 ms → 1.112 ms (+99.6 %) 300.90 ms → 1.218 ms (+99.6 %), 8 beats/7 missed → 1688 beats/1 missed
Cycle throughput 9.971 → 9.95 /s (−0.2 %) 9.968 → 9.912 /s (−0.6 %)
Stop latency p50 (T=1000) 1702 → 150.6 ms (+91.6 %) 1701.7 → 101.0 ms (+94.1 %)
Process exit with a wait in flight (T=0) 10020 ms (watchdog-killed) → 113.7 ms 10015.8 ms (watchdog-killed) → 164.3 ms

The baselines land within a few tenths of a percent of each other across two different GPUs, drivers, Python versions, bindings majors and CUDA header versions, which is what you would expect if these quantities are set by the native wait semantics rather than by host specifics.

@0z5a
0z5a marked this pull request as ready for review September 21, 2026 08:22
@mdboom
mdboom self-requested a review September 22, 2026 15:13
DeviceEvents.wait_async() and RegisteredSystemEvents.wait_async() wait for an
NVML event without blocking the event loop.  The native wait is issued in
slices of at most 100 ms against a monotonic deadline, which keeps
cancellation bounded: a cancelled task drains the in-flight slice instead of
parking a worker thread and the event-set handle for the caller's whole
timeout (indefinitely, when the timeout is 0).

The wait state enforces a single consumer per event set, so at most one native
wait is in flight, and it keeps an event that a cancelled or expired slice had
already consumed in a pending slot for the next consumer; a system-event batch
is handed over one buffer_size at a time so no event is dropped.  The worker
pool used for the blocking native call is private, bounded and created on
first use, so importing cuda.core.system still starts no threads and needs no
CUDA context.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
Deterministic tests drive the wait state with a fake native wait to pin the
slice budget, the anti-busy-wait behaviour of an early native timeout, single
consumer ownership, cancellation during the in-flight slice, repeated
cancellation, and the hand-off of an event that a cancelled slice consumed.
Native tests cover a real event set on the installed GPUs: timeout budget,
cancel-to-drain latency, two independent event sets, and the sync/async lease
conflict.  A subprocess test pins that importing cuda.core.system starts no
threads.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
The work item handed to the worker pool can be cancelled while it is still
queued, and its future then raises CancelledError from exception().  The drain
path has to read that as "this slice never borrowed the event set" instead of
letting a second exception type escape the cancellation path.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
A thread pool priced per request dominated the cost of a short wait: every
slice paid for a work item, an executor future and a chained asyncio future,
and a queued slice could wait behind a long one belonging to another event set.
Slices now go through a private queue served by workers that are created on
demand - one per in-flight slice, up to the bound - and reused across waits,
and the result is delivered straight to the awaiting future.

A slice that times out is now a value rather than an exception: only the
caller's own deadline raises, once, with the project's NVML timeout type.  The
result of a slice is kept in a slot on the wait state so that a cancellation
can still read it after the future carrying it was cancelled.

Measured on 2x L20 with a 1 ms timeout cycle: 926 -> 821 cycles/s against a
blocking baseline (was 761 before this change), at 176 us CPU per cycle.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
A drain that waits for the borrowed slice through the worker queue can queue up
behind another event set's slice, which doubled the measured stop latency.  The
drain now registers a waiter on the outcome slot and the worker that finishes
the slice completes it, so draining costs the rest of the slice and no more.
Workers also retire sooner once they go idle, which keeps a process that stops
waiting from paying for them at exit.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
The extension methods wrapped the wait state machine in one more coroutine,
which cost a frame on every wait: on a 1 ms timeout cycle it was ~14 us of the
~28 us the async path still spent over a hand-rolled non-blocking handoff.  The
methods now hand back the state machine's coroutine with the result converter,
so the state machine stays a single pure-Python frame; `await` is unchanged.

The conversion of a consumed payload moves into the state, which lets the
system-event path turn both a fresh batch and a parked one into SystemEvents
through one helper.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
stubgen-pyx has no way to annotate a coroutine that a plain def returns, and
pre-commit is the gate for review, so DeviceEvents.wait_async and
RegisteredSystemEvents.wait_async are real async def methods again: the stubs
keep their EventData / SystemEvents return types, and iscoroutinefunction()
reports what callers expect.  The cost is one more coroutine frame per wait.

Cython's annotation typing turns that cdef-class return annotation into a C type
the coroutine wrapper cannot return, so the two annotations are marked as hints
with @cython.annotation_typing(False), which keeps the generated stubs accurate.

Also fixes what pre-commit reported: import order, a bound method that kept the
test owner alive through a closure that a later del made look undefined, and a
subprocess call that now says why its argv is fixed.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
mypy asks for annotations on the queue and the worker list, and the request slot
is cleared to None so an idle worker does not pin an event set.  Types only: no
runtime behaviour changes.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
Four cases that the contract implies but the suite did not exercise yet:

* a failure after the native event set exists (device_register_events rejects)
  frees it exactly once instead of leaking it or freeing it twice;
* with the worker bound reached, slices are serialised, the queued slice does
  not reach the driver out of turn, and the queue wait stays inside the
  caller's budget;
* the same event set is reusable across consecutive event loops, while a second
  loop waiting on it concurrently is rejected by the lease instead of racing
  the driver for the same event set;
* type and range errors come from the annotated signatures (sync and async),
  including that an async argument error only surfaces when awaited.

The suite is 27 cases now.

Signed-off-by: 0z5a <Dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the cp1458-async-nvml-events branch from d6bff84 to ab5b1a7 Compare September 24, 2026 06:37

@mdboom mdboom left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this PR!

In addition to the code changes below, this will need release notes -- both for the (minor) breaking changes to wait and to mention the new feature of the async APIs.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a pre-existing documentation bug, but it's worth fixing it now.

Suggested change
The timeout in milliseconds. A default value of 0 means to skip waiting.

Comment thread cuda_core/cuda/core/system/_event.pxi Outdated

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is a pre-existing documentation bug, but it's worth fixing it now.

Suggested change
The timeout in milliseconds. A default value of 0 means to skip waiting.

except asyncio.CancelledError:
result = await _drain(outcome, loop)
if result is not None:
self.park(result)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From my agent:

cuda_core/cuda/core/system/_system_events.pyx (wait / _take_batch) — HIGH — correctness (event loss)

A cancelled wait_async parks the raw SystemEventData_v1 (park(result) at _async_events.py:304). Sync wait() converts parked data with _take_batch, which assumes a (batch, index) tuple. Only _batch_result handles both shapes, and only the async path uses it. Unpacking a raw SystemEventData_v1 raises ValueError for 1 event and 3 events, and mis-unpacks silently for exactly 2 (confirmed experimentally). take_pending() has already cleared the slot, so the batch is lost.

  • Failure scenario: cancel (or time out) a wait_async whose slice consumed a batch, then call wait(). It raises instead of returning the events, and the events are gone.
  • Fix: use _batch_result in both paths, or park a normalized (batch, 0) tuple.
  • Related: raw parked batches are returned whole by _batch_result, ignoring a smaller buffer_size. That contradicts the docstring ("one buffer_size slice at a time"). Any exception from convert, such as the buffer_size < 1 ValueError in _take_batch, also happens after the pending slot was cleared, so it loses the event too.
  • Test gap: test_system_batch_remainder_is_preserved only feeds tuples, so none of this is covered.

try:
await asyncio.shield(waiting)
except asyncio.CancelledError:
continue

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From my agent:

If the drained slice failed (for example GpuIsLostError), _drain raises that error from the except CancelledError handler. The task then ends with GpuIsLostError rather than CancelledError (reproduced). This breaks the usual contract that task.cancel() ends in CancelledError, and it confuses asyncio.timeout() and TaskGroup. Consider logging or parking the error and re-raising CancelledError.

Comment on lines +256 to +263
if self._busy:
raise RuntimeError("an event wait is already in flight for this event set")
self._busy = True
try:
pending, claimed = self._claim(convert)
return pending if claimed else native_wait(timeout_ms)
finally:
self._busy = False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

From my agent:

The lease check-and-set (if self._busy: ...; self._busy = True) is unlocked. Two threads calling wait(), or one calling wait() and one wait_async() from another loop, can both pass the check. test_event_set_is_reusable_across_event_loops only tests the deterministic ordering



def test_event_set_is_reusable_across_event_loops():
"""N22: serial reuse across loops is fine; a concurrent loop is rejected."""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"""N22: serial reuse across loops is fine; a concurrent loop is rejected."""
"""serial reuse across loops is fine; a concurrent loop is rejected."""



def test_parameter_bounds_match_the_signatures():
"""N23: the type and range errors come from the annotated signatures."""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
"""N23: the type and range errors come from the annotated signatures."""
"""the type and range errors come from the annotated signatures."""

with pytest.raises(system.TimeoutError):
asyncio.run(events.wait_async(timeout_ms=300))
elapsed = time.monotonic() - started
assert 0.3 <= elapsed < 1.0, elapsed

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
assert 0.3 <= elapsed < 1.0, elapsed
# Timings may be potentially flaky on loaded runners.
# Remove if we see flaky tests in CI.
assert 0.3 <= elapsed < 1.0, elapsed

elapsed = time.monotonic() - started

asyncio.run(main())
assert elapsed < 0.5, f"cancel returned only after {elapsed:.3f}s"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
assert elapsed < 0.5, f"cancel returned only after {elapsed:.3f}s"
# Timings may be potentially flaky on loaded runners.
# Remove if we see flaky tests in CI.
assert elapsed < 0.5, f"cancel returned only after {elapsed:.3f}s"

@@ -0,0 +1,315 @@
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
# SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.

@mdboom mdboom self-assigned this Oct 5, 2026
@mdboom mdboom added bug Something isn't working enhancement Any code-related improvements labels Oct 5, 2026
@mdboom mdboom added this to the cuda.core 1.3.0 milestone Oct 5, 2026
0z5a added 3 commits October 6, 2026 03:07
Serialize waits across threads and event loops, preserve batches and native errors across cancellation, and retain consumed payloads when conversion fails. Cover early timeouts and closed loops, correct synchronous timeout documentation, and document the async APIs and wait changes in the release notes.

Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@mdboom

mdboom commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/ok to test 5199f5e

@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/ok to test 5199f5e

@mdboom, there was an error processing your request: E2

See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/

@mdboom

mdboom commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/ok to test a0faf4a

@copy-pr-bot

copy-pr-bot Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/ok to test a0faf4a

@mdboom, there was an error processing your request: E2

See the following link for more information: https://docs.gha-runners.nvidia.com/cpr/e/2/

@coderabbitai

coderabbitai Bot commented Oct 6, 2026

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA/cuda-python/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: f05b23cf-6b2b-4942-99a3-61b91aafe339
📥 Commits

Reviewing files that changed from the base of the PR and between 4c893de and a0faf4a.

📒 Files selected for processing (8)
  • cuda_core/cuda/core/system/_async_events.py
  • cuda_core/cuda/core/system/_device.pyi
  • cuda_core/cuda/core/system/_device.pyx
  • cuda_core/cuda/core/system/_event.pxi
  • cuda_core/cuda/core/system/_system_events.pyi
  • cuda_core/cuda/core/system/_system_events.pyx
  • cuda_core/docs/source/release/1.3.0-notes.rst
  • cuda_core/tests/system/test_system_events_async.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review.


📝 Summary

Summary by CodeRabbit

  • New Features

    • Added asynchronous event waiting for device and registered system events, allowing waits without blocking the event loop.
    • Added configurable timeouts and buffered system-event batches. Async waits with a zero timeout wait indefinitely; synchronous device-event waits with a zero timeout skip waiting.
    • Synchronous and asynchronous waits on the same event set are serialized. Events and errors encountered during cancellation are preserved for the next wait.
  • Bug Fixes

    • Registered system-event waits now reject buffer sizes below one before consuming events.

Walkthrough

Adds asynchronous NVML event waits for device and system event sets. Waits on each event set are serialized. Bounded native wait slices support cancellation, and consumed events or errors can be handed to a later wait. System-event results are delivered in buffer-sized batches.

Changes

NVML event wait coordination

Layer / File(s) Summary
Wait coordination and cancellation
cuda_core/cuda/core/system/_async_events.py, cuda_core/tests/system/test_system_events_async.py
Adds a bounded worker dispatcher and EventSetWaiting coordinator. Waits use deadline-limited slices, serialize access to an event set, and preserve results or errors from abandoned waits. Tests cover timeouts, cancellation, worker limits, serialization, and lease handling.
Device event wait API
cuda_core/cuda/core/system/_device.pyi, cuda_core/cuda/core/system/_device.pyx, cuda_core/cuda/core/system/_event.pxi, cuda_core/tests/system/test_system_events_async.py
Routes synchronous DeviceEvents.wait through the coordinator and adds wait_async. Synchronous timeout_ms=0 skips waiting; asynchronous timeout_ms=0 waits indefinitely. The Cython import and typing import ordering also change.
System event batches and validation
cuda_core/cuda/core/system/_system_events.pyi, cuda_core/cuda/core/system/_system_events.pyx, cuda_core/tests/system/test_system_events_async.py, cuda_core/docs/source/release/1.3.0-notes.rst
Adds synchronous and asynchronous RegisteredSystemEvents waits. Rejects buffer_size < 1 and parks surplus events for later batches. Tests cover batch preservation, validation, and cancellation. Release notes describe the new APIs and wait behavior.

Suggested reviewers: andy-jost

Priority: ⬇️ Low

Change: Feature

Merge Risk: ⚪ Minimal · up to a0faf

No merge-blocking issue is established for the async NVML waits. Proceed with normal checks, including native GPU validation when available.

  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@mdboom

mdboom commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/ok to test a0faf4a

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Comment thread cuda_core/cuda/core/system/_device.pyi

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working cuda.core Everything related to the cuda.core module enhancement Any code-related improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants