Repository navigation
bug: "Sent message larger than max" error is retryable #621
Description
Activity
Analysis
Root Cause
The bug originates from the unconditional classification of CLIENT_RESOURCE_EXHAUSTED as retryable. There are three layers where this is hardcoded:
-
StatusCode.java:67-74—CLIENT_RESOURCE_EXHAUSTEDis included inRETRYABLE_STATUSESEnumSet, makingisRetryable()returntruefor all ResourceExhausted errors without distinction. -
YdbRetryConfig.java:42-44—CLIENT_RESOURCE_EXHAUSTEDis mapped toslowbackoff retry policy. ThegetStatusRetryPolicy(Status status)method receives the fullStatusobject (which includes issues/error messages), but the implementation only inspectsstatus.getCode()and ignores the message content entirely. -
SessionRetryContext.java(bothquery/andtable/modules) —canRetry()delegates toStatusCode.isRetryable(), andbackoffTimeMillis()assignsCLIENT_RESOURCE_EXHAUSTEDto slow backoff (lines 125-128 in both files). These contexts have their own retry loop independent ofRetryConfig.
The error path is: gRPC returns RESOURCE_EXHAUSTED → GrpcStatuses.getStatusCode() maps it to CLIENT_RESOURCE_EXHAUSTED → retry logic treats it as retryable → the same oversized request is retried indefinitely until timeout.
Impact on Services
- Query service (DoTx/Exec) and Table service — Both
SessionRetryContextimplementations retryCLIENT_RESOURCE_EXHAUSTEDwith slow backoff. With defaultmaxRetries=10and slow backoff (slot=500ms, ceiling=6), the caller waits through all retries before getting the error. This creates a perceived "hang" for large payloads. - Topic service — PR Added support of RetryConfig to topic writers #637 (merged 2026-04-22) added
RetryConfigsupport to topic writers, and PR Added support of large buffer and message sizes #623 (merged 2026-04-08) added large buffer support. However, these changes handle topic-specific splitting/buffering — the coreYdbRetryConfigstill classifiesCLIENT_RESOURCE_EXHAUSTEDas retryable with slow backoff. - BulkUpsert — Goes through standard gRPC error handling without explicit message-size checking or splitting.
Severity Assessment: Medium-High
- Functional impact: Users sending payloads exceeding the gRPC limit (~4 MiB default outbound) experience effective hangs instead of immediate error feedback. The operation will never succeed on retry since the payload size is deterministic.
- User experience: The behavior is confusing — the caller has no indication that retries are futile, and must wait for timeout/cancellation.
- Workaround exists: Users can set
maxRetries(0)onSessionRetryContext.Builderor useRetryConfig.noRetries(), but this disables retries for all errors, not just this one. - Scope: Affects any service using
SessionRetryContextorYdbRetryConfig(Query, Table, and potentially others). - No data loss risk: The operation is never partially executed — it fails at the gRPC transport level before reaching the server.
Proposed Solutions
Option A (Recommended): Inspect error description in GrpcStatuses and map to a non-retryable status
Modify GrpcStatuses.getStatusCode() to check the gRPC status description for the "message larger than max" pattern when the code is RESOURCE_EXHAUSTED, and map it to a non-retryable status (e.g., a new CLIENT_RESOURCE_EXHAUSTED_NON_RETRYABLE status code, or simply BAD_REQUEST):
// In GrpcStatuses.getStatusCode() or toStatus()
case RESOURCE_EXHAUSTED:
if (status.getDescription() != null
&& status.getDescription().contains("message larger than max")) {
return StatusCode.BAD_REQUEST; // non-retryable
}
return StatusCode.CLIENT_RESOURCE_EXHAUSTED; // retryable (server overload)This approach is clean because it distinguishes at the source: "message too large" (client-side, deterministic, non-retryable) vs. server-side resource exhaustion (transient, retryable).
Option B: Inspect error message in YdbRetryConfig.getStatusRetryPolicy()
Since getStatusRetryPolicy(Status status) already receives the full Status object (including Issue array with the error message), it can check the message content:
case CLIENT_RESOURCE_EXHAUSTED:
if (hasMessageTooLargeIssue(status)) {
return null; // non-retryable
}
return slow;This would also need to be applied to both SessionRetryContext implementations in query/ and table/ modules, which have independent retry logic that doesn't use RetryConfig.
Option C: Harmonize retry logic across all services
The SessionRetryContext classes (in both query/ and table/ modules) implement their own retry logic independent of RetryConfig/YdbRetryConfig. Refactoring them to use RetryConfig would centralize the retry policy and make fixes like this one effective across all services. This is a larger change but would prevent similar issues in the future.
Note on Branch Scope
Only the master branch exists in the repository. All analysis applies to the current master branch.
Related Issues
- Go SDK equivalent: ydb-platform/ydb-go-sdk#2024 — same problem reported for the Go SDK
- Go SDK precedent: ydb-platform/ydb-go-sdk#1660 — similar issue was fixed for Topic writer in Go SDK
- Java SDK PR Added support of large buffer and message sizes #623 — addressed large buffer support for Topic writer but not the core retry classification
- Java SDK PR Added support of RetryConfig to topic writers #637 — added RetryConfig to topic writers but didn't change how CLIENT_RESOURCE_EXHAUSTED is classified
Bug Report
When sending a large payload (e.g. UPSERT with $data parameter exceeding gRPC message limit), the driver returns ResourceExhausted ("trying to send message larger than max"). The problem is that this error is currently treated as retryable. Retrying does not change the situation — the payload size is the same, so every retry fails again with the same error. From the caller's perspective the process effectively hangs (repeated retries until timeout or context cancel), instead of failing fast so the client can reduce batch size or handle the error.
A similar situation was addressed for BulkUpsert and Topic writer. For the Query service (DoTx / Exec with large params), this specific subtype of ResourceExhausted ("message larger than max") should be treated as non-retryable, so that the error is returned to the client immediately. Alternatively, the SDK could handle large messages (e.g. split or document the limit) like for BulkUpsert/Topic.
YDB Java SDK version:
Any version
Environment
Any environment
Current behavior:
Message retires according to retry configuration
Expected behavior:
Steps to reproduce:
Create an upsert query with huge parameters size
Other information:
Similar issue for go SDK: ydb-platform/ydb-go-sdk#2024