Skip to content

bug: "Sent message larger than max" error is retryable #621

Description

@weres-sa

Bug Report

When sending a large payload (e.g. UPSERT with $data parameter exceeding gRPC message limit), the driver returns ResourceExhausted ("trying to send message larger than max"). The problem is that this error is currently treated as retryable. Retrying does not change the situation — the payload size is the same, so every retry fails again with the same error. From the caller's perspective the process effectively hangs (repeated retries until timeout or context cancel), instead of failing fast so the client can reduce batch size or handle the error.

A similar situation was addressed for BulkUpsert and Topic writer. For the Query service (DoTx / Exec with large params), this specific subtype of ResourceExhausted ("message larger than max") should be treated as non-retryable, so that the error is returned to the client immediately. Alternatively, the SDK could handle large messages (e.g. split or document the limit) like for BulkUpsert/Topic.

YDB Java SDK version:

Any version

Environment

Any environment

Current behavior:

Message retires according to retry configuration

Expected behavior:

  • Option A (recommended): Treat ResourceExhausted with description "message larger than max" (and similar parameter size limits) as non-retryable, so the error is returned immediately and the client can reduce batch size or react.
  • Option B: Or handle large messages in the SDK (e.g. split before send, or document the limit) like for BulkUpsert/Topic.

Steps to reproduce:

Create an upsert query with huge parameters size

Other information:

Similar issue for go SDK: ydb-platform/ydb-go-sdk#2024

Activity

robot-vibe-db commented on Apr 27, 2026

@robot-vibe-db

Analysis

Root Cause

The bug originates from the unconditional classification of CLIENT_RESOURCE_EXHAUSTED as retryable. There are three layers where this is hardcoded:

  1. StatusCode.java:67-74 — CLIENT_RESOURCE_EXHAUSTED is included in RETRYABLE_STATUSES EnumSet, making isRetryable() return true for all ResourceExhausted errors without distinction.

  2. YdbRetryConfig.java:42-44 — CLIENT_RESOURCE_EXHAUSTED is mapped to slow backoff retry policy. The getStatusRetryPolicy(Status status) method receives the full Status object (which includes issues/error messages), but the implementation only inspects status.getCode() and ignores the message content entirely.

  3. SessionRetryContext.java (both query/ and table/ modules) — canRetry() delegates to StatusCode.isRetryable(), and backoffTimeMillis() assigns CLIENT_RESOURCE_EXHAUSTED to slow backoff (lines 125-128 in both files). These contexts have their own retry loop independent of RetryConfig.

The error path is: gRPC returns RESOURCE_EXHAUSTED → GrpcStatuses.getStatusCode() maps it to CLIENT_RESOURCE_EXHAUSTED → retry logic treats it as retryable → the same oversized request is retried indefinitely until timeout.

Impact on Services

  • Query service (DoTx/Exec) and Table service — Both SessionRetryContext implementations retry CLIENT_RESOURCE_EXHAUSTED with slow backoff. With default maxRetries=10 and slow backoff (slot=500ms, ceiling=6), the caller waits through all retries before getting the error. This creates a perceived "hang" for large payloads.
  • Topic service — PR Added support of RetryConfig to topic writers #637 (merged 2026-04-22) added RetryConfig support to topic writers, and PR Added support of large buffer and message sizes #623 (merged 2026-04-08) added large buffer support. However, these changes handle topic-specific splitting/buffering — the core YdbRetryConfig still classifies CLIENT_RESOURCE_EXHAUSTED as retryable with slow backoff.
  • BulkUpsert — Goes through standard gRPC error handling without explicit message-size checking or splitting.

Severity Assessment: Medium-High

  • Functional impact: Users sending payloads exceeding the gRPC limit (~4 MiB default outbound) experience effective hangs instead of immediate error feedback. The operation will never succeed on retry since the payload size is deterministic.
  • User experience: The behavior is confusing — the caller has no indication that retries are futile, and must wait for timeout/cancellation.
  • Workaround exists: Users can set maxRetries(0) on SessionRetryContext.Builder or use RetryConfig.noRetries(), but this disables retries for all errors, not just this one.
  • Scope: Affects any service using SessionRetryContext or YdbRetryConfig (Query, Table, and potentially others).
  • No data loss risk: The operation is never partially executed — it fails at the gRPC transport level before reaching the server.

Proposed Solutions

Option A (Recommended): Inspect error description in GrpcStatuses and map to a non-retryable status

Modify GrpcStatuses.getStatusCode() to check the gRPC status description for the "message larger than max" pattern when the code is RESOURCE_EXHAUSTED, and map it to a non-retryable status (e.g., a new CLIENT_RESOURCE_EXHAUSTED_NON_RETRYABLE status code, or simply BAD_REQUEST):

// In GrpcStatuses.getStatusCode() or toStatus()
case RESOURCE_EXHAUSTED:
    if (status.getDescription() != null 
            && status.getDescription().contains("message larger than max")) {
        return StatusCode.BAD_REQUEST; // non-retryable
    }
    return StatusCode.CLIENT_RESOURCE_EXHAUSTED; // retryable (server overload)

This approach is clean because it distinguishes at the source: "message too large" (client-side, deterministic, non-retryable) vs. server-side resource exhaustion (transient, retryable).

Option B: Inspect error message in YdbRetryConfig.getStatusRetryPolicy()

Since getStatusRetryPolicy(Status status) already receives the full Status object (including Issue array with the error message), it can check the message content:

case CLIENT_RESOURCE_EXHAUSTED:
    if (hasMessageTooLargeIssue(status)) {
        return null; // non-retryable
    }
    return slow;

This would also need to be applied to both SessionRetryContext implementations in query/ and table/ modules, which have independent retry logic that doesn't use RetryConfig.

Option C: Harmonize retry logic across all services

The SessionRetryContext classes (in both query/ and table/ modules) implement their own retry logic independent of RetryConfig/YdbRetryConfig. Refactoring them to use RetryConfig would centralize the retry policy and make fixes like this one effective across all services. This is a larger change but would prevent similar issues in the future.

Note on Branch Scope

Only the master branch exists in the repository. All analysis applies to the current master branch.

Related Issues

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions