Problem
A simple request for today's paid-account count took about 4m18s. Junior created one Hex thread, then used repeated model turns to poll for its result. The user expected Junior to pass the question to Hex and return the answer.
The investigation recovered these timings from Sentry spans:
- One
create_thread call: about 0.96s.
- Ten
get_thread calls: about 165.78s total. Most took about 18s.
- Thirteen main-model calls: about 51.66s total.
- Eleven nested model calls: about 34.61s total.
- No
continue_thread calls.
The reported token total was about 420k, mostly cached. The exact breakdown was not independently verified. Do not treat this as 420k uncached or generated tokens.
This was not repeated SQL exploration. Each poll returned control to Junior's model and tool-review path with accumulated context. Hex execution time may remain substantial, but polling should not add model calls or repeated conversation content.
A related concern is result size: Junior needs the answer, not the full Hex agent conversation. We have not yet measured how much of the observed token total came from Hex response content versus repeated Junior context.
Goals
- Expose one model-visible operation for a direct Hex question.
- Pass through the question with only the context needed to resolve it, including the date and timezone when relevant.
- Poll a fixed Hex thread ID in ordinary code, not through repeated model decisions.
- Return a compact final answer, metric definition and date window when available, warnings or errors, and a Hex source URL.
- Keep Hex thinking, tool history, and full transcripts out of Junior's model context.
- Preserve safety and action review. Do not fix this by globally disabling review.
- Set limits for wait time and expected transient retries. Support cancellation. Let unexpected failures reach the owning boundary.
- Reuse the existing thread after a timeout or resume. Do not create duplicate Hex work.
- Measure model calls, cached and uncached token use, context size, response bytes, total latency, and Hex execution time separately.
Contracts found in the investigation
Hex documents separate response and history endpoints:
POST /v1/threads starts work asynchronously and returns a thread ID and URL.
GET /v1/threads/{id} returns RUNNING, IDLE, or ERROR, with a response array and URL. Response content can include text and references to cells, projects, and tasks.
GET /v1/threads/{threadId}/messages retrieves conversation history. full=true includes thinking, tool calls, and results.
No blocking execution endpoint or completion webhook was found in the reviewed documentation. This is a documentation finding, not proof that Hex cannot provide one.
REST access must be checked before choosing an implementation. Availability can vary by organization and may require Hex support. MCP access does not prove REST access, and MCP OAuth credentials must not be assumed to work as a REST personal token.
The active Hex MCP exposes create_thread, get_thread, and continue_thread. Its get_thread tool has no result-only option. Its description directs the model to poll up to 30 times, while the Hex skill directs up to 10 polls with roughly 20-second waits. The skill also requests query/pattern/context inputs and full raw output. These instructions do not fit a simple pass-through question well.
Possible approaches
A. Preferred: a Hex-owned, result-focused operation
Implement the behavior in packages/junior-hex:
- Accept the question and necessary context.
- Create one thread and retain its ID.
- Poll that ID in code within a bounded wait budget.
- Use the REST response endpoint if access is available.
- Extract only the answer and useful references.
- On timeout, return a pending result with the existing thread ID and URL. Resume the same work rather than starting again.
Do not call continue_thread to poll or to repair output formatting. Keep Hex routes, response parsing, errors, and policy in the Hex package. Add only a small provider-neutral contract to core if one is needed.
Before implementation, verify REST access, representative response payloads, the completion rule, and safe handling of an uncertain create result. A timeout during creation must not cause a blind retry that starts a second thread.
B. Interim: wrap the existing MCP tools
If REST access is not available, a Hex-owned wrapper could call create/get internally and return one compact result to the model.
This removes repeated model turns, but parsing MCP-rendered output may be less stable than using a typed REST response. Confirm the output contract. The MCP client inspected during the investigation does not pass an AbortSignal through callTool, so cancellation needs work. Preserve operation review rather than exposing an unreviewed path.
C. Watches and asynchronous completion
Return a pending response, then wake the same Conversation when Hex finishes.
Watches can route a completion event to a Conversation, but no Hex completion event is currently enabled. Without a Hex callback, the Hex plugin would still need bounded background polling and would publish one completion event itself.
This adds durable work, event ingestion, deduplication, expiry, and failure handling. Prefer it if normal Hex execution exceeds a safe synchronous wait budget. Do not create a scheduled automation just to poll each query. Check with Hex for a supported completion callback before building background polling.
D. Longer term: precomputed results for common metrics
Frequently requested metrics could use a precomputed result or cache with a clear freshness limit. This is not the first fix for arbitrary questions. It also needs date, timezone, freshness, and permission rules.
Acceptance criteria
- A direct question creates at most one Hex thread, including timeout/resume handling.
- The model invokes one result-focused operation in the normal path. Increasing the number of internal status polls does not increase Junior model or review calls.
- The returned content includes the answer and source URL, not intermediate agent history. Large results have an explicit size limit and a clear indication when content is omitted.
- Running, completed, failed, canceled, and timed-out work have clear outcomes. Pending work keeps the original thread reference.
- Conflicting skill and MCP polling guidance no longer drives the model into a polling loop.
- A before/after run separates Hex execution time from Junior overhead and reports the token breakdown. Hex latency itself is not promised to disappear.
Extend the primary owning integration scenario through real Junior wiring for product behavior. Use an eval for whether Junior selects the operation and answers concisely. Use telemetry for diagnostics, not as a product behavior assertion.
Source references
Source was inspected at commit 03005313a033ccd05a1a06c6104df9cb3556da31; recheck before implementation:
No implementation is selected or shipped by this issue. The proposed first step is to verify REST access and result payloads, then choose the smallest Hex-owned operation that meets these goals.
via David Cramer.
--
View Junior Session [Sentry]
Problem
A simple request for today's paid-account count took about 4m18s. Junior created one Hex thread, then used repeated model turns to poll for its result. The user expected Junior to pass the question to Hex and return the answer.
The investigation recovered these timings from Sentry spans:
create_threadcall: about 0.96s.get_threadcalls: about 165.78s total. Most took about 18s.continue_threadcalls.The reported token total was about 420k, mostly cached. The exact breakdown was not independently verified. Do not treat this as 420k uncached or generated tokens.
This was not repeated SQL exploration. Each poll returned control to Junior's model and tool-review path with accumulated context. Hex execution time may remain substantial, but polling should not add model calls or repeated conversation content.
A related concern is result size: Junior needs the answer, not the full Hex agent conversation. We have not yet measured how much of the observed token total came from Hex response content versus repeated Junior context.
Goals
Contracts found in the investigation
Hex documents separate response and history endpoints:
POST /v1/threadsstarts work asynchronously and returns a thread ID and URL.GET /v1/threads/{id}returnsRUNNING,IDLE, orERROR, with aresponsearray and URL. Response content can include text and references to cells, projects, and tasks.GET /v1/threads/{threadId}/messagesretrieves conversation history.full=trueincludes thinking, tool calls, and results.No blocking execution endpoint or completion webhook was found in the reviewed documentation. This is a documentation finding, not proof that Hex cannot provide one.
REST access must be checked before choosing an implementation. Availability can vary by organization and may require Hex support. MCP access does not prove REST access, and MCP OAuth credentials must not be assumed to work as a REST personal token.
The active Hex MCP exposes
create_thread,get_thread, andcontinue_thread. Itsget_threadtool has no result-only option. Its description directs the model to poll up to 30 times, while the Hex skill directs up to 10 polls with roughly 20-second waits. The skill also requests query/pattern/context inputs and full raw output. These instructions do not fit a simple pass-through question well.Possible approaches
A. Preferred: a Hex-owned, result-focused operation
Implement the behavior in
packages/junior-hex:Do not call
continue_threadto poll or to repair output formatting. Keep Hex routes, response parsing, errors, and policy in the Hex package. Add only a small provider-neutral contract to core if one is needed.Before implementation, verify REST access, representative response payloads, the completion rule, and safe handling of an uncertain create result. A timeout during creation must not cause a blind retry that starts a second thread.
B. Interim: wrap the existing MCP tools
If REST access is not available, a Hex-owned wrapper could call create/get internally and return one compact result to the model.
This removes repeated model turns, but parsing MCP-rendered output may be less stable than using a typed REST response. Confirm the output contract. The MCP client inspected during the investigation does not pass an
AbortSignalthroughcallTool, so cancellation needs work. Preserve operation review rather than exposing an unreviewed path.C. Watches and asynchronous completion
Return a pending response, then wake the same Conversation when Hex finishes.
Watches can route a completion event to a Conversation, but no Hex completion event is currently enabled. Without a Hex callback, the Hex plugin would still need bounded background polling and would publish one completion event itself.
This adds durable work, event ingestion, deduplication, expiry, and failure handling. Prefer it if normal Hex execution exceeds a safe synchronous wait budget. Do not create a scheduled automation just to poll each query. Check with Hex for a supported completion callback before building background polling.
D. Longer term: precomputed results for common metrics
Frequently requested metrics could use a precomputed result or cache with a clear freshness limit. This is not the first fix for arbitrary questions. It also needs date, timezone, freshness, and permission rules.
Acceptance criteria
Extend the primary owning integration scenario through real Junior wiring for product behavior. Use an eval for whether Junior selects the operation and answers concisely. Use telemetry for diagnostics, not as a product behavior assertion.
Source references
Source was inspected at commit
03005313a033ccd05a1a06c6104df9cb3556da31; recheck before implementation:approvalMode: "auto"still uses review machinery.signalandtoolCallId; verify whether these meet cancellation and duplicate-work needs.No implementation is selected or shipped by this issue. The proposed first step is to verify REST access and result payloads, then choose the smallest Hex-owned operation that meets these goals.
via David Cramer.
--
View Junior Session [Sentry]