Skip to content

Expose stream retry budgets and transport timeouts as configuration #6700

Description

@7jrxt42BxFZo4iAnN4CX

Problem

The parameters that decide how long the runtime tolerates a flaky network are
compiled in as const, with no configuration surface. Operators on proxy'd or
otherwise unreliable networks cannot tune them, and the only recourse is
patching the binary.

Current state by category:

Configurable today

  • [retry] — enabled, max_retries, initial_delay, max_delay,
    exponential_base (crates/tui/src/config.rs:2327-2333)
  • [tui].stream_chunk_timeout_secs — per-chunk idle timeout

Env var only

  • CODEWHALE_STREAM_OPEN_TIMEOUT_SECS — response-header wait, default 45s,
    clamped 5..=300 (crates/tui/src/client/stream_entry.rs:22)
  • CODEWHALE_FORCE_HTTP1 — pin HTTP/1.1

Not configurable at all (compiled constants)

Constant Default File
MAX_STREAM_RETRIES 3 crates/tui/src/core/engine/streaming.rs:77
MAX_TRANSPARENT_STREAM_RETRIES 2 crates/tui/src/core/engine/streaming.rs:53
MAX_STREAM_ERRORS_BEFORE_FAIL 5 crates/tui/src/core/engine/streaming.rs:48
STREAM_MAX_DURATION_SECS 1800 crates/tui/src/core/engine/streaming.rs:43
connect_timeout 30s crates/tui/src/client.rs:2028
tcp_keepalive 30s crates/tui/src/client.rs:2029
http2_keep_alive_interval 15s crates/tui/src/client.rs:2030
http2_keep_alive_timeout 20s crates/tui/src/client.rs:2031
jitter / jitter_factor true / 0.1 crates/tui/src/llm_client/mod.rs:993, :996, :1021-1022

The gap is not uniform. [retry] exposes the HTTP-request retry schedule, but
the stream-level budgets that consume it are fixed, so raising
[retry].max_retries has no effect on a stream-open failure — the turn never
reaches that layer. Meanwhile the env var that does control the stream-open
window is invisible to codewhale config dump.

jitter and jitter_factor exist as public fields of RetryConfig
(crates/tui/src/llm_client/mod.rs:993-996) but are never read from config:
From<RetryPolicy> for RetryConfig (crates/tui/src/llm_client/mod.rs:1111)
transfers only the five policy fields and takes the rest from Default::default().

Proposed solution

Extend the existing tables rather than introduce a new one. Concretely:

[retry]
enabled = true
max_retries = 3
initial_delay = 1.0
max_delay = 60.0
exponential_base = 2.0
# new
respect_retry_after = true
jitter = true
jitter_factor = 0.1

[stream]
open_timeout_secs = 45      # replaces CODEWHALE_STREAM_OPEN_TIMEOUT_SECS
force_http1 = false         # replaces CODEWHALE_FORCE_HTTP1
max_resumes = 3             # MAX_STREAM_RETRIES
max_transparent_retries = 2 # MAX_TRANSPARENT_STREAM_RETRIES
max_stream_errors = 5       # MAX_STREAM_ERRORS_BEFORE_FAIL
max_duration_secs = 1800    # STREAM_MAX_DURATION_SECS

A [stream] table groups the stream-layer concerns that [tui] and [retry]
do not currently own, and keeps the transport knobs (which are network
topology, not TUI preferences) out of [tui].

Env vars would keep working as overrides so existing setups are unaffected.

Use case

Operating behind a corporate proxy, v2rayN, or a satellite/mobile link, where
the first streaming request occasionally hangs until the 45s header timeout.
Today the only options are: accept the failure, or edit and rebuild the
binary. With [stream].max_resumes and open_timeout_secs, the retry posture
becomes a config change.

A second case: a provider that takes 60-90s to emit response headers on a cold
start. open_timeout_secs alone (already available via env) fixes this, but is
undiscoverable — it does not appear in codewhale config dump or
docs/CONFIGURATION.md's [retry] section, only in the env-var appendix.

Alternatives considered

  • Env vars for everything. Rejected: the budgets are policy, not
    deployment topology, and the codebase already has a config path for the
    adjacent retry schedule. codewhale config dump would still not show them.
  • A single network table. Considered and rejected as too broad for
    this change; it would mix proxy/TLS concerns into the same edit.

Impact

Unlocks recovery from transient provider/proxy stalls without a rebuild. Also
makes the effective retry posture inspectable, which matters for diagnosing
"it retried for 15 minutes and then gave up" reports.

Additional context

Related: the turn-level resume budgets below this layer are never consulted for
stream-open failures — see the companion bug report on a turn failing with no
retry when the SSE request never receives response headers.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageNew external report awaiting maintainer triage; repro, logs and version output help

    Projects

    • Status
      Backlog

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions