Problem
The parameters that decide how long the runtime tolerates a flaky network are
compiled in as const, with no configuration surface. Operators on proxy'd or
otherwise unreliable networks cannot tune them, and the only recourse is
patching the binary.
Current state by category:
Configurable today
[retry] — enabled, max_retries, initial_delay, max_delay,
exponential_base (crates/tui/src/config.rs:2327-2333)
[tui].stream_chunk_timeout_secs — per-chunk idle timeout
Env var only
CODEWHALE_STREAM_OPEN_TIMEOUT_SECS — response-header wait, default 45s,
clamped 5..=300 (crates/tui/src/client/stream_entry.rs:22)
CODEWHALE_FORCE_HTTP1 — pin HTTP/1.1
Not configurable at all (compiled constants)
| Constant |
Default |
File |
MAX_STREAM_RETRIES |
3 |
crates/tui/src/core/engine/streaming.rs:77 |
MAX_TRANSPARENT_STREAM_RETRIES |
2 |
crates/tui/src/core/engine/streaming.rs:53 |
MAX_STREAM_ERRORS_BEFORE_FAIL |
5 |
crates/tui/src/core/engine/streaming.rs:48 |
STREAM_MAX_DURATION_SECS |
1800 |
crates/tui/src/core/engine/streaming.rs:43 |
connect_timeout |
30s |
crates/tui/src/client.rs:2028 |
tcp_keepalive |
30s |
crates/tui/src/client.rs:2029 |
http2_keep_alive_interval |
15s |
crates/tui/src/client.rs:2030 |
http2_keep_alive_timeout |
20s |
crates/tui/src/client.rs:2031 |
jitter / jitter_factor |
true / 0.1 |
crates/tui/src/llm_client/mod.rs:993, :996, :1021-1022 |
The gap is not uniform. [retry] exposes the HTTP-request retry schedule, but
the stream-level budgets that consume it are fixed, so raising
[retry].max_retries has no effect on a stream-open failure — the turn never
reaches that layer. Meanwhile the env var that does control the stream-open
window is invisible to codewhale config dump.
jitter and jitter_factor exist as public fields of RetryConfig
(crates/tui/src/llm_client/mod.rs:993-996) but are never read from config:
From<RetryPolicy> for RetryConfig (crates/tui/src/llm_client/mod.rs:1111)
transfers only the five policy fields and takes the rest from Default::default().
Proposed solution
Extend the existing tables rather than introduce a new one. Concretely:
[retry]
enabled = true
max_retries = 3
initial_delay = 1.0
max_delay = 60.0
exponential_base = 2.0
# new
respect_retry_after = true
jitter = true
jitter_factor = 0.1
[stream]
open_timeout_secs = 45 # replaces CODEWHALE_STREAM_OPEN_TIMEOUT_SECS
force_http1 = false # replaces CODEWHALE_FORCE_HTTP1
max_resumes = 3 # MAX_STREAM_RETRIES
max_transparent_retries = 2 # MAX_TRANSPARENT_STREAM_RETRIES
max_stream_errors = 5 # MAX_STREAM_ERRORS_BEFORE_FAIL
max_duration_secs = 1800 # STREAM_MAX_DURATION_SECS
A [stream] table groups the stream-layer concerns that [tui] and [retry]
do not currently own, and keeps the transport knobs (which are network
topology, not TUI preferences) out of [tui].
Env vars would keep working as overrides so existing setups are unaffected.
Use case
Operating behind a corporate proxy, v2rayN, or a satellite/mobile link, where
the first streaming request occasionally hangs until the 45s header timeout.
Today the only options are: accept the failure, or edit and rebuild the
binary. With [stream].max_resumes and open_timeout_secs, the retry posture
becomes a config change.
A second case: a provider that takes 60-90s to emit response headers on a cold
start. open_timeout_secs alone (already available via env) fixes this, but is
undiscoverable — it does not appear in codewhale config dump or
docs/CONFIGURATION.md's [retry] section, only in the env-var appendix.
Alternatives considered
- Env vars for everything. Rejected: the budgets are policy, not
deployment topology, and the codebase already has a config path for the
adjacent retry schedule. codewhale config dump would still not show them.
- A single
network table. Considered and rejected as too broad for
this change; it would mix proxy/TLS concerns into the same edit.
Impact
Unlocks recovery from transient provider/proxy stalls without a rebuild. Also
makes the effective retry posture inspectable, which matters for diagnosing
"it retried for 15 minutes and then gave up" reports.
Additional context
Related: the turn-level resume budgets below this layer are never consulted for
stream-open failures — see the companion bug report on a turn failing with no
retry when the SSE request never receives response headers.
Problem
The parameters that decide how long the runtime tolerates a flaky network are
compiled in as
const, with no configuration surface. Operators on proxy'd orotherwise unreliable networks cannot tune them, and the only recourse is
patching the binary.
Current state by category:
Configurable today
[retry]—enabled,max_retries,initial_delay,max_delay,exponential_base(crates/tui/src/config.rs:2327-2333)[tui].stream_chunk_timeout_secs— per-chunk idle timeoutEnv var only
CODEWHALE_STREAM_OPEN_TIMEOUT_SECS— response-header wait, default 45s,clamped
5..=300(crates/tui/src/client/stream_entry.rs:22)CODEWHALE_FORCE_HTTP1— pin HTTP/1.1Not configurable at all (compiled constants)
MAX_STREAM_RETRIEScrates/tui/src/core/engine/streaming.rs:77MAX_TRANSPARENT_STREAM_RETRIEScrates/tui/src/core/engine/streaming.rs:53MAX_STREAM_ERRORS_BEFORE_FAILcrates/tui/src/core/engine/streaming.rs:48STREAM_MAX_DURATION_SECScrates/tui/src/core/engine/streaming.rs:43connect_timeoutcrates/tui/src/client.rs:2028tcp_keepalivecrates/tui/src/client.rs:2029http2_keep_alive_intervalcrates/tui/src/client.rs:2030http2_keep_alive_timeoutcrates/tui/src/client.rs:2031jitter/jitter_factortrue/0.1crates/tui/src/llm_client/mod.rs:993,:996,:1021-1022The gap is not uniform.
[retry]exposes the HTTP-request retry schedule, butthe stream-level budgets that consume it are fixed, so raising
[retry].max_retrieshas no effect on a stream-open failure — the turn neverreaches that layer. Meanwhile the env var that does control the stream-open
window is invisible to
codewhale config dump.jitterandjitter_factorexist as public fields ofRetryConfig(
crates/tui/src/llm_client/mod.rs:993-996) but are never read from config:From<RetryPolicy> for RetryConfig(crates/tui/src/llm_client/mod.rs:1111)transfers only the five policy fields and takes the rest from
Default::default().Proposed solution
Extend the existing tables rather than introduce a new one. Concretely:
A
[stream]table groups the stream-layer concerns that[tui]and[retry]do not currently own, and keeps the transport knobs (which are network
topology, not TUI preferences) out of
[tui].Env vars would keep working as overrides so existing setups are unaffected.
Use case
Operating behind a corporate proxy, v2rayN, or a satellite/mobile link, where
the first streaming request occasionally hangs until the 45s header timeout.
Today the only options are: accept the failure, or edit and rebuild the
binary. With
[stream].max_resumesandopen_timeout_secs, the retry posturebecomes a config change.
A second case: a provider that takes 60-90s to emit response headers on a cold
start.
open_timeout_secsalone (already available via env) fixes this, but isundiscoverable — it does not appear in
codewhale config dumpordocs/CONFIGURATION.md's[retry]section, only in the env-var appendix.Alternatives considered
deployment topology, and the codebase already has a config path for the
adjacent retry schedule.
codewhale config dumpwould still not show them.networktable. Considered and rejected as too broad forthis change; it would mix proxy/TLS concerns into the same edit.
Impact
Unlocks recovery from transient provider/proxy stalls without a rebuild. Also
makes the effective retry posture inspectable, which matters for diagnosing
"it retried for 15 minutes and then gave up" reports.
Additional context
Related: the turn-level resume budgets below this layer are never consulted for
stream-open failures — see the companion bug report on a turn failing with no
retry when the SSE request never receives response headers.