v0.10.1 integration: land the ready PRs together - #6672
Merged
Merged
Conversation
The Runtime API bearer / x-codewhale-runtime-token / legacy
x-deepseek-runtime-token checks and the computer-display owner checks
used `==`, and constant_time_eq existed in three private copies
(runtime_api/mobile.rs, runtime_api/web.rs, app-server). One shared
codewhale_core::secret_eq::constant_time_eq (the length-hiding
app-server variant) now serves every credential comparison; the copies
are deleted.
Tests: codewhale-core secret_eq 2 passed / 0 failed (new);
codewhale-tui runtime_api::{auth,web,mobile} 11 passed / 0 failed;
codewhale-app-server lib test build compiles.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
Logout deletes every provider API key, OAuth login, the account session and the Daytona token at once, with no prompt. It now asks for an explicit `yes`/`y` on a terminal and refuses (deleting nothing) when stdin is not a terminal unless `--yes`/`-y` is passed. Session revocation on logout is unchanged (not in this slice). Tests: codewhale-cli logout filter 8 passed / 0 failed, including new logout_confirmation_accepts_only_an_explicit_yes and logout_parses_yes_flag. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…and interact calls The "approve for session" key for a shell command was its arity-dictionary prefix, falling back to the first word, so a grant for `cd` or `bash` also covered `cd x && rm -rf ~` or `bash -c ...`, `rm tmp/x` covered `rm -rf ~`, and every exec_*_interact / wait call shared one `shell:<empty>` key. Now only a simple, dictionary-known, non-wrapper command keeps its family grant (`git status` still covers `git status -s`, not `git push`). Compound commands (chaining, pipes, redirects, substitution, `$`), wrappers/interpreters (bash/sh/env/sudo/xargs/python/node/find/ssh, docker run/exec, uv run, ...), env-assignment prefixes and unknown commands are keyed by the full normalized command (whitespace collapsed only outside quotes). Interact and wait calls are keyed by the exact call. Documented in RUNTIME_API.md. Tests: codewhale-tui approval_cache 26 passed / 0 failed, including new shell_grants_fail_closed_on_compound_wrapper_and_unknown_commands and shell_interact_grants_are_per_exact_call. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
A repo-level `.codewhale/config.toml` (or legacy `.deepseek/`) could set `notes_path = "~/.zshrc"`, and the auto-approved `note` tool would then append model text to the user's shell rc file. - Project-scope `notes_path` is honoured only as a relative path inside the workspace (resolved against it); `~`, `$`, absolute, rooted and `..` paths are ignored with a warning. User-config `notes_path` is unchanged. - The note tool refuses a notes target that is a symlink or not a regular file, and one placed in the workspace whose existing directory resolves outside it after following symlinks (a committed `notes.md -> ~/.zshrc` or `notes/ -> ~`). Tests: codewhale-tui project_config_tests 22 passed / 0 failed (new project_overlay_keeps_notes_path_inside_the_workspace); tools::shell::tests::note_tool_refuses_symlinked_targets_that_leave_the_workspace 1 passed / 0 failed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
… budget slot An approval card that expired unanswered reached the model as "Tool 'x' denied by user", and the Runtime API path (GPUI/web) called deny_tool_call on its timeout, recording the user's denial. Denied, timed-out, cancelled and unavailable calls also kept their per-turn tool-call budget slot although they never ran. - ApprovalResult gains TimedOut; the engine tells the model the request timed out and the user did not deny it (ExecutionFailed, not the PermissionDenied "denied by user" marker), and audits decision=timeout. - runtime_threads' approval timeout sends deny_tool_call_timed_out, so the receipt and the model both say timeout (the approval.decided event shape for clients is unchanged). - Any approval outcome that stops the call (denied, timed out, cancelled/unavailable) refunds its ToolCallBudget slot (#5170 already refunds planning-time gates). Tests: codewhale-tui approval filter 31 passed / 0 failed and tool_call_budget 4 passed / 0 failed, including new core::engine::tests::approval_timeout_is_reported_as_timeout_and_refunds_the_budget and the updated runtime_threads approval_timeout_denies_clears_ui_and_next_turn_can_start. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
Repository skill roots (`.agents/skills`, `skills/`, `.opencode`, `.claude`, `.cursor`, `.codewhale/skills` under the workspace) were read into runtime discovery without workspace trust, so a cloned repo's SKILL.md reached the model's skill list (and shadowed a same-named global skill) before the user ever ran /trust. Runtime discovery now skips project-scope roots until is_workspace_trusted(workspace) — the same predicate hooks and MCP use — checked lazily, only when such a root exists. Skipping is not silent: discovery adds a warning naming the skipped directories and `/trust`, visible in /skills and the skills block. Once trusted, project skills keep their precedence and a shadowed global skill keeps the existing "is shadowed by" warning. Audit/mutation catalogs are unchanged. Tests: codewhale-tui skills::tests::project_skills_require_workspace_trust 1 passed / 0 failed (new). The wider skill/plugin sweep did not run: the shared build volume filled (ENOSPC) before it could link. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…ted built-ins Markdown commands from a repository's `.codewhale/`, `.deepseek/`, `.claude/` or `.cursor/commands` (and `.codewhale/workflows`) loaded in any workspace and were dispatched ahead of the built-ins, so a cloned repo's `trust.md` or `undo.md` answered `/trust` or `/undo` with its own prompt, sent to the model as the user's own message, before the user ever ran /trust. - commands_dirs/workflow_dirs include workspace directories only when is_workspace_trusted(workspace) holds (the predicate hooks, MCP and project skills use). User-global commands are unchanged. - A workspace command (or alias) never stands in for a protected built-in: the ones that grant or revoke authority, hold credentials, or undo and discard work (/trust, /undo, /permissions, /mode, /config, /login, /logout, /restore, /hooks, /mcp, ... and the /jihua, /zidong mode aliases). The definition is skipped with a load error naming the rename; other built-ins (/help, /review) stay shadowable as FEAT-011/012 specify. - Test fixtures that exercise workspace commands and project skills now trust their temp workspace (test_support::trust_workspace), which also repairs the fixtures broken by the project-skill trust gate (294b616). - docs/architecture/command-dispatch.md states the gate. Tests: codewhale-tui focused run (user_registry, user_commands, epic_dispatch/discovery, command_palette, skills::, slash_completion, views::help, command_catalog, skill_lifecycle, commands::tests, tools::skill, codemode, cached_skills, prompts::tests, plus the B1 filters) 638 passed / 1 failed; the one failure was the new runtime B1 fixture in the next commit, fixed and rerun 1/0. New: workspace_command_never_replaces_a_protected_builtin, workspace_alias_never_replaces_a_protected_builtin, untrusted_workspace_commands_do_not_load. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
… and scrub for stored secrets Session JSON stored tool output verbatim: redaction ran only when a model request was built (client.rs), so a `cat ~/.codex/auth.json` or a printed bearer token was written to ~/.codewhale/sessions live (sweep B1). - The engine masks credentials in ToolResult content (and text content_blocks) once, in add_session_message, with the same masking the request boundary applies. The transcript is what session files, checkpoints and later requests are built from. The double-confirmed `[redaction] model_bound = "disabled"` opt-out is honored the same way. - Runtime API threads mask tool output before it reaches the durable tool item and the event log built from it (always on: receipts fan out to UI clients). - `codewhale doctor` gains a "Stored Sessions" finding: it scans the newest 50 session/checkpoint files and says how many still hold credentials in stored tool output. It never rewrites anything. - `codewhale sessions scrub-secrets` reports every affected file; `--apply` rewrites them atomically, touching only tool_result text. Documented under docs/CONFIGURATION.md "Stored sessions", with the reminder that masking a copy does not un-leak a credential: rotate it. Not covered: runtime item/event files written by older builds are not scrubbed by the command (it walks the sessions store only). Tests: codewhale-tui core::engine::tests::tool_output_credentials_are_redacted_when_they_enter_the_transcript, session_secret_scrub (2), runtime_threads::tests::runtime_tool_items_store_tool_output_with_credentials_masked all pass (in the 638/1 focused run above; the runtime fixture failed once on a single-line JSON fixture and passes 1/0 after the fix); runtime_threads:: 226 passed / 0 failed / 2 ignored. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
The note tool wrote through tokio's File and returned without flushing, so the append could still be in flight on a blocking thread when the tool reported "Note appended". Under parallel load the notes-symlink test read the file before the bytes landed (2 of 4 runs of the approval/notes filter). Flush before reporting success. Tests: codewhale-tui filter `approval tool_call_budget constant_time logout note project_overlay auth::` 483 passed / 1 failed, twice; the note test passes in both. The one failure is task_manager::tests::pending_approval_suspends_idle_and_timeout_denial_settles_failed, a pre-existing race: runtime_threads tests set the process-global TEST_APPROVAL_DECISION_TIMEOUT_MS that this test's drive reads; it passes alone (1/0). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 6 to 7. - [Release notes](https://github.com/actions/setup-python/releases) - [Commits](actions/setup-python@v6...v7) --- updated-dependencies: - dependency-name: actions/setup-python dependency-version: '7' dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
…ession lock Review findings on 000ab6b: - Runtime API tool items copied the tool's metadata unmasked. exec_shell metadata carries summary/stdout_summary/stderr_summary (first stdout lines), so `cat auth.json` on a desktop or web thread still stored and emitted the token. Metadata now goes through a model-bound JSON redactor (new redact_model_bound_json_secrets); the test uses real shell-shaped metadata. - The redactor missed credentials in compact JSON and query strings ({"tokens":{"access_token":"eyJ…"}}, jq -c, curl, ?access_token=…): the whole document is one word whose first key is harmless. Such words are now split on their structure and each member gets the same keyed and bare-token checks, delimiters byte-exact; already-masked values (`token=***`) are left alone. doctor no longer reports a false clean. - `sessions scrub-secrets --apply` re-reads and rewrites each file under the per-session lock every save takes, so it cannot lose a concurrent save. Invalid ids are left untouched. - docs: say what is still stored unmasked (spillover files, shell evidence artifacts, tool-call inputs). Tests: codewhale-secrets 77/0 (1 ignored); codewhale-config lib 695/0 (1 ignored, incl. new compact_json_and_query_strings_mask_their_credentials); codewhale-tui lib focused (runtime_tool_items_store session_secret_scrub approval_cache user_registry memory::note redact export) 252/0 (1 ignored). Wider tui set (user_commands skills:: project_config_tests note tool_call_budget core::engine::tests command_palette epic_ runtime_threads::tests tools::shell client::tests) 1379/2: the two core::engine::tests::mcp_boot_* failures reproduce identically with these edits stashed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…pace; remote commands protected Review findings on ecdc840, d0e3158 and 9253eff: - Shell grants ignored flags, so `git status` covered `git -ccore.fsmonitor=./evil status`, `cargo build` covered `cargo build --config build.rustc-wrapper=…`, and `git config user.name` covered `git config core.fsmonitor`. Any `=` argument, a code-running option (-c, -C, --config, --exec, -f, --file, -e, …) or a config-style family (config/set/remote) now keys the grant on the full command. - `/note clear|edit|remove|list` followed a committed `notes.md -> ~/.zshrc` or a symlinked notes directory. They now apply the note tool's rule: no symlinked file, directory must resolve inside the workspace. - rc/remote-control, relay, remote-env, profile and share are protected built-ins: a trusted repo's rc.md can no longer answer `/rc off`. Tests: codewhale-tui lib focused (approval_cache user_registry memory::note ...) 252/0 (1 ignored), run together with the previous commit's set. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
Resolves crates/secrets/src/redact.rs against main's redact_json_model_bound_secrets (auto-review, 2c672a9): the branch's own model-bound JSON helper is dropped and Runtime API tool metadata now uses main's, so there is one JSON redactor per policy. The structured compact-JSON word pass is kept. Tests after the merge: codewhale-secrets 77/0 (1 ignored), codewhale-config lib 696/0 (1 ignored), codewhale-tui lib focused (runtime_tool_items_store session_secret_scrub approval_cache user_registry user_commands memory::note redact export auto_review skills:: project_config_tests tool_call_budget core::engine::tests::approval command_palette epic_ runtime_threads::tests tools::shell client::tests) 1242/0 (3 ignored). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
turn_loop.rs conflicted: main (2b440fa, #6584) removed the questions_allowed parameter, and this branch added tool_call_budget beside it. Kept tool_call_budget and dropped questions_allowed. Checks: cargo check -p codewhale-tui --tests clean; focused approval|turn_loop|redact: 462 passed, 1 failed. The failure is task_manager pending_approval_suspends_idle_and_timeout_denial_settles_failed, which is unchanged from main and passes alone (120 ms idle window under parallel load). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…s/setup-python-7 chore(deps): bump actions/setup-python from 6 to 7
Merge main at 24a7b99 (runtime split #6586, code mode MCP default #6583). Conflicts: - lib.rs: keep this branch's `mod session_secret_scrub`; `session_tree` now comes from codewhale_runtime. - turn_loop.rs: execute_planned_tools takes main's NestedGateEnv, which already carries the tool-call budget; this branch's approval refunds go through nested_gate_env.tool_call_budget. - main's new code-mode nested approval gate did not handle this branch's ApprovalResult::TimedOut (E0004). A timed-out nested call now reports NestedDecision::TimedOut with the same "did not run, user did not deny it" error as a direct call, and every refused nested approval hands its budget slot back like the direct path. CI fixes: - Test (ubuntu/windows): `codewhale doctor` created CODEWHALE_HOME via SessionManager::default_location() in the stored-secrets scan. It now resolves the sessions dir with the read-path resolver (resolve_state_dir), so the diagnostic stays read-only. - Test: apply_slash_menu_selection_honors_user_argument_metadata_and_builtin_override (new on main) writes workspace commands without trusting the workspace; this PR intentionally loads workspace commands only in a trusted workspace, so the test now trusts its temp workspace. - Lint blocking-calls: session_secret_scrub.rs is a synchronous module (doctor calls it inside spawn_blocking; `sessions scrub-secrets` is a one-shot sync subcommand like `sessions list`); budget raised by its 2 std::fs sites. - CodeQL cleartext-logging: the printed count/path fields were named *_with_secrets; renamed to flagged_tool_results / flagged_files. They never held secret values. Local: clippy -p codewhale-tui --all-targets --all-features (CI flags) clean; cargo fmt --check, git diff --check clean; codewhale-tui lib focused 93+1577 passed, 1 failed (task_manager pending_approval_suspends_idle_and_timeout_denial_settles_failed: unchanged from main, passes 3/3 alone, wall-clock flake under load); integration diagnostic_read_only 13/13; codewhale-cli diagnostic_dispatch_read_only 1/1; blocking-calls, dead-code, command-crate-boundaries, module_graph --check, provider-registry, command-migration-manifest, reqwest-builders, locale checks pass. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…ice B) The offline seed now carries MiMo 2.6's canonical facts under upstream Models.dev keys (xiaomi/mimo-v2.6-pro, xiaomi/mimo-v2.6-flash), projected unedited from https://models.dev/catalog.json (2026-09-25) onto the fields ModelsDevModel reads. Slice A normalized a namespaced key's vendor only on live refresh, so the same entry offline surfaced as a non-route `xiaomi` provider. The vendor namespace is never a CodeWhale id, so it now normalizes in both modes, and gap-filling compares on the normalized provider identity so a verbatim bundled provider row still shadows its namespaced twin. Kept on purpose: the providers.xiaomi-mimo 2.6 rows (they still win and carry the verified text-only qualification; retiring them would widen MiMo to image/audio/video input and drop the bare-id rows that reasoning_support, provider_model and the fleet label read), and the bare deepseek-v4-pro/-flash keys (every hosted DeepSeek base_model, the route layer and pricing name those ids; renaming them is slice C). No write-capable generator exists: scripts/catalog_models_dev.py fails closed on --write by design, so the entries were staged per docs/CATALOG_REFRESH.md and validated with its snapshot --check and drift. Tests: codewhale-config --lib 698 passed, 0 failed, 1 ignored (catalog filter: 54 passed incl. 2 new); codewhale-models 44 passed; scripts/catalog_models_dev_test.py 8 passed; snapshot --check ok; cargo fmt --check, blocking-calls, dead-code, command-crate-boundaries and module_graph --check all pass. Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
These changes move to a separate change. The protected built-in command list from the same commit stays in this PR. Reverts the approval_cache.rs and commands/groups/memory/note.rs hunks of a8755ae. Tests: none run for this commit alone; the focused set runs on the final head of this PR (see the fix(ci) commit that follows). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This change moves to a separate change. This reverts commit d0e3158 (lib.rs, tools/shell.rs and tools/shell/tests.rs). The note-flush fix from 133a8a2 stays. Tests: none run for this commit alone; the focused set runs on the final head of this PR (see the fix(ci) commit that follows). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This change moves to a separate change. This reverts commit ecdc840 (tools/approval_cache.rs and docs/RUNTIME_API.md). Tests: none run for this commit alone; the focused set runs on the final head of this PR (see the fix(ci) commit that follows). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-only scan, synthetic JWT from parts - The runtime-contract fixture puts its representative skill in the workspace's project skills directory. Project skills now load only in a trusted workspace, so the fixture trusts its workspace and keeps measuring a loaded skill. The budget is unchanged from main. - `sessions scrub-secrets` takes its printed counts from the read-only scan. The rewrite under the per-session lock re-reads the file and returns only success, so no value flows back from the lock call into the report. - The redaction test's synthetic, unsigned JWT is assembled from its three segments at runtime. Tests (local, macOS): - codewhale-config lib `persistence`: 28 passed / 0 failed. - codewhale-tui lib focused: approval_cache 24/0, session_secret_scrub 2/0, user_registry 31/0, user_commands 26/0, memory::note 14/0, tools::shell 160/0 (1 ignored), skills:: 225/0, project_config_tests 33/0, tool_call_budget 4/0, core::engine::tests::approval 2/0, runtime_threads::tests::approval 8/0 with --test-threads=1. In one shared process, approval_wait_heartbeat_is_never_sequenced_after_the_decision fails when it runs beside approval_timeout_denies_clears_ui_and_next_turn_can_start: both tests use the process-global test approval timeout. Both tests are on main unchanged, and CI's nextest runs each test in its own process. - scripts/check-runtime-contract-budget.py: PASS, all 55 metrics at budget. - cargo clippy -p codewhale-config -p codewhale-tui --all-targets --all-features --locked with CI's flags: clean. cargo fmt --check: clean. - check-blocking-calls-budget, check-dead-code-budget, check-command-crate-boundaries, split/module_graph.py --check: pass. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
) The offline seed carried Codewhale's deliberate holds as hand edits: prices left out where a flat rate would mislead, DeepSeek's output limit kept at its published 384K. A live Models.dev refresh replaces seed rows wholesale, so every hold vanished on a networked install: live rows priced MiMo and DeepSeek, listed Alibaba plan rows at $0 (read as free), and raised DeepSeek's output limit to 393216. The holds now live in `crates/config/assets/catalog_corrections.json` as field patches in the cloud-facts `ModelFact` shape. The signed layer's patch code applies them to every Models.dev row as it is hydrated, bundled seed and live refresh alike. So a correction ranks above both Models.dev layers and below signed cloud facts (which may still correct it) and provider rosters. - `ModelFact.pricing_withheld` (optional reason) clears the row's price so the route reports `PricingSku::UnknownOrStale`. It is an optional field; payloads without it parse unchanged. - `catalog_patch::apply_patches` is shared: the signed layer calls it with row materialization, corrections without, so they never add a row. The loader also refuses hide, deprecate, `allow_unlisted`, labels, a patch that changes nothing, and any limit patch without a reason. - A corrected row keeps its own `source`. `CatalogCompiler::with_live` and other code sort rows by source, and a relabelled live row was placed above signed facts (an existing cloud-facts test caught it). The price a correction owns reports `CatalogSource::CodewhaleBundled` through `cost_source`. - The declared-but-unused `CatalogCompiler.codewhale_bundled` field is removed; the layer docs (catalog.rs, CATALOG_REFRESH.md, CLOUD_FACTS.md) now describe the rank where corrections actually apply. Each hold was checked again on its merits: - kept: DeepSeek native pricing (time-aware table in tui pricing.rs), xAI (rates double past 200K), MiMo (PAYG and Token Plan keys look the same), the four Model Studio plan providers (quota), StepFun (plan and PAYG share ids), DeepSeek native output 384000 (published 384K; which value the API accepts is unverified, and the lower bound is always accepted); - added: MiniMax-M3 (upstream price tiers at 512K, same rule as Grok; the seed already listed it unpriced without saying why); - dropped: aggregator-hosted DeepSeek (an aggregator's published rate is its own price, as it is for every other aggregator row we price), OpenRouter dots :free (a free model is free), zai GLM-5.3 (the zai route is the pay-as-you-go API, not the Coding Plan the hold cited), MiMo's text-only image hold (made harmless by the previous commit) and MiMo's 1,000,000 context (the only citation was our own rows). Tests (focused): - codewhale-config --lib: 706 passed, 0 failed, 1 ignored; --test configured_models: 9 passed - codewhale-tui `-- vision image badge capabilit pricing provider_lake models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi mimo modelstudio grok stepfun minimax`: 872 passed, 0 failed, 2 ignored (one earlier run failed route_budget::uncatalogued_remote_model_keeps_a_conservative_ceiling; it passes alone and on rerun, and reads output-limit env vars without the env lock) - cargo fmt --check clean; clippy -p codewhale-config -p codewhale-tui --all-targets --all-features with CI flags clean; budget, boundary and module-graph gates pass; changelog chain passes (public-copy 6/6) Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…6396) Regenerating the offline seed from Models.dev (#6612) showed a second kind of offline-only hold. The seed carried the reasoning controls Codewhale relies on, but live Models.dev rows replace them: - Grok: xAI documents `high` as the default effort (6be755f / #6501). Upstream lists the ladder with no default. - MiniMax: M2.7 always thinks; M3 is adaptive or disabled, with a different default per dialect (47163da). Upstream lists a toggle or nothing. - qwen3.8-max: thinking-only (the owner's console, ec4a5ae). Upstream lists a toggle. - Muse Spark: also takes the `none` and `ultra` efforts (a50b653). - Step 3.5: base model uses adaptive reasoning (StepFun guide, recorded 2026-09-19). Upstream lists low/high. The picker, `/effort` and the effort clamp read these, so today a networked install loses them. `ModelFact.reasoning_options` (optional) replaces a row's reasoning controls and keeps any signed annotation already on the row. The 18 rows above get it as a correction, each with its reason. GLM-5.3's high/max ladder was left out: the seed only inherited it from GLM-5.2 as a placeholder, and upstream now publishes low/high/max for 5.3. The Coding Plan qwen `budget_tokens` entries were left out too: they were copies of Token Plan rows, not a Codewhale decision. Tests (focused): - codewhale-config --lib: 707 passed, 0 failed, 1 ignored; --test configured_models: 9 passed - codewhale-tui `-- vision image badge capabilit pricing provider_lake models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi mimo modelstudio grok stepfun minimax effort reasoning thinking muse`: 1169 passed, 0 failed, 2 ignored - cargo fmt --check clean; clippy -p codewhale-config --all-targets --all-features with CI flags clean; changelog chain passes (public-copy 6/6) Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first commit dropped the aggregator-hosted DeepSeek price hold on the grounds that an aggregator's published rate is its own price. Running the TUI pricing tests against a seed generated from Models.dev (#6612) showed the hold's real purpose: `crates/tui/src/pricing.rs` prices these routes from a reviewed provider-docs table, and a catalog rate outranks that table. For Fireworks V4 Pro, for example, the table has 1.74 per 1M input and Models.dev has 1.20. The hold is restored for the nine hosted DeepSeek rows, with that reason. Also from the generated seed: - Novita DeepSeek rows: max_output 384000. Models.dev lists 393216; this is the same published-384K question as native DeepSeek. - grok-4.3: reasoning_options []. xAI documents no effort control for it, so Codewhale sends none (#6501); Models.dev lists a ladder. All three are no-ops against the current hand seed. They take effect on live refresh today, and offline once the seed is generated. Tests (focused): codewhale-config --lib 707 passed, 0 failed, 1 ignored. Changelog chain: public-copy 6/6. Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cause: the 80-column footer reserved only 24 columns for route identity. gpt-5.6 with thinking: high fit, but xhigh and longer requested-to-effective labels exceeded the budget, so the entire effort field disappeared. Let the route budget floor grow from 24 to 32 columns with terminal width. This preserves all reasoning labels at 80 columns while retaining context and performance readings on narrower rows. Add a rendered-footer regression covering all nine ReasoningEffort variants and sync the Unreleased fix note. Regression proof with the original budget (Minimal, XHigh, Ultra missing): CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true cargo test -p codewhale-tui --lib -- footer_keeps_reasoning_label_for_every_effort_tier test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 13774 filtered out; finished in 0.27s Final focused validation, including regression, narrow layouts and hitboxes: CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true cargo test -p codewhale-tui --lib -- tui::ui::frame:: test result: ok. 30 passed; 0 failed; 0 ignored; 0 measured; 13745 filtered out; finished in 0.67s CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true cargo fmt --all sh scripts/sync-changelog.sh sh scripts/sync-changelog.sh --check git diff --check Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Conflict resolution: - memory/note.rs: kept #6601's lowercased command dispatch, ensure_notes_target_in_workspace pre-check and notes_path resolver, and threaded #6685's workspace root through every subcommand. read_notes stays pub(crate) (dock NOTES view) but now takes the workspace and reads through fs_confined::read_to_string; tui/workspace_context.rs passes the workspace. Restored the std::fs import the pre-check needs. - skills/install.rs: test-module add/add; kept #6679's download_with_cap server and test alongside #6685's registry-name tests (braces verified). - CHANGELOG.md: kept both Security bullets; crates/tui/CHANGELOG.md regenerated with scripts/sync-changelog.sh. Verification: cargo check -p codewhale-tui --tests clean. cargo test -p codewhale-tui --lib -- memory::note anchor skills::install fs_confined compaction::last_round project_context: 183 passed; 0 failed. workspace_context: 14 passed; 0 failed. Codex read-only review: no findings. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cause: App initialization could retain an options-default model despite a resolved config route. Local Ollama adoption treated missing credentials as permission to replace an explicit provider/model. The context meter counted the startup system prompt against the fallback window before any conversation. Use the configured startup model and retain whether the resolved config named a route before runtime synchronization materializes defaults. Reject both probing and late adoption for explicit routes, including missing-key recovery. Show zero context usage until messages or provider usage exist; retain normal pressure readings afterwards. Update and sync the Unreleased changelog. Validation (all cargo commands used CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true): cargo fmt --all sh scripts/sync-changelog.sh git diff --check cargo test -p codewhale-tui --lib -- first_run_route_ Original production code with the new regression tests: test result: FAILED. 1 passed; 3 failed; 0 ignored; 0 measured; 13773 filtered out; finished in 0.28s Restored fix: test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 13773 filtered out; finished in 0.23s cargo test -p codewhale-tui --lib -- adopting_a_local_model test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 13776 filtered out; finished in 0.41s cargo test -p codewhale-tui --lib -- local_ollama_probe_leaves test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 13776 filtered out; finished in 0.11s cargo test -p codewhale-tui --lib -- context_usage test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 13771 filtered out; finished in 0.12s cargo test -p codewhale-tui --lib -- startup_provider test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 13775 filtered out; finished in 0.08s Local library-test evidence only; no packaged runtime, CI, or release claim. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…dpoints Cause: #6687 treated any default_text_model as an explicit route, but the generated first-launch config.toml writes default_text_model = DEFAULT_TEXT_MODEL, so an unconfigured person who first launched without Ollama lost local discovery on every later launch. Endpoint-only routes (legacy top-level base_url, [providers.<active>].base_url, or DEEPSEEK_BASE_URL/CODEWHALE_BASE_URL) with a missing key were still replaceable by a live Ollama. The empty-session meter shortcut also zeroed a submitted first turn before the engine mirrored messages back (one_owner_tests::context_cap_warns_once_in_the_posture_bar), and with no messages the reported prompt was dropped in favour of the system-prompt-only estimate. Fix: a template default_text_model (no provider, value DEFAULT_TEXT_MODEL) is not a choice; explicit env provider/model overrides and a configured active-route endpoint (new Config::active_route_endpoint_configured) are. The meter treats a loading turn or a user history cell as a started conversation, and with no messages lifts the estimate to the reported prompt. model_picker fleet_case_distinct_pins test encoded the superseded options-default startup model: the configured default_text_model is now the startup route and leads the list, so the test asserts pin order by position. test result (CARGO_BUILD_JOBS=4 cargo test -p codewhale-tui --lib -- context_cap_warns_once_in_the_posture_bar fleet_case_distinct_pins first_run_route_ context_usage): test result: ok. 16 passed; 0 failed; 0 ignored; 0 measured; 13783 filtered out; finished in 1.31s Same filter with init.rs/frame.rs reverted to f45d1df: test result: FAILED. 5 passed; 4 failed; 0 ignored; 0 measured; 13790 filtered out; finished in 0.40s CARGO_BUILD_JOBS=4 cargo check --workspace --tests: Finished sh scripts/sync-changelog.sh --check: crates/tui/CHANGELOG.md is up to date Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: runtime-contract ratchet failed on PR #6672: 6fae324 added a 53-byte refspec rule to the Git tool schema (plan/act/operate) and #6644 added 119 bytes of overwrite warning to revert_turn (act/operate). Fix: drop the Git action description's duplicate of the fetch sentence already in the tool description (the refspec rule stays, and the validator's error names the fix); tighten revert_turn wording while keeping the whole-workspace overwrite warning. Lock in the lower ceilings. test result: python3 scripts/check-runtime-contract-budget.py [runtime-contract-budget] PASS: all 55 metrics are exactly at budget. (before --update: PASS: 55 ceilings respected; 6 can be tightened; plan -13 bytes, act/operate -33 bytes) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: Version drift failed in .github/scripts/release-workflows.test.js: "docs/zh_hans/FLEET.md feeds Rust and must classify heavy". fad53fa added include_str!("../../../../docs/zh_hans/FLEET.md") in crates/tui/src/fleet/host.rs, but the ci.yml change filter still sent that path to the docs-only light arm. Fix: add it to the heavy arm next to the other embedded docs. test result: node .github/scripts/release-workflows.test.js Change detection OK: 134 Rust-consumed non-.rs paths classify heavy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: CodeQL py/clear-text-logging-sensitive-data flagged die() in scripts/catalog_models_dev.py because data returned by a function named scrub_secrets() flowed into an error naming a provider/model row id. The value is already-scrubbed public models.dev data, not a secret. Fix: rename scrub_secrets -> strip_sensitive_fields; behaviour unchanged. test result: python3 scripts/catalog_models_dev_test.py Ran 16 tests ... OK Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
First-run recovery treated a fresh home as having no saved route, even when config selected one. A usable route that skipped onboarding left no provider step receipt, and setup treated saved-but-unchecked credentials as verified. Recognize explicit startup routes in the recovery predicate. Record an unchecked first-run route as configured through the existing setup owner; reserve verified for observed success. Keep the unconfigured picker and local Ollama discovery, prior setup decisions, and Fleet verification requirements. Generalize the existing locked sidecar updater and migrate every caller so the new startup receipt preserves privacy decisions and rejects corrupt/busy state. Extend and synchronize the existing first-launch changelog bullet. Validation (all cargo commands used CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true): cargo fmt --all sh scripts/sync-changelog.sh sh scripts/sync-changelog.sh --check git diff --check cargo build -p codewhale-tui cargo test -p codewhale-tui --lib -- first_run_ test result: ok. 18 passed; 0 failed; 0 ignored; 0 measured; 13786 filtered out; finished in 1.06s Negative control: restore the old recovery predicate and remove startup receipt writing while retaining the new tests and type vocabulary. All three configured route regressions fail; both production files were restored afterward. The first control attempt stopped at unused-import compilation; remove only the now-unused imports in the control, without weakening warnings, and rerun: cargo test -p codewhale-tui --lib -- first_run_configured_ test result: FAILED. 0 passed; 3 failed; 0 ignored; 0 measured; 13801 filtered out; finished in 0.13s Restored fix and adjacent setup/provider receipt checks: cargo test -p codewhale-tui --lib -- first_run_ configured_provider_receipt_only_verifies_observed_success tui::setup::progressive_tests doctor_verdict_tests provider_switch_model_override_updates_target_provider_model_slot provider_switch_auth_error_restores_previous_provider_and_model model_picker_apply_is_session_local_until_startup_default_is_requested test result: ok. 33 passed; 0 failed; 0 ignored; 0 measured; 13771 filtered out; finished in 1.92s cargo test -p codewhale-config --lib -- setup_state::tests telemetry_metadata_update telemetry_disclosure_records test result: ok. 21 passed; 0 failed; 0 ignored; 0 measured; 723 filtered out; finished in 0.01s Direct PTY, rebuilt debug binary, fresh config-only homes, 130x40: - Supplied OpenAI/loopback/gpui-fixture config opens at the composer, no picker. Footer: OpenAI-compatible / gpui-fixture. Persisted provider_model status: configured; result: provider=openai, model=gpui-fixture; configured, not checked. - Configured OpenAI route without credentials opens the navigable picker with OpenAI-compatible selected. - Both processes exited normally. No inference request was sent. Reproduction boundary: a freshly built baseline at e4c94c8 also skipped the picker in a direct PTY, but produced no provider-step receipt. The reported forced-DeepSeek screen was not reproduced here. The full fixture plus running Ollama tmux check remains pending: tmux socket creation failed with Operation not permitted in the sandbox. No hosted CI, packaged release, or push claim. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: OpenRouter's /v1/models now lists `~`-prefixed "latest" alias ids (`~deepseek/deepseek-pro-latest`). They fail `valid_catalog_model_id`, and one invalid row fails the whole roster closed, so the OpenRouter lake scope was permanently `failed/invalid_response` and `fresh_provider_live_pricing_quote_at` never returned a rate: every OpenRouter turn showed "rate unavailable". Separately, the main interactive turn froze its dispatch quote only from the lake and never consulted the operator-declared `[[custom_models]]` rate that background/review envelopes (`effective_route_envelope`) already honor. Fix: `parse_openrouter_models_response` drops `~` alias rows (moving aliases, not billing identities); the fail-closed validator is unchanged for every other id. A new `CodewhaleClient::configured_pricing_quote_at` is shared by `effective_route_envelope` and the turn-loop dispatch boundary, which now prefers a declared rate (scoped to the installed client's endpoint fingerprint) before the lake quote. Not covered: server-tag suffixed ids (`:nitro`, `:free`) still need an exact catalog or declared row; their price can differ from the base id, so they are not stripped. Tests (each fails with its fix hunk reverted, passes with it): - client::tests::fetch_catalog_delta_skips_openrouter_latest_aliases_and_keeps_prices - core::engine::tests::main_turn_dispatch_freezes_declared_custom_model_rate test result: ok. 2 passed; 0 failed (reverted: 0 passed; 2 failed) Related filters (fetch_catalog, openrouter, configured_model_client_tests, exact_turn_snapshot, effective_route, provider_live_pricing): test result: ok. 46 passed; 0 failed cost_status, provider_catalog_live, core::engine::tests::exact: test result: ok. 60 passed; 0 failed Refs #6690 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: `codewhale exec` accepted its prompt only as positional argv, so a prompt past the kernel's per-argument ceiling (131071 bytes on Linux) failed with E2BIG before Codewhale started, with no alternative transport. Fix: add `--prompt-file <PATH>` (`-` = stdin), mutually exclusive with the positional prompt, which is now `required_unless_present = "prompt_file"`; a lone positional `-` also reads stdin. `resolve_exec_prompt` produces the effective prompt before any model call and fails loudly on a missing, unreadable, non-UTF-8 or empty source. Reading goes through a shared `read_capped_text` (also used by `apply`'s stdin patch reader) with a named 64 MiB byte limit. `-` is refused with `--parent-death-watch`, which owns stdin for Fleet workers. Implicit stdin reading when no prompt is given is not added: exec callers (Fleet) inherit pipes on stdin. An argv-side byte guard is not possible: the kernel rejects the exec before Codewhale runs. Tests: - cargo test -p codewhale-tui --lib -- exec_prompt_file read_capped_text exec_accepts_split exec_keeps test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 13795 filtered out - With resolve_exec_prompt reverted to the pre-fix argv join: test result: FAILED. 0 passed; 1 failed (left: "" vs the 540 KB file body) - Built binary: 540001-byte file via --prompt-file and via `cat | exec -` resolve and reach provider routing; empty stdin -> "exec prompt from stdin is empty"; missing file -> "failed to open --prompt-file nope.txt"; positional + --prompt-file -> clap ArgumentConflict. Fixes #6688 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: with `provider = "openai"`, `default_text_model = "gpui-fixture"` and no `[providers.openai].model`, the root alias is the active route's fallback and nothing records which provider it belongs to. - TUI: a session-local `/provider <other>` left the alias at the root, so `/provider openai` afterwards resolved with `openai != api_provider()` and fell through to the catalog default `gpt-5.6`. - Engine: `reconcile_root_model_aliases` relocated the alias only when the incoming route failed validation. A pass-through incoming route (openrouter, zai, ...) inherited `gpui-fixture`, and `GET /v1/providers` advertised OpenAI's `default_model` as `gpt-5.6`, which is what the desktop model chip picked for "OpenAI-compatible". Fix: `Config::root_model_alias_owned_by_outgoing` names the outgoing route and value when that route has no leaf and was actually resolving the alias. The persisted writer moves it onto the outgoing leaf and clears the root; the TUI switch applies the same move in memory. The resolution order is then providers.<p>.model, the root alias while its own provider is active, then the catalog default. Aliases the outgoing route never resolved (shadowed by a leaf, `auto`, foreign ids dropped by `default_model`) keep their previous handling. Two persistence tests that asserted the alias stays at the root in the no-leaf case now assert it moves to the outgoing leaf; switching back is still checked. Tests (focused, CARGO_BUILD_JOBS=4): - new: tui::ui::tests::provider_switch_back_lands_on_root_default_owned_by_that_provider and runtime_api::tests::switch_provider_away_and_back_keeps_the_root_default_with_its_route fail without the fix (left: "gpt-5.6" right: "gpui-fixture"; openrouter body model "gpui-fixture"): test result: FAILED. 3 passed; 2 failed; 0 ignored; 0 measured; 13796 filtered out - with the fix, filter root_default: test result: ok. 5 passed; 0 failed; 0 ignored; 0 measured; 13796 filtered out - filters config_persistence default_text_model alias route_save save_route provider_picker route_preferences provider_switch switch_provider root_default: test result: ok. 469 passed; 0 failed; 0 ignored; 0 measured; 13332 filtered out - filters config::tests model_inventory route_runtime runtime_api::tests::switch runtime_api::tests::get_config provider: test result: ok. 1302 passed; 0 failed; 1 ignored; 0 measured; 12498 filtered out Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Review follow-up for #6691. The live OpenRouter /v1/models lists six routers (openrouter/auto, auto-beta, fusion, pareto-code, bodybuilder, typesafe/jev-router) at "prompt":"-1","completion":"-1". parse_price rejected negatives and the conversion collected with `?`, so the whole refresh still failed and every turn stayed "rate unavailable". - A negative OpenRouter price now means "no fixed rate": the row is kept with cost None. Rows that do not decode, carry an id the catalog cannot hold, or an unreadable/implausible price are skipped and counted with a warning. A response that is not a `{"data": [...]}` list, or whose rows all fail, is still rejected. - declared_or_catalog_quote: a [[custom_models]] row with no rates yields to the exact endpoint's catalog price at every dispatch boundary (main turn, background envelope, EffectiveRouteEnvelope::capture); it stays frozen rate-less only when the catalog has none. The turn-loop logic moved into client::main_turn_pricing_quote_at so it is testable. - A pinned OpenRouter vendor no longer blocks an explicit declared rate for the exact route; the aggregate catalog price stays blocked. Against the saved live list (458 rows): refresh succeeds with 440 rows, 434 priced; deepseek/deepseek-v4-pro prices at 0.95526/1.91052 per M. Focused (cargo test -p codewhale-tui --lib): fixed: test result: ok. 10 passed; 0 failed; 0 ignored; 0 measured; 13793 filtered out reverted (neg price, rate-less fallthrough, vendor pin): test result: FAILED. 0 passed; 3 failed; 0 ignored; 0 measured; 13800 filtered out reverted (per-row decode skip): test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 13802 filtered out related (fetch_catalog openrouter configured_model effective_route provider_live_pricing exact_turn_snapshot main_turn cost_status provider_catalog_live): test result: ok. 120 passed; 0 failed; 0 ignored; 0 measured; 13683 filtered out Fixes #6690 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…ault Review finding on #6693: moving `default_text_model` onto the outgoing route's leaf removed only that key. `Config::load` then copies the legacy root `model` into `default_text_model` whenever the latter is unset, so a stale legacy value that the root default had been shadowing on every route came back on the incoming route. With `model = "deepseek-v4-flash"`, `default_text_model = "gpui-fixture"`, `provider = "openai"`, switching to openrouter resolved `deepseek/deepseek-v4-flash`. Fix: every place the route writer removes the root alias now goes through `unset_root_model_aliases`, which clears `default_text_model` and the legacy root `model` together (the legacy key resolved on no route while the alias existed). `root_model_alias_owned_by_outgoing` also treats a legacy-only root `model` as the alias, the same fallback `default_model` applies, so a legacy-only choice moves onto the outgoing leaf instead of being inherited. The switch-away-and-back persistence test now expects that move for the legacy `model` key as well. The TUI session-local switch needed no change: it writes the resolved model onto the incoming leaf in memory, which shadows the legacy value. Tests (focused, CARGO_BUILD_JOBS=4): - new: config_persistence::tests::moving_root_default_off_the_root_does_not_resurrect_a_shadowed_legacy_model and runtime_api::tests::switch_provider_does_not_resurrect_a_shadowed_legacy_root_model fail without the fix (openrouter body model "deepseek/deepseek-v4-flash"): test result: FAILED. 6 passed; 2 failed; 0 ignored; 0 measured; 13796 filtered out - with the fix, filters resurrect config_persistence default_text_model alias route_save save_route provider_picker route_preferences provider_switch switch_provider root_default legacy_model: test result: ok. 477 passed; 0 failed; 0 ignored; 0 measured; 13326 filtered out - filters config::tests model_inventory route_runtime runtime_api::tests::get_config provider: test result: ok. 1303 passed; 0 failed; 1 ignored; 0 measured; 12499 filtered out Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…ch (#6688) Review of #6692: 1. A lone positional `-` read stdin, which changed behavior for raw-prompt callers: cloud dispatch builds `codewhale exec --auto <job.prompt>`, so a job whose prompt is `-` would block on or fail reading stdin. Only `--prompt-file -` reads stdin now; a positional `-` is prompt text again. Help text, docs/MODES.md (+ zh_hans) and the CHANGELOG say so. 2. The tests called `resolve_exec_prompt` directly and would pass with the exec dispatch call site reverted. A new integration test runs the real `codewhale-tui exec` binary against a wiremock provider and checks the user message the model receives: a 200 KiB `--prompt-file` body, a `--prompt-file -` stdin body, and a literal `-` with stdin piped (never read). Tests: - cargo test -p codewhale-tui --lib -- exec_prompt_file read_capped_text exec_accepts_split exec_keeps test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 13795 filtered out - cargo test -p codewhale-tui --test integration -- exec_turn_usage test result: ok. 9 passed; 0 failed; 0 ignored; 0 measured; 179 filtered out - Dispatch call site reverted to `join_prompt_parts(&args.prompt)`: exec_prompt_reaches_the_model_from_prompt_file_and_stdin test result: FAILED. 0 passed; 1 failed (left: "") - Lone positional `-` arm restored: the unit test blocked reading stdin (killed after 9 minutes), which is the regression itself. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
… on a keyed route The reported first-launch symptom (a configured openai/gpui-fixture home opening on "Choose your model provider" with DeepSeek focused) was not a config-resolution bug. The tmux server used for the repro carries CODEWHALE_PROVIDER=deepseek and CODEWHALE_MODEL=deepseek-v4-flash in its global environment, so every new session inherited them. Config::load then resolved deepseek per the documented env > file precedence (instrumented: provider=deepseek already at load_config_from_cli_with_effective_profile). `codewhale model resolve` from a shell without those vars reported openai, which hid the difference. With the vars unset, the unmodified release build aaaadfe starts on the composer with "OpenAI-compatible · gpui-fixture". It did expose a real bug. After Esc on the recovery picker, `/provider openai` switched to a keyed route, but switch_provider never recomputed app.onboarding_needs_api_key. Only complete_provider_picker_onboarding clears it, and that runs only while onboarding == Provider. The stale true kept the footer on "model not connected" (frame.rs info_segments) and left should_adopt_live_local_ollama armed against the provider the user had just chosen. Fix: on a successful switch, set onboarding_needs_api_key from has_api_key(config) for the new route, and clear onboarding_missing_key_recovery when the key is present. The auth-failure rollback already restores the previous flag. Tests (focused, local): - new first_run_switch_to_keyed_route_clears_launch_missing_key_state (real Config::load + App::new startup with the inherited env pair, Esc, switch_provider): passes; fails without the fix at the !onboarding_needs_api_key assert. - provider_switch|switch_provider|local_ollama|missing_key|onboarding: 109 passed, 0 failed. - first_run_|provider_picker: 161 passed, 1 failed. The one failure, first_run_ollama_choice_survives_restart_from_canonical_config ("index not found"), also fails on a clean aaaadfe tree and is not touched here. - first_run_route_env_guards now also removes CODEWHALE_PROVIDER. E2E (tmux 130x40, debug build, fixture at 127.0.0.1:4880): - configured home, env clean: composer, footer "OpenAI-compatible · gpui-fixture · thinking: max". - configured home with the inherited deepseek env: picker (correct per precedence), then Esc and /provider openai: footer "OpenAI-compatible · gpt-5.6 ..." (release build: "model not connected"). - unconfigured home with local Ollama unreachable: "Choose your model provider" picker. With the live local Ollama it auto-adopts ollama/qwen3:4b, same as the release build. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…6693 Hosted CI on 776e3fe (PR #6672) failed two codewhale-tui tests. 1. client::catalog_tests::raw_rows_keep_duplicate_known_fields_invalid_in_existing_parsers Cause: #6691 (79d9815) made OpenRouter parsing per-row by decoding `data` as Vec<serde_json::Value>. A Value map keeps the last of a duplicated key, so `{"id":"second","id":"replacement"}` was silently accepted as "replacement" (and a duplicated pricing.prompt took the last rate) instead of being rejected. Fix (product): keep rows as Box<RawValue> and decode each with from_str, so serde's duplicate-field error still fires and the row is skipped and counted as malformed like any other undecodable row. The test now asserts the per-row contract: the ambiguous row is dropped, "first" survives; Baseten still rejects the whole response. 2. tui::ui::tests::first_run_ollama_choice_survives_restart_from_canonical_config Cause: #6693 (1394c29, eb48db7) moves a root default_text_model onto the leaf of the outgoing route that was resolving it. The test seeded default_text_model = "deepseek-v4-pro" with DeepSeek active and no leaf, then asserted the root key survived the switch to Ollama ("index not found"). That encoded superseded behavior. Fix (test): assert the root alias is gone and [providers.deepseek].model = "deepseek-v4-pro", which is the #6693 contract (switching back keeps the choice; Ollama never inherits it). Tests (CARGO_BUILD_JOBS=4, cargo test -p codewhale-tui --lib): - raw_rows_keep_duplicate_known_fields_invalid_in_existing_parsers: test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 13814 filtered out - first_run_ollama_choice_survives_restart_from_canonical_config: test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 13814 filtered out - filters catalog openrouter fetch_catalog first_run_ provider_picker local_ollama: test result: ok. 450 passed; 0 failed; 2 ignored; 0 measured; 13363 filtered out (a first run of the same filters had 1 failure in provider_picker::tests::credential_draft_is_masked_and_escape_drops_it_without_persistence, which passed 3/3 alone and in the rerun; it counts files in a home dir and is untouched by this change) Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Test (windows-latest) failed runtime_surface_review_documented_ssh_host_loads
("documented SSH worker example"): Windows checkouts give
docs/zh_hans/FLEET.md CRLF line endings, so splitting on "```json\n"
found no block. The test now normalizes line endings before parsing.
Checks: cargo test -p codewhale-tui --lib -- fleet::host::tests:
21 passed; 0 failed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Integration PR for v0.10.1. It merges the ready PRs together so CI runs once on the combination, instead of each PR re-running after every CHANGELOG conflict. Merging this with a merge commit makes each PR's head reachable from
main, so GitHub marks them merged.Included
Round 1 (each had its required checks green and no unaddressed review comments):
Round 2, in merge order:
Contributor PRs #6664 (@gaord) and #6666 (@aboimpinto) are merged as they are. Their commits and authorship are unchanged.
Hand-resolved conflicts
Each merge commit message gives the full reasoning.
thread.id.thread_session_id(id)returns the thread id.save_current_sessionkeeps fix(sessions): stop orphaning sessions at their writers and repair existing orphans (#6144) #6640's guards and fix(runtime): threads own the restore points on their turns; undo restores or refuses (#6621) #6645's choice of target document.RuntimeApiStategets both new fields,shutdown(forstream.end) andgit_writes(serialized guarded git writes). Both are set in both constructors.RuntimeProcessOwnerLock::acquirekeeps fix(sessions): stop orphaning sessions at their writers and repair existing orphans (#6144) #6640's retry loop for a lock file moved away mid-open. Both "held" exits now return fix(trust): credentials masked at rest, honest approval timeouts, workspace trust #6601's typedWouldBlockerror. The credential scrub uses that error to skip a busy store instead of failing.(Timeout, decided_by: None)and returnsApprovalResult::TimedOut. Both snapshot re-exports are kept.*whose decoded word prefix is empty is refused. The oldhas_leading_globfunction is removed.shell_command_grant_scope. A family grant needs both rules to allow it.read_only_git_commandwith fix: harden project config, repository reads, fetch refspecs and downloads #6679'sREVIEW_DIFF_ARGS. Both new install tests are kept.## 📋 …heading and the pinned-anchors preamble. fix: keep undo, resume and requirements honest; bound stream lines #6680's new test now iterates the legacy markers.codewhale_core::secret_eq::constant_time_eq.secured_assettakesimpl IntoResponse.SUMMARY_HEADERis kept and madepub(crate)for fix(tui): a fork continues when a turn lost its tool call #6664's recovery tests.Semantic fixes after merging (commits 983bcf2, 68a8c93)
SnapshotRepo::changed_paths_between(fix(runtime): threads own the restore points on their turns; undo restores or refuses (#6621) #6645/fix(tui): scope /undo to the paths the undone step changed (#6644) #6682 vs feat(receipts): list what a session did, from the records it already keeps #6591). The receipts version is renamedpath_changes_between.TimedOut. It now refunds the tool-call budget on every refusal: timeout, denial, policy retry or error.ApiError.codewas missing at fix(sessions): stop orphaning sessions at their writers and repair existing orphans (#6144) #6640's 409. Also fixed fix(tui): start model-run child processes from the scrubbed environment #6671'srun_gatetest call.install.rsneedscrate::utils. The capped body reader moved toutils/response_body.rs, which the harness includes.&&,||and;(fix(fleet): read-only agents run read-only commands; refusals are actionable (#6015) #6637 × fix: harden project config, repository reads, fetch refspecs and downloads #6679). Before,pwd && git diffran without filter overrides.run_git_review_command) is removed.crate::tui::historymoved into atui::historytest. This brings the runtime→UI ratchet back from 41 to 40.runtime_api/git.rsgoes from 5 to 6. runtime_api: optional preconditions on git stage/unstage/discard/commit #6648's PR body justified this category: fingerprint reads in sync functions reached only fromspawn_blocking. Its own branch already has 6 such sites.subagent/worktree.rstightens from 7 to 6.Local evidence on 68a8c93
cargo check --workspace --tests: clean, no warnings.cargo test -p codewhale-execpolicy: 221 + 1 + 5 + 7 + 1 passed, 0 failed.cargo test -p codewhale-tui --test integration: 187 passed, 0 failed.cargo test -p codewhale-tui --lib -- runtime_api:: runtime_threads::: 654 passed, 0 failed on the final run. Earlier runs under load each had 2–4 different timing-deadline failures, and each of those passed when run alone.codewhale-tui --libsets for session, compaction, approval, readonly, receipts, snapshot, verify, debug, shell and git:runtime_chat_relaydeadline). Theruntime_chat_relay::rerun passed 17/17, 3 times.core::engine::approval::rerun passed 15/15, twice.check-command-crate-boundaries.py,check-blocking-calls-budget.py,check-dead-code-budget.py(260, at budget),split/module_graph.py --checkandcargo fmt --checkall pass.sync-changelog.sh --check: up to date.check-versions.sh --range-audit-advisory: OK.npm test: 636 passed, 0 failed (68 + 16 + 50 node tests, 502 web).npm run check:web: passed.Not run locally: the full workspace test suite. CI covers it.
🤖 Generated with Claude Code