Skip to content

v0.10.1 integration: land the ready PRs together - #6672

Merged
Hmbown merged 227 commits into
mainfrom
integrate/0.10.1
Sep 28, 2026
Merged

Hmbown merged 227 commits into
mainfrom
integrate/0.10.1

Conversation

@Hmbown

@Hmbown Hmbown commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Integration PR for v0.10.1. It merges the ready PRs together so CI runs once on the combination, instead of each PR re-running after every CHANGELOG conflict. Merging this with a merge commit makes each PR's head reachable from main, so GitHub marks them merged.

Included

Round 1 (each had its required checks green and no unaddressed review comments):

Round 2, in merge order:

Contributor PRs #6664 (@gaord) and #6666 (@aboimpinto) are merged as they are. Their commits and authorship are unchanged.

Hand-resolved conflicts

Each merge commit message gives the full reasoning.

Semantic fixes after merging (commits 983bcf2, 68a8c93)

Local evidence on 68a8c93

  • cargo check --workspace --tests: clean, no warnings.
  • cargo test -p codewhale-execpolicy: 221 + 1 + 5 + 7 + 1 passed, 0 failed.
  • cargo test -p codewhale-tui --test integration: 187 passed, 0 failed.
  • cargo test -p codewhale-tui --lib -- runtime_api:: runtime_threads::: 654 passed, 0 failed on the final run. Earlier runs under load each had 2–4 different timing-deadline failures, and each of those passed when run alone.
  • Focused codewhale-tui --lib sets for session, compaction, approval, readonly, receipts, snapshot, verify, debug, shell and git:
    • First run: 2331 passed, 1 failed (a runtime_chat_relay deadline). The runtime_chat_relay:: rerun passed 17/17, 3 times.
    • After the review fixes: 1133 passed, 1 failed (an approval event deadline). The core::engine::approval:: rerun passed 15/15, twice.
  • check-command-crate-boundaries.py, check-blocking-calls-budget.py, check-dead-code-budget.py (260, at budget), split/module_graph.py --check and cargo fmt --check all pass.
  • sync-changelog.sh --check: up to date.
  • check-versions.sh --range-audit-advisory: OK.
  • npm test: 636 passed, 0 failed (68 + 16 + 50 node tests, 502 web). npm run check:web: passed.
  • Codex ran a read-only review of every hand-resolved merge. It reported 5 findings (1 high, 2 medium, 2 low). All 5 were checked against the code and fixed in 68a8c93.

Not run locally: the full workspace test suite. CI covers it.

🤖 Generated with Claude Code

CodeWhale Bot and others added 30 commits September 25, 2026 05:37
The Runtime API bearer / x-codewhale-runtime-token / legacy
x-deepseek-runtime-token checks and the computer-display owner checks
used `==`, and constant_time_eq existed in three private copies
(runtime_api/mobile.rs, runtime_api/web.rs, app-server). One shared
codewhale_core::secret_eq::constant_time_eq (the length-hiding
app-server variant) now serves every credential comparison; the copies
are deleted.

Tests: codewhale-core secret_eq 2 passed / 0 failed (new);
codewhale-tui runtime_api::{auth,web,mobile} 11 passed / 0 failed;
codewhale-app-server lib test build compiles.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
Logout deletes every provider API key, OAuth login, the account session
and the Daytona token at once, with no prompt. It now asks for an
explicit `yes`/`y` on a terminal and refuses (deleting nothing) when
stdin is not a terminal unless `--yes`/`-y` is passed.

Session revocation on logout is unchanged (not in this slice).

Tests: codewhale-cli logout filter 8 passed / 0 failed, including new
logout_confirmation_accepts_only_an_explicit_yes and
logout_parses_yes_flag.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…and interact calls

The "approve for session" key for a shell command was its arity-dictionary
prefix, falling back to the first word, so a grant for `cd` or `bash` also
covered `cd x && rm -rf ~` or `bash -c ...`, `rm tmp/x` covered `rm -rf ~`,
and every exec_*_interact / wait call shared one `shell:<empty>` key.

Now only a simple, dictionary-known, non-wrapper command keeps its family
grant (`git status` still covers `git status -s`, not `git push`).
Compound commands (chaining, pipes, redirects, substitution, `$`),
wrappers/interpreters (bash/sh/env/sudo/xargs/python/node/find/ssh,
docker run/exec, uv run, ...), env-assignment prefixes and unknown
commands are keyed by the full normalized command (whitespace collapsed
only outside quotes). Interact and wait calls are keyed by the exact call.
Documented in RUNTIME_API.md.

Tests: codewhale-tui approval_cache 26 passed / 0 failed, including new
shell_grants_fail_closed_on_compound_wrapper_and_unknown_commands and
shell_interact_grants_are_per_exact_call.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
A repo-level `.codewhale/config.toml` (or legacy `.deepseek/`) could set
`notes_path = "~/.zshrc"`, and the auto-approved `note` tool would then
append model text to the user's shell rc file.

- Project-scope `notes_path` is honoured only as a relative path inside
  the workspace (resolved against it); `~`, `$`, absolute, rooted and
  `..` paths are ignored with a warning. User-config `notes_path` is
  unchanged.
- The note tool refuses a notes target that is a symlink or not a
  regular file, and one placed in the workspace whose existing directory
  resolves outside it after following symlinks (a committed
  `notes.md -> ~/.zshrc` or `notes/ -> ~`).

Tests: codewhale-tui project_config_tests 22 passed / 0 failed (new
project_overlay_keeps_notes_path_inside_the_workspace);
tools::shell::tests::note_tool_refuses_symlinked_targets_that_leave_the_workspace
1 passed / 0 failed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
… budget slot

An approval card that expired unanswered reached the model as
"Tool 'x' denied by user", and the Runtime API path (GPUI/web) called
deny_tool_call on its timeout, recording the user's denial. Denied,
timed-out, cancelled and unavailable calls also kept their per-turn
tool-call budget slot although they never ran.

- ApprovalResult gains TimedOut; the engine tells the model the request
  timed out and the user did not deny it (ExecutionFailed, not the
  PermissionDenied "denied by user" marker), and audits decision=timeout.
- runtime_threads' approval timeout sends deny_tool_call_timed_out, so
  the receipt and the model both say timeout (the approval.decided event
  shape for clients is unchanged).
- Any approval outcome that stops the call (denied, timed out,
  cancelled/unavailable) refunds its ToolCallBudget slot (#5170 already
  refunds planning-time gates).

Tests: codewhale-tui approval filter 31 passed / 0 failed and
tool_call_budget 4 passed / 0 failed, including new
core::engine::tests::approval_timeout_is_reported_as_timeout_and_refunds_the_budget
and the updated runtime_threads approval_timeout_denies_clears_ui_and_next_turn_can_start.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
Repository skill roots (`.agents/skills`, `skills/`, `.opencode`,
`.claude`, `.cursor`, `.codewhale/skills` under the workspace) were read
into runtime discovery without workspace trust, so a cloned repo's
SKILL.md reached the model's skill list (and shadowed a same-named global
skill) before the user ever ran /trust.

Runtime discovery now skips project-scope roots until
is_workspace_trusted(workspace) — the same predicate hooks and MCP use —
checked lazily, only when such a root exists. Skipping is not silent:
discovery adds a warning naming the skipped directories and `/trust`,
visible in /skills and the skills block. Once trusted, project skills
keep their precedence and a shadowed global skill keeps the existing
"is shadowed by" warning. Audit/mutation catalogs are unchanged.

Tests: codewhale-tui skills::tests::project_skills_require_workspace_trust
1 passed / 0 failed (new). The wider skill/plugin sweep did not run: the
shared build volume filled (ENOSPC) before it could link.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…ted built-ins

Markdown commands from a repository's `.codewhale/`, `.deepseek/`,
`.claude/` or `.cursor/commands` (and `.codewhale/workflows`) loaded in any
workspace and were dispatched ahead of the built-ins, so a cloned repo's
`trust.md` or `undo.md` answered `/trust` or `/undo` with its own prompt,
sent to the model as the user's own message, before the user ever ran
/trust.

- commands_dirs/workflow_dirs include workspace directories only when
  is_workspace_trusted(workspace) holds (the predicate hooks, MCP and
  project skills use). User-global commands are unchanged.
- A workspace command (or alias) never stands in for a protected built-in:
  the ones that grant or revoke authority, hold credentials, or undo and
  discard work (/trust, /undo, /permissions, /mode, /config, /login,
  /logout, /restore, /hooks, /mcp, ... and the /jihua, /zidong mode
  aliases). The definition is skipped with a load error naming the rename;
  other built-ins (/help, /review) stay shadowable as FEAT-011/012 specify.
- Test fixtures that exercise workspace commands and project skills now
  trust their temp workspace (test_support::trust_workspace), which also
  repairs the fixtures broken by the project-skill trust gate (294b616).
- docs/architecture/command-dispatch.md states the gate.

Tests: codewhale-tui focused run (user_registry, user_commands,
epic_dispatch/discovery, command_palette, skills::, slash_completion,
views::help, command_catalog, skill_lifecycle, commands::tests,
tools::skill, codemode, cached_skills, prompts::tests, plus the B1
filters) 638 passed / 1 failed; the one failure was the new runtime B1
fixture in the next commit, fixed and rerun 1/0. New:
workspace_command_never_replaces_a_protected_builtin,
workspace_alias_never_replaces_a_protected_builtin,
untrusted_workspace_commands_do_not_load.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
… and scrub for stored secrets

Session JSON stored tool output verbatim: redaction ran only when a model
request was built (client.rs), so a `cat ~/.codex/auth.json` or a printed
bearer token was written to ~/.codewhale/sessions live (sweep B1).

- The engine masks credentials in ToolResult content (and text
  content_blocks) once, in add_session_message, with the same masking the
  request boundary applies. The transcript is what session files,
  checkpoints and later requests are built from. The double-confirmed
  `[redaction] model_bound = "disabled"` opt-out is honored the same way.
- Runtime API threads mask tool output before it reaches the durable tool
  item and the event log built from it (always on: receipts fan out to UI
  clients).
- `codewhale doctor` gains a "Stored Sessions" finding: it scans the newest
  50 session/checkpoint files and says how many still hold credentials in
  stored tool output. It never rewrites anything.
- `codewhale sessions scrub-secrets` reports every affected file;
  `--apply` rewrites them atomically, touching only tool_result text.
  Documented under docs/CONFIGURATION.md "Stored sessions", with the
  reminder that masking a copy does not un-leak a credential: rotate it.

Not covered: runtime item/event files written by older builds are not
scrubbed by the command (it walks the sessions store only).

Tests: codewhale-tui
core::engine::tests::tool_output_credentials_are_redacted_when_they_enter_the_transcript,
session_secret_scrub (2), runtime_threads::tests::runtime_tool_items_store_tool_output_with_credentials_masked
all pass (in the 638/1 focused run above; the runtime fixture failed once
on a single-line JSON fixture and passes 1/0 after the fix);
runtime_threads:: 226 passed / 0 failed / 2 ignored.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
The note tool wrote through tokio's File and returned without flushing, so
the append could still be in flight on a blocking thread when the tool
reported "Note appended". Under parallel load the notes-symlink test read
the file before the bytes landed (2 of 4 runs of the approval/notes
filter). Flush before reporting success.

Tests: codewhale-tui filter `approval tool_call_budget constant_time
logout note project_overlay auth::` 483 passed / 1 failed, twice; the
note test passes in both. The one failure is
task_manager::tests::pending_approval_suspends_idle_and_timeout_denial_settles_failed,
a pre-existing race: runtime_threads tests set the process-global
TEST_APPROVAL_DECISION_TIMEOUT_MS that this test's drive reads; it passes
alone (1/0).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
Bumps [actions/setup-python](https://github.com/actions/setup-python) from 6 to 7.
- [Release notes](https://github.com/actions/setup-python/releases)
- [Commits](actions/setup-python@v6...v7)

---
updated-dependencies:
- dependency-name: actions/setup-python
  dependency-version: '7'
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
…ession lock

Review findings on 000ab6b:

- Runtime API tool items copied the tool's metadata unmasked. exec_shell
  metadata carries summary/stdout_summary/stderr_summary (first stdout
  lines), so `cat auth.json` on a desktop or web thread still stored and
  emitted the token. Metadata now goes through a model-bound JSON
  redactor (new redact_model_bound_json_secrets); the test uses real
  shell-shaped metadata.
- The redactor missed credentials in compact JSON and query strings
  ({"tokens":{"access_token":"eyJ…"}}, jq -c, curl, ?access_token=…):
  the whole document is one word whose first key is harmless. Such words
  are now split on their structure and each member gets the same keyed
  and bare-token checks, delimiters byte-exact; already-masked values
  (`token=***`) are left alone. doctor no longer reports a false clean.
- `sessions scrub-secrets --apply` re-reads and rewrites each file under
  the per-session lock every save takes, so it cannot lose a concurrent
  save. Invalid ids are left untouched.
- docs: say what is still stored unmasked (spillover files, shell
  evidence artifacts, tool-call inputs).

Tests: codewhale-secrets 77/0 (1 ignored); codewhale-config lib 695/0
(1 ignored, incl. new compact_json_and_query_strings_mask_their_credentials);
codewhale-tui lib focused (runtime_tool_items_store session_secret_scrub
approval_cache user_registry memory::note redact export) 252/0 (1 ignored).
Wider tui set (user_commands skills:: project_config_tests note
tool_call_budget core::engine::tests command_palette epic_
runtime_threads::tests tools::shell client::tests) 1379/2: the two
core::engine::tests::mcp_boot_* failures reproduce identically with these
edits stashed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…pace; remote commands protected

Review findings on ecdc840, d0e3158 and 9253eff:

- Shell grants ignored flags, so `git status` covered
  `git -ccore.fsmonitor=./evil status`, `cargo build` covered
  `cargo build --config build.rustc-wrapper=…`, and `git config user.name`
  covered `git config core.fsmonitor`. Any `=` argument, a code-running
  option (-c, -C, --config, --exec, -f, --file, -e, …) or a config-style
  family (config/set/remote) now keys the grant on the full command.
- `/note clear|edit|remove|list` followed a committed
  `notes.md -> ~/.zshrc` or a symlinked notes directory. They now apply
  the note tool's rule: no symlinked file, directory must resolve inside
  the workspace.
- rc/remote-control, relay, remote-env, profile and share are protected
  built-ins: a trusted repo's rc.md can no longer answer `/rc off`.

Tests: codewhale-tui lib focused (approval_cache user_registry
memory::note ...) 252/0 (1 ignored), run together with the previous
commit's set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
Resolves crates/secrets/src/redact.rs against main's
redact_json_model_bound_secrets (auto-review, 2c672a9): the branch's
own model-bound JSON helper is dropped and Runtime API tool metadata now
uses main's, so there is one JSON redactor per policy. The structured
compact-JSON word pass is kept.

Tests after the merge: codewhale-secrets 77/0 (1 ignored),
codewhale-config lib 696/0 (1 ignored), codewhale-tui lib focused
(runtime_tool_items_store session_secret_scrub approval_cache
user_registry user_commands memory::note redact export auto_review
skills:: project_config_tests tool_call_budget
core::engine::tests::approval command_palette epic_ runtime_threads::tests
tools::shell client::tests) 1242/0 (3 ignored).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
turn_loop.rs conflicted: main (2b440fa, #6584) removed the
questions_allowed parameter, and this branch added tool_call_budget beside
it. Kept tool_call_budget and dropped questions_allowed.

Checks: cargo check -p codewhale-tui --tests clean; focused
approval|turn_loop|redact: 462 passed, 1 failed. The failure is
task_manager pending_approval_suspends_idle_and_timeout_denial_settles_failed,
which is unchanged from main and passes alone (120 ms idle window under
parallel load).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…s/setup-python-7

chore(deps): bump actions/setup-python from 6 to 7
Merge main at 24a7b99 (runtime split #6586, code mode MCP default #6583).

Conflicts:
- lib.rs: keep this branch's `mod session_secret_scrub`; `session_tree`
  now comes from codewhale_runtime.
- turn_loop.rs: execute_planned_tools takes main's NestedGateEnv, which
  already carries the tool-call budget; this branch's approval refunds
  go through nested_gate_env.tool_call_budget.
- main's new code-mode nested approval gate did not handle this branch's
  ApprovalResult::TimedOut (E0004). A timed-out nested call now reports
  NestedDecision::TimedOut with the same "did not run, user did not deny
  it" error as a direct call, and every refused nested approval hands its
  budget slot back like the direct path.

CI fixes:
- Test (ubuntu/windows): `codewhale doctor` created CODEWHALE_HOME via
  SessionManager::default_location() in the stored-secrets scan. It now
  resolves the sessions dir with the read-path resolver
  (resolve_state_dir), so the diagnostic stays read-only.
- Test: apply_slash_menu_selection_honors_user_argument_metadata_and_builtin_override
  (new on main) writes workspace commands without trusting the
  workspace; this PR intentionally loads workspace commands only in a
  trusted workspace, so the test now trusts its temp workspace.
- Lint blocking-calls: session_secret_scrub.rs is a synchronous module
  (doctor calls it inside spawn_blocking; `sessions scrub-secrets` is a
  one-shot sync subcommand like `sessions list`); budget raised by its 2
  std::fs sites.
- CodeQL cleartext-logging: the printed count/path fields were named
  *_with_secrets; renamed to flagged_tool_results / flagged_files. They
  never held secret values.

Local: clippy -p codewhale-tui --all-targets --all-features (CI flags)
clean; cargo fmt --check, git diff --check clean; codewhale-tui lib
focused 93+1577 passed, 1 failed (task_manager
pending_approval_suspends_idle_and_timeout_denial_settles_failed:
unchanged from main, passes 3/3 alone, wall-clock flake under load);
integration diagnostic_read_only 13/13; codewhale-cli
diagnostic_dispatch_read_only 1/1; blocking-calls, dead-code,
command-crate-boundaries, module_graph --check, provider-registry,
command-migration-manifest, reqwest-builders, locale checks pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R9rJEoMuUSRWznU6QjkE7h
…ice B)

The offline seed now carries MiMo 2.6's canonical facts under upstream
Models.dev keys (xiaomi/mimo-v2.6-pro, xiaomi/mimo-v2.6-flash), projected
unedited from https://models.dev/catalog.json (2026-09-25) onto the fields
ModelsDevModel reads.

Slice A normalized a namespaced key's vendor only on live refresh, so the
same entry offline surfaced as a non-route `xiaomi` provider. The vendor
namespace is never a CodeWhale id, so it now normalizes in both modes, and
gap-filling compares on the normalized provider identity so a verbatim
bundled provider row still shadows its namespaced twin.

Kept on purpose: the providers.xiaomi-mimo 2.6 rows (they still win and
carry the verified text-only qualification; retiring them would widen
MiMo to image/audio/video input and drop the bare-id rows that
reasoning_support, provider_model and the fleet label read), and the bare
deepseek-v4-pro/-flash keys (every hosted DeepSeek base_model, the route
layer and pricing name those ids; renaming them is slice C).

No write-capable generator exists: scripts/catalog_models_dev.py fails
closed on --write by design, so the entries were staged per
docs/CATALOG_REFRESH.md and validated with its snapshot --check and drift.

Tests: codewhale-config --lib 698 passed, 0 failed, 1 ignored (catalog
filter: 54 passed incl. 2 new); codewhale-models 44 passed;
scripts/catalog_models_dev_test.py 8 passed; snapshot --check ok;
cargo fmt --check, blocking-calls, dead-code, command-crate-boundaries
and module_graph --check all pass.

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
These changes move to a separate change. The protected built-in
command list from the same commit stays in this PR.

Reverts the approval_cache.rs and commands/groups/memory/note.rs
hunks of a8755ae.

Tests: none run for this commit alone; the focused set runs on the
final head of this PR (see the fix(ci) commit that follows).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This change moves to a separate change.

This reverts commit d0e3158 (lib.rs, tools/shell.rs and
tools/shell/tests.rs). The note-flush fix from 133a8a2 stays.

Tests: none run for this commit alone; the focused set runs on the
final head of this PR (see the fix(ci) commit that follows).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This change moves to a separate change.

This reverts commit ecdc840 (tools/approval_cache.rs and
docs/RUNTIME_API.md).

Tests: none run for this commit alone; the focused set runs on the
final head of this PR (see the fix(ci) commit that follows).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…-only scan, synthetic JWT from parts

- The runtime-contract fixture puts its representative skill in the
  workspace's project skills directory. Project skills now load only in a
  trusted workspace, so the fixture trusts its workspace and keeps
  measuring a loaded skill. The budget is unchanged from main.
- `sessions scrub-secrets` takes its printed counts from the read-only
  scan. The rewrite under the per-session lock re-reads the file and
  returns only success, so no value flows back from the lock call into
  the report.
- The redaction test's synthetic, unsigned JWT is assembled from its three
  segments at runtime.

Tests (local, macOS):
- codewhale-config lib `persistence`: 28 passed / 0 failed.
- codewhale-tui lib focused: approval_cache 24/0, session_secret_scrub 2/0,
  user_registry 31/0, user_commands 26/0, memory::note 14/0,
  tools::shell 160/0 (1 ignored), skills:: 225/0, project_config_tests
  33/0, tool_call_budget 4/0, core::engine::tests::approval 2/0,
  runtime_threads::tests::approval 8/0 with --test-threads=1.
  In one shared process, approval_wait_heartbeat_is_never_sequenced_after_the_decision
  fails when it runs beside approval_timeout_denies_clears_ui_and_next_turn_can_start:
  both tests use the process-global test approval timeout. Both tests
  are on main unchanged, and CI's nextest runs each test in its own process.
- scripts/check-runtime-contract-budget.py: PASS, all 55 metrics at budget.
- cargo clippy -p codewhale-config -p codewhale-tui --all-targets
  --all-features --locked with CI's flags: clean. cargo fmt --check: clean.
- check-blocking-calls-budget, check-dead-code-budget,
  check-command-crate-boundaries, split/module_graph.py --check: pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
)

The offline seed carried Codewhale's deliberate holds as hand edits: prices
left out where a flat rate would mislead, DeepSeek's output limit kept at its
published 384K. A live Models.dev refresh replaces seed rows wholesale, so
every hold vanished on a networked install: live rows priced MiMo and
DeepSeek, listed Alibaba plan rows at $0 (read as free), and raised DeepSeek's
output limit to 393216.

The holds now live in `crates/config/assets/catalog_corrections.json` as field
patches in the cloud-facts `ModelFact` shape. The signed layer's patch code
applies them to every Models.dev row as it is hydrated, bundled seed and live
refresh alike. So a correction ranks above both Models.dev layers and below
signed cloud facts (which may still correct it) and provider rosters.

- `ModelFact.pricing_withheld` (optional reason) clears the row's price so the
  route reports `PricingSku::UnknownOrStale`. It is an optional field;
  payloads without it parse unchanged.
- `catalog_patch::apply_patches` is shared: the signed layer calls it with
  row materialization, corrections without, so they never add a row. The
  loader also refuses hide, deprecate, `allow_unlisted`, labels, a patch that
  changes nothing, and any limit patch without a reason.
- A corrected row keeps its own `source`. `CatalogCompiler::with_live` and
  other code sort rows by source, and a relabelled live row was placed above
  signed facts (an existing cloud-facts test caught it). The price a
  correction owns reports `CatalogSource::CodewhaleBundled` through
  `cost_source`.
- The declared-but-unused `CatalogCompiler.codewhale_bundled` field is
  removed; the layer docs (catalog.rs, CATALOG_REFRESH.md, CLOUD_FACTS.md)
  now describe the rank where corrections actually apply.

Each hold was checked again on its merits:
- kept: DeepSeek native pricing (time-aware table in tui pricing.rs), xAI
  (rates double past 200K), MiMo (PAYG and Token Plan keys look the same),
  the four Model Studio plan providers (quota), StepFun (plan and PAYG share
  ids), DeepSeek native output 384000 (published 384K; which value the API
  accepts is unverified, and the lower bound is always accepted);
- added: MiniMax-M3 (upstream price tiers at 512K, same rule as Grok; the
  seed already listed it unpriced without saying why);
- dropped: aggregator-hosted DeepSeek (an aggregator's published rate is its
  own price, as it is for every other aggregator row we price), OpenRouter
  dots :free (a free model is free), zai GLM-5.3 (the zai route is the
  pay-as-you-go API, not the Coding Plan the hold cited), MiMo's text-only
  image hold (made harmless by the previous commit) and MiMo's 1,000,000
  context (the only citation was our own rows).

Tests (focused):
- codewhale-config --lib: 706 passed, 0 failed, 1 ignored;
  --test configured_models: 9 passed
- codewhale-tui `-- vision image badge capabilit pricing provider_lake
  models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi
  mimo modelstudio grok stepfun minimax`: 872 passed, 0 failed, 2 ignored
  (one earlier run failed
  route_budget::uncatalogued_remote_model_keeps_a_conservative_ceiling; it
  passes alone and on rerun, and reads output-limit env vars without the
  env lock)
- cargo fmt --check clean; clippy -p codewhale-config -p codewhale-tui
  --all-targets --all-features with CI flags clean; budget, boundary and
  module-graph gates pass; changelog chain passes (public-copy 6/6)

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
# Conflicts:
#	crates/config/src/catalog/tests.rs
…6396)

Regenerating the offline seed from Models.dev (#6612) showed a second kind
of offline-only hold. The seed carried the reasoning controls Codewhale
relies on, but live Models.dev rows replace them:
- Grok: xAI documents `high` as the default effort (6be755f / #6501).
  Upstream lists the ladder with no default.
- MiniMax: M2.7 always thinks; M3 is adaptive or disabled, with a different
  default per dialect (47163da). Upstream lists a toggle or nothing.
- qwen3.8-max: thinking-only (the owner's console, ec4a5ae). Upstream
  lists a toggle.
- Muse Spark: also takes the `none` and `ultra` efforts (a50b653).
- Step 3.5: base model uses adaptive reasoning (StepFun guide, recorded
  2026-09-19). Upstream lists low/high.

The picker, `/effort` and the effort clamp read these, so today a networked
install loses them.

`ModelFact.reasoning_options` (optional) replaces a row's reasoning
controls and keeps any signed annotation already on the row. The 18 rows
above get it as a correction, each with its reason. GLM-5.3's high/max
ladder was left out: the seed only inherited it from GLM-5.2 as a
placeholder, and upstream now publishes low/high/max for 5.3. The Coding
Plan qwen `budget_tokens` entries were left out too: they were copies of
Token Plan rows, not a Codewhale decision.

Tests (focused):
- codewhale-config --lib: 707 passed, 0 failed, 1 ignored;
  --test configured_models: 9 passed
- codewhale-tui `-- vision image badge capabilit pricing provider_lake
  models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi
  mimo modelstudio grok stepfun minimax effort reasoning thinking muse`:
  1169 passed, 0 failed, 2 ignored
- cargo fmt --check clean; clippy -p codewhale-config --all-targets
  --all-features with CI flags clean; changelog chain passes
  (public-copy 6/6)

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first commit dropped the aggregator-hosted DeepSeek price hold on the
grounds that an aggregator's published rate is its own price. Running the
TUI pricing tests against a seed generated from Models.dev (#6612) showed
the hold's real purpose: `crates/tui/src/pricing.rs` prices these routes
from a reviewed provider-docs table, and a catalog rate outranks that table.
For Fireworks V4 Pro, for example, the table has 1.74 per 1M input and
Models.dev has 1.20. The hold is restored for the nine hosted DeepSeek rows,
with that reason.

Also from the generated seed:
- Novita DeepSeek rows: max_output 384000. Models.dev lists 393216; this is
  the same published-384K question as native DeepSeek.
- grok-4.3: reasoning_options []. xAI documents no effort control for it,
  so Codewhale sends none (#6501); Models.dev lists a ladder.

All three are no-ops against the current hand seed. They take effect on
live refresh today, and offline once the seed is generated.

Tests (focused): codewhale-config --lib 707 passed, 0 failed, 1 ignored.
Changelog chain: public-copy 6/6.

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CodeWhale Bot and others added 22 commits September 27, 2026 07:16
Cause: the 80-column footer reserved only 24 columns for route identity.
gpt-5.6 with thinking: high fit, but xhigh and longer requested-to-effective
labels exceeded the budget, so the entire effort field disappeared.

Let the route budget floor grow from 24 to 32 columns with terminal width.
This preserves all reasoning labels at 80 columns while retaining context
and performance readings on narrower rows. Add a rendered-footer regression
covering all nine ReasoningEffort variants and sync the Unreleased fix note.

Regression proof with the original budget (Minimal, XHigh, Ultra missing):
CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true cargo test -p codewhale-tui --lib -- footer_keeps_reasoning_label_for_every_effort_tier
test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 13774 filtered out; finished in 0.27s

Final focused validation, including regression, narrow layouts and hitboxes:
CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true cargo test -p codewhale-tui --lib -- tui::ui::frame::
test result: ok. 30 passed; 0 failed; 0 ignored; 0 measured; 13745 filtered out; finished in 0.67s

CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true cargo fmt --all
sh scripts/sync-changelog.sh
sh scripts/sync-changelog.sh --check
git diff --check

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Conflict resolution:
- memory/note.rs: kept #6601's lowercased command dispatch,
  ensure_notes_target_in_workspace pre-check and notes_path resolver, and
  threaded #6685's workspace root through every subcommand. read_notes stays
  pub(crate) (dock NOTES view) but now takes the workspace and reads through
  fs_confined::read_to_string; tui/workspace_context.rs passes the workspace.
  Restored the std::fs import the pre-check needs.
- skills/install.rs: test-module add/add; kept #6679's download_with_cap
  server and test alongside #6685's registry-name tests (braces verified).
- CHANGELOG.md: kept both Security bullets; crates/tui/CHANGELOG.md
  regenerated with scripts/sync-changelog.sh.

Verification: cargo check -p codewhale-tui --tests clean.
cargo test -p codewhale-tui --lib -- memory::note anchor skills::install
fs_confined compaction::last_round project_context: 183 passed; 0 failed.
workspace_context: 14 passed; 0 failed. Codex read-only review: no findings.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Cause: App initialization could retain an options-default model despite a
resolved config route. Local Ollama adoption treated missing credentials as
permission to replace an explicit provider/model. The context meter counted
the startup system prompt against the fallback window before any conversation.

Use the configured startup model and retain whether the resolved config named
a route before runtime synchronization materializes defaults. Reject both
probing and late adoption for explicit routes, including missing-key recovery.
Show zero context usage until messages or provider usage exist; retain normal
pressure readings afterwards. Update and sync the Unreleased changelog.

Validation (all cargo commands used CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true):
cargo fmt --all
sh scripts/sync-changelog.sh
git diff --check

cargo test -p codewhale-tui --lib -- first_run_route_
Original production code with the new regression tests:
test result: FAILED. 1 passed; 3 failed; 0 ignored; 0 measured; 13773 filtered out; finished in 0.28s
Restored fix:
test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 13773 filtered out; finished in 0.23s

cargo test -p codewhale-tui --lib -- adopting_a_local_model
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 13776 filtered out; finished in 0.41s
cargo test -p codewhale-tui --lib -- local_ollama_probe_leaves
test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 13776 filtered out; finished in 0.11s
cargo test -p codewhale-tui --lib -- context_usage
test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 13771 filtered out; finished in 0.12s
cargo test -p codewhale-tui --lib -- startup_provider
test result: ok. 2 passed; 0 failed; 0 ignored; 0 measured; 13775 filtered out; finished in 0.08s

Local library-test evidence only; no packaged runtime, CI, or release claim.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…dpoints

Cause: #6687 treated any default_text_model as an explicit route, but the
generated first-launch config.toml writes default_text_model =
DEFAULT_TEXT_MODEL, so an unconfigured person who first launched without
Ollama lost local discovery on every later launch. Endpoint-only routes
(legacy top-level base_url, [providers.<active>].base_url, or
DEEPSEEK_BASE_URL/CODEWHALE_BASE_URL) with a missing key were still
replaceable by a live Ollama. The empty-session meter shortcut also zeroed
a submitted first turn before the engine mirrored messages back
(one_owner_tests::context_cap_warns_once_in_the_posture_bar), and with no
messages the reported prompt was dropped in favour of the system-prompt-only
estimate.

Fix: a template default_text_model (no provider, value DEFAULT_TEXT_MODEL)
is not a choice; explicit env provider/model overrides and a configured
active-route endpoint (new Config::active_route_endpoint_configured) are.
The meter treats a loading turn or a user history cell as a started
conversation, and with no messages lifts the estimate to the reported
prompt. model_picker fleet_case_distinct_pins test encoded the superseded
options-default startup model: the configured default_text_model is now the
startup route and leads the list, so the test asserts pin order by position.

test result (CARGO_BUILD_JOBS=4 cargo test -p codewhale-tui --lib --
context_cap_warns_once_in_the_posture_bar fleet_case_distinct_pins
first_run_route_ context_usage):
test result: ok. 16 passed; 0 failed; 0 ignored; 0 measured; 13783 filtered out; finished in 1.31s
Same filter with init.rs/frame.rs reverted to f45d1df:
test result: FAILED. 5 passed; 4 failed; 0 ignored; 0 measured; 13790 filtered out; finished in 0.40s
CARGO_BUILD_JOBS=4 cargo check --workspace --tests: Finished
sh scripts/sync-changelog.sh --check: crates/tui/CHANGELOG.md is up to date

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: runtime-contract ratchet failed on PR #6672: 6fae324 added a
53-byte refspec rule to the Git tool schema (plan/act/operate) and #6644
added 119 bytes of overwrite warning to revert_turn (act/operate).

Fix: drop the Git action description's duplicate of the fetch sentence
already in the tool description (the refspec rule stays, and the
validator's error names the fix); tighten revert_turn wording while
keeping the whole-workspace overwrite warning. Lock in the lower ceilings.

test result: python3 scripts/check-runtime-contract-budget.py
[runtime-contract-budget] PASS: all 55 metrics are exactly at budget.
(before --update: PASS: 55 ceilings respected; 6 can be tightened; plan -13 bytes, act/operate -33 bytes)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: Version drift failed in .github/scripts/release-workflows.test.js:
"docs/zh_hans/FLEET.md feeds Rust and must classify heavy". fad53fa
added include_str!("../../../../docs/zh_hans/FLEET.md") in
crates/tui/src/fleet/host.rs, but the ci.yml change filter still sent that
path to the docs-only light arm.

Fix: add it to the heavy arm next to the other embedded docs.

test result: node .github/scripts/release-workflows.test.js
Change detection OK: 134 Rust-consumed non-.rs paths classify heavy.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: CodeQL py/clear-text-logging-sensitive-data flagged die() in
scripts/catalog_models_dev.py because data returned by a function named
scrub_secrets() flowed into an error naming a provider/model row id. The
value is already-scrubbed public models.dev data, not a secret.

Fix: rename scrub_secrets -> strip_sensitive_fields; behaviour unchanged.

test result: python3 scripts/catalog_models_dev_test.py
Ran 16 tests ... OK

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
First-run recovery treated a fresh home as having no saved route, even when
config selected one. A usable route that skipped onboarding left no provider
step receipt, and setup treated saved-but-unchecked credentials as verified.

Recognize explicit startup routes in the recovery predicate. Record an
unchecked first-run route as configured through the existing setup owner;
reserve verified for observed success. Keep the unconfigured picker and local
Ollama discovery, prior setup decisions, and Fleet verification requirements.
Generalize the existing locked sidecar updater and migrate every caller so the
new startup receipt preserves privacy decisions and rejects corrupt/busy state.
Extend and synchronize the existing first-launch changelog bullet.

Validation (all cargo commands used CARGO_BUILD_JOBS=4 CARGO_NET_OFFLINE=true):
cargo fmt --all
sh scripts/sync-changelog.sh
sh scripts/sync-changelog.sh --check
git diff --check
cargo build -p codewhale-tui

cargo test -p codewhale-tui --lib -- first_run_
test result: ok. 18 passed; 0 failed; 0 ignored; 0 measured; 13786 filtered out; finished in 1.06s

Negative control: restore the old recovery predicate and remove startup receipt
writing while retaining the new tests and type vocabulary. All three configured
route regressions fail; both production files were restored afterward. The
first control attempt stopped at unused-import compilation; remove only the
now-unused imports in the control, without weakening warnings, and rerun:
cargo test -p codewhale-tui --lib -- first_run_configured_
test result: FAILED. 0 passed; 3 failed; 0 ignored; 0 measured; 13801 filtered out; finished in 0.13s

Restored fix and adjacent setup/provider receipt checks:
cargo test -p codewhale-tui --lib -- first_run_ configured_provider_receipt_only_verifies_observed_success tui::setup::progressive_tests doctor_verdict_tests provider_switch_model_override_updates_target_provider_model_slot provider_switch_auth_error_restores_previous_provider_and_model model_picker_apply_is_session_local_until_startup_default_is_requested
test result: ok. 33 passed; 0 failed; 0 ignored; 0 measured; 13771 filtered out; finished in 1.92s

cargo test -p codewhale-config --lib -- setup_state::tests telemetry_metadata_update telemetry_disclosure_records
test result: ok. 21 passed; 0 failed; 0 ignored; 0 measured; 723 filtered out; finished in 0.01s

Direct PTY, rebuilt debug binary, fresh config-only homes, 130x40:
- Supplied OpenAI/loopback/gpui-fixture config opens at the composer, no picker.
  Footer: OpenAI-compatible / gpui-fixture. Persisted provider_model status:
  configured; result: provider=openai, model=gpui-fixture; configured, not checked.
- Configured OpenAI route without credentials opens the navigable picker with
  OpenAI-compatible selected.
- Both processes exited normally. No inference request was sent.

Reproduction boundary: a freshly built baseline at e4c94c8 also skipped the
picker in a direct PTY, but produced no provider-step receipt. The reported
forced-DeepSeek screen was not reproduced here. The full fixture plus running
Ollama tmux check remains pending: tmux socket creation failed with Operation
not permitted in the sandbox. No hosted CI, packaged release, or push claim.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: OpenRouter's /v1/models now lists `~`-prefixed "latest" alias ids
(`~deepseek/deepseek-pro-latest`). They fail `valid_catalog_model_id`, and
one invalid row fails the whole roster closed, so the OpenRouter lake scope
was permanently `failed/invalid_response` and
`fresh_provider_live_pricing_quote_at` never returned a rate: every
OpenRouter turn showed "rate unavailable". Separately, the main interactive
turn froze its dispatch quote only from the lake and never consulted the
operator-declared `[[custom_models]]` rate that background/review
envelopes (`effective_route_envelope`) already honor.

Fix: `parse_openrouter_models_response` drops `~` alias rows (moving
aliases, not billing identities); the fail-closed validator is unchanged
for every other id. A new `CodewhaleClient::configured_pricing_quote_at`
is shared by `effective_route_envelope` and the turn-loop dispatch
boundary, which now prefers a declared rate (scoped to the installed
client's endpoint fingerprint) before the lake quote.

Not covered: server-tag suffixed ids (`:nitro`, `:free`) still need an
exact catalog or declared row; their price can differ from the base id,
so they are not stripped.

Tests (each fails with its fix hunk reverted, passes with it):
- client::tests::fetch_catalog_delta_skips_openrouter_latest_aliases_and_keeps_prices
- core::engine::tests::main_turn_dispatch_freezes_declared_custom_model_rate
test result: ok. 2 passed; 0 failed (reverted: 0 passed; 2 failed)
Related filters (fetch_catalog, openrouter, configured_model_client_tests,
exact_turn_snapshot, effective_route, provider_live_pricing):
test result: ok. 46 passed; 0 failed
cost_status, provider_catalog_live, core::engine::tests::exact:
test result: ok. 60 passed; 0 failed

Refs #6690

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: `codewhale exec` accepted its prompt only as positional argv, so a
prompt past the kernel's per-argument ceiling (131071 bytes on Linux) failed
with E2BIG before Codewhale started, with no alternative transport.

Fix: add `--prompt-file <PATH>` (`-` = stdin), mutually exclusive with the
positional prompt, which is now `required_unless_present = "prompt_file"`; a
lone positional `-` also reads stdin. `resolve_exec_prompt` produces the
effective prompt before any model call and fails loudly on a missing,
unreadable, non-UTF-8 or empty source. Reading goes through a shared
`read_capped_text` (also used by `apply`'s stdin patch reader) with a named
64 MiB byte limit. `-` is refused with `--parent-death-watch`, which owns
stdin for Fleet workers. Implicit stdin reading when no prompt is given is
not added: exec callers (Fleet) inherit pipes on stdin. An argv-side byte
guard is not possible: the kernel rejects the exec before Codewhale runs.

Tests:
- cargo test -p codewhale-tui --lib -- exec_prompt_file read_capped_text
  exec_accepts_split exec_keeps
  test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 13795 filtered out
- With resolve_exec_prompt reverted to the pre-fix argv join:
  test result: FAILED. 0 passed; 1 failed (left: "" vs the 540 KB file body)
- Built binary: 540001-byte file via --prompt-file and via `cat | exec -`
  resolve and reach provider routing; empty stdin -> "exec prompt from stdin
  is empty"; missing file -> "failed to open --prompt-file nope.txt";
  positional + --prompt-file -> clap ArgumentConflict.

Fixes #6688

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Cause: with `provider = "openai"`, `default_text_model = "gpui-fixture"`
and no `[providers.openai].model`, the root alias is the active route's
fallback and nothing records which provider it belongs to.
- TUI: a session-local `/provider <other>` left the alias at the root, so
  `/provider openai` afterwards resolved with `openai != api_provider()`
  and fell through to the catalog default `gpt-5.6`.
- Engine: `reconcile_root_model_aliases` relocated the alias only when the
  incoming route failed validation. A pass-through incoming route
  (openrouter, zai, ...) inherited `gpui-fixture`, and `GET /v1/providers`
  advertised OpenAI's `default_model` as `gpt-5.6`, which is what the
  desktop model chip picked for "OpenAI-compatible".

Fix: `Config::root_model_alias_owned_by_outgoing` names the outgoing
route and value when that route has no leaf and was actually resolving
the alias. The persisted writer moves it onto the outgoing leaf and
clears the root; the TUI switch applies the same move in memory. The
resolution order is then providers.<p>.model, the root alias while its
own provider is active, then the catalog default. Aliases the outgoing
route never resolved (shadowed by a leaf, `auto`, foreign ids dropped by
`default_model`) keep their previous handling. Two persistence tests that
asserted the alias stays at the root in the no-leaf case now assert it
moves to the outgoing leaf; switching back is still checked.

Tests (focused, CARGO_BUILD_JOBS=4):
- new: tui::ui::tests::provider_switch_back_lands_on_root_default_owned_by_that_provider
  and runtime_api::tests::switch_provider_away_and_back_keeps_the_root_default_with_its_route
  fail without the fix (left: "gpt-5.6" right: "gpui-fixture"; openrouter
  body model "gpui-fixture"):
  test result: FAILED. 3 passed; 2 failed; 0 ignored; 0 measured; 13796 filtered out
- with the fix, filter root_default:
  test result: ok. 5 passed; 0 failed; 0 ignored; 0 measured; 13796 filtered out
- filters config_persistence default_text_model alias route_save save_route
  provider_picker route_preferences provider_switch switch_provider root_default:
  test result: ok. 469 passed; 0 failed; 0 ignored; 0 measured; 13332 filtered out
- filters config::tests model_inventory route_runtime runtime_api::tests::switch
  runtime_api::tests::get_config provider:
  test result: ok. 1302 passed; 0 failed; 1 ignored; 0 measured; 12498 filtered out

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Review follow-up for #6691. The live OpenRouter /v1/models lists six
routers (openrouter/auto, auto-beta, fusion, pareto-code, bodybuilder,
typesafe/jev-router) at "prompt":"-1","completion":"-1". parse_price
rejected negatives and the conversion collected with `?`, so the whole
refresh still failed and every turn stayed "rate unavailable".

- A negative OpenRouter price now means "no fixed rate": the row is kept
  with cost None. Rows that do not decode, carry an id the catalog cannot
  hold, or an unreadable/implausible price are skipped and counted with a
  warning. A response that is not a `{"data": [...]}` list, or whose rows
  all fail, is still rejected.
- declared_or_catalog_quote: a [[custom_models]] row with no rates yields
  to the exact endpoint's catalog price at every dispatch boundary (main
  turn, background envelope, EffectiveRouteEnvelope::capture); it stays
  frozen rate-less only when the catalog has none. The turn-loop logic
  moved into client::main_turn_pricing_quote_at so it is testable.
- A pinned OpenRouter vendor no longer blocks an explicit declared rate
  for the exact route; the aggregate catalog price stays blocked.

Against the saved live list (458 rows): refresh succeeds with 440 rows,
434 priced; deepseek/deepseek-v4-pro prices at 0.95526/1.91052 per M.

Focused (cargo test -p codewhale-tui --lib):
fixed:    test result: ok. 10 passed; 0 failed; 0 ignored; 0 measured; 13793 filtered out
reverted (neg price, rate-less fallthrough, vendor pin):
          test result: FAILED. 0 passed; 3 failed; 0 ignored; 0 measured; 13800 filtered out
reverted (per-row decode skip):
          test result: FAILED. 0 passed; 1 failed; 0 ignored; 0 measured; 13802 filtered out
related (fetch_catalog openrouter configured_model effective_route
provider_live_pricing exact_turn_snapshot main_turn cost_status
provider_catalog_live):
          test result: ok. 120 passed; 0 failed; 0 ignored; 0 measured; 13683 filtered out

Fixes #6690

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…ault

Review finding on #6693: moving `default_text_model` onto the outgoing
route's leaf removed only that key. `Config::load` then copies the legacy
root `model` into `default_text_model` whenever the latter is unset, so a
stale legacy value that the root default had been shadowing on every route
came back on the incoming route. With
`model = "deepseek-v4-flash"`, `default_text_model = "gpui-fixture"`,
`provider = "openai"`, switching to openrouter resolved
`deepseek/deepseek-v4-flash`.

Fix: every place the route writer removes the root alias now goes through
`unset_root_model_aliases`, which clears `default_text_model` and the legacy
root `model` together (the legacy key resolved on no route while the alias
existed). `root_model_alias_owned_by_outgoing` also treats a legacy-only
root `model` as the alias, the same fallback `default_model` applies, so a
legacy-only choice moves onto the outgoing leaf instead of being inherited.
The switch-away-and-back persistence test now expects that move for the
legacy `model` key as well.

The TUI session-local switch needed no change: it writes the resolved model
onto the incoming leaf in memory, which shadows the legacy value.

Tests (focused, CARGO_BUILD_JOBS=4):
- new: config_persistence::tests::moving_root_default_off_the_root_does_not_resurrect_a_shadowed_legacy_model
  and runtime_api::tests::switch_provider_does_not_resurrect_a_shadowed_legacy_root_model
  fail without the fix (openrouter body model "deepseek/deepseek-v4-flash"):
  test result: FAILED. 6 passed; 2 failed; 0 ignored; 0 measured; 13796 filtered out
- with the fix, filters resurrect config_persistence default_text_model alias
  route_save save_route provider_picker route_preferences provider_switch
  switch_provider root_default legacy_model:
  test result: ok. 477 passed; 0 failed; 0 ignored; 0 measured; 13326 filtered out
- filters config::tests model_inventory route_runtime runtime_api::tests::get_config provider:
  test result: ok. 1303 passed; 0 failed; 1 ignored; 0 measured; 12499 filtered out

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…ch (#6688)

Review of #6692:

1. A lone positional `-` read stdin, which changed behavior for raw-prompt
   callers: cloud dispatch builds `codewhale exec --auto <job.prompt>`, so a
   job whose prompt is `-` would block on or fail reading stdin. Only
   `--prompt-file -` reads stdin now; a positional `-` is prompt text again.
   Help text, docs/MODES.md (+ zh_hans) and the CHANGELOG say so.
2. The tests called `resolve_exec_prompt` directly and would pass with the
   exec dispatch call site reverted. A new integration test runs the real
   `codewhale-tui exec` binary against a wiremock provider and checks the
   user message the model receives: a 200 KiB `--prompt-file` body, a
   `--prompt-file -` stdin body, and a literal `-` with stdin piped (never
   read).

Tests:
- cargo test -p codewhale-tui --lib -- exec_prompt_file read_capped_text
  exec_accepts_split exec_keeps
  test result: ok. 6 passed; 0 failed; 0 ignored; 0 measured; 13795 filtered out
- cargo test -p codewhale-tui --test integration -- exec_turn_usage
  test result: ok. 9 passed; 0 failed; 0 ignored; 0 measured; 179 filtered out
- Dispatch call site reverted to `join_prompt_parts(&args.prompt)`:
  exec_prompt_reaches_the_model_from_prompt_file_and_stdin
  test result: FAILED. 0 passed; 1 failed (left: "")
- Lone positional `-` arm restored: the unit test blocked reading stdin
  (killed after 9 minutes), which is the regression itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Comment thread crates/tui/src/cost_status.rs Dismissed
CodeWhale Bot and others added 4 commits September 27, 2026 21:03
… on a keyed route

The reported first-launch symptom (a configured openai/gpui-fixture home
opening on "Choose your model provider" with DeepSeek focused) was not a
config-resolution bug. The tmux server used for the repro carries
CODEWHALE_PROVIDER=deepseek and CODEWHALE_MODEL=deepseek-v4-flash in its
global environment, so every new session inherited them. Config::load then
resolved deepseek per the documented env > file precedence (instrumented:
provider=deepseek already at load_config_from_cli_with_effective_profile).
`codewhale model resolve` from a shell without those vars reported openai,
which hid the difference. With the vars unset, the unmodified release build
aaaadfe starts on the composer with "OpenAI-compatible · gpui-fixture".

It did expose a real bug. After Esc on the recovery picker, `/provider openai`
switched to a keyed route, but switch_provider never recomputed
app.onboarding_needs_api_key. Only complete_provider_picker_onboarding clears
it, and that runs only while onboarding == Provider. The stale true kept the
footer on "model not connected" (frame.rs info_segments) and left
should_adopt_live_local_ollama armed against the provider the user had just
chosen.

Fix: on a successful switch, set onboarding_needs_api_key from
has_api_key(config) for the new route, and clear
onboarding_missing_key_recovery when the key is present. The auth-failure
rollback already restores the previous flag.

Tests (focused, local):
- new first_run_switch_to_keyed_route_clears_launch_missing_key_state
  (real Config::load + App::new startup with the inherited env pair, Esc,
  switch_provider): passes; fails without the fix at the
  !onboarding_needs_api_key assert.
- provider_switch|switch_provider|local_ollama|missing_key|onboarding:
  109 passed, 0 failed.
- first_run_|provider_picker: 161 passed, 1 failed. The one failure,
  first_run_ollama_choice_survives_restart_from_canonical_config ("index not
  found"), also fails on a clean aaaadfe tree and is not touched here.
- first_run_route_env_guards now also removes CODEWHALE_PROVIDER.

E2E (tmux 130x40, debug build, fixture at 127.0.0.1:4880):
- configured home, env clean: composer, footer
  "OpenAI-compatible · gpui-fixture · thinking: max".
- configured home with the inherited deepseek env: picker (correct per
  precedence), then Esc and /provider openai: footer
  "OpenAI-compatible · gpt-5.6 ..." (release build: "model not connected").
- unconfigured home with local Ollama unreachable: "Choose your model
  provider" picker. With the live local Ollama it auto-adopts ollama/qwen3:4b,
  same as the release build.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
…6693

Hosted CI on 776e3fe (PR #6672) failed two codewhale-tui tests.

1. client::catalog_tests::raw_rows_keep_duplicate_known_fields_invalid_in_existing_parsers
   Cause: #6691 (79d9815) made OpenRouter parsing per-row by decoding
   `data` as Vec<serde_json::Value>. A Value map keeps the last of a
   duplicated key, so `{"id":"second","id":"replacement"}` was silently
   accepted as "replacement" (and a duplicated pricing.prompt took the
   last rate) instead of being rejected.
   Fix (product): keep rows as Box<RawValue> and decode each with
   from_str, so serde's duplicate-field error still fires and the row is
   skipped and counted as malformed like any other undecodable row. The
   test now asserts the per-row contract: the ambiguous row is dropped,
   "first" survives; Baseten still rejects the whole response.

2. tui::ui::tests::first_run_ollama_choice_survives_restart_from_canonical_config
   Cause: #6693 (1394c29, eb48db7) moves a root default_text_model
   onto the leaf of the outgoing route that was resolving it. The test
   seeded default_text_model = "deepseek-v4-pro" with DeepSeek active and
   no leaf, then asserted the root key survived the switch to Ollama
   ("index not found"). That encoded superseded behavior.
   Fix (test): assert the root alias is gone and
   [providers.deepseek].model = "deepseek-v4-pro", which is the #6693
   contract (switching back keeps the choice; Ollama never inherits it).

Tests (CARGO_BUILD_JOBS=4, cargo test -p codewhale-tui --lib):
- raw_rows_keep_duplicate_known_fields_invalid_in_existing_parsers:
  test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 13814 filtered out
- first_run_ollama_choice_survives_restart_from_canonical_config:
  test result: ok. 1 passed; 0 failed; 0 ignored; 0 measured; 13814 filtered out
- filters catalog openrouter fetch_catalog first_run_ provider_picker local_ollama:
  test result: ok. 450 passed; 0 failed; 2 ignored; 0 measured; 13363 filtered out
  (a first run of the same filters had 1 failure in
  provider_picker::tests::credential_draft_is_masked_and_escape_drops_it_without_persistence,
  which passed 3/3 alone and in the rerun; it counts files in a home dir
  and is untouched by this change)

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
Test (windows-latest) failed runtime_surface_review_documented_ssh_host_loads
("documented SSH worker example"): Windows checkouts give
docs/zh_hans/FLEET.md CRLF line endings, so splitting on "```json\n"
found no block. The test now normalizes line endings before parsing.

Checks: cargo test -p codewhale-tui --lib -- fleet::host::tests:
21 passed; 0 failed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014ZwqatxgVFxHvovngywnks
@Hmbown
Hmbown merged commit 0bfe04e into main Sep 28, 2026
37 of 38 checks passed
@Hmbown
Hmbown deleted the integrate/0.10.1 branch September 28, 2026 08:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants