Observed on clean upstream main
The shared-process full-workspace gate is red on clean main@0bfe04e163e16ecd583a7f771ee2cd787cd6ff27, after integration PR #6672. No FEAT-029 commits or edits are in this checkout. This is an observed failure on the merged source, not a claim that #6672 caused every failure.
The fixes from #6666 are present: 45017fa532bf6bb45b11b0b4d77aa4ad1b2d6258 is an ancestor of main, its four changes remain in source, and all corresponding affected tests passed in this run. The current nine failing identities have zero overlap with the original nine FEAT-029 Phase 8 failures. Earlier baseline repairs #6581/#6666 are not missing or reverted.
Exact reproduction and result
# From a clean checkout of the revision above. Use a short, private TMPDIR.
DiagnosticTmpDir=$(mktemp -d /var/tmp/cw29m.XXXXXX)
TMPDIR="$DiagnosticTmpDir" sh scripts/with-hermetic-test-home.sh cargo test --workspace --all-features --locked -- --format=pretty
Environment: Linux 7.0.0-31-generic x86_64, Rust 1.98.0 (88d9e12ae 2026-08-18), 28 logical CPUs, default libtest concurrency, no custom RUSTFLAGS/CARGO_TARGET_DIR/test-thread override, fresh hermetic HOME. Cargo commands ran strictly sequentially. The wrapper sets RUST_MIN_STACK=16777216. TMPDIR is on the disk filesystem with ample space and inodes.
Native result:
test result: FAILED. 13792 passed; 9 failed; 20 ignored; 0 measured; 0 filtered out; finished in 322.04s
error: test failed, to rerun pass `-p codewhale-tui --lib`
Cargo exited 101 and stopped at the TUI library; later integration targets and doctests were not reached.
Failing identities:
remote_control::tests::classic_recovery_uses_persisted_seq_floor_and_ignores_older_terminal
remote_control::tests::actual_start_reclaims_runtime_chat_writer_before_worker_spawn
remote_control::tests::separate_predispatch_crashes_on_one_run_get_distinct_recovery_turn_ids
runtime_api::tests::events_endpoint_respects_since_seq_cursor
runtime_threads::tests::approval_required_awaits_external_decision_allow
runtime_threads::tests::approval_remember_grants_tool_class_without_changing_posture
tools::subagent::budget_handback_tests::budget_handback_inflight_wall_timeout_persists_unreported_usage
tools::subagent::tests::child_permission_gate::wall_deadline_ends_pending_wait_with_receipt
tools::subagent::tests::resume_keeps_recorded_reasoning_in_manifest_and_request_after_parent_changes
The captured failure stdout in this run includes timed out waiting for terminal turn status for events_endpoint_respects_since_seq_cursor. The native output did not retain stdout blocks for the other eight failures; those errors are not invented.
Controlled comparisons, with all outcomes retained
- A preceding run on the same source reported
13,783 passed; 18 failed; 20 ignored. Eight failures were a diagnostic runner mistake: a long TMPDIR made Unix socket paths exceed SUN_LEN. A controlled repeat changed only TMPDIR to a short path. All eight socket cases then passed; the nine failures above remained. That first result is retained, not counted as a valid socket regression.
- That preceding run also recorded classic-journal reopen errors:
This saved session still owns an unfinished account turn in its original workspace; reconnect that workspace or create a new local session. The affected journal identities differed across runs.
- Other preceding failure evidence: child approval prompt
Elapsed(()); budget-handback result missing the expected in-flight wall-time exhaustion reason; a crash-dump fixture observed two files instead of one; an active-workspace restore fixture received 500 instead of 409. These describe that preceding run, not missing stdout invented for the second.
- All 14 distinct non-socket failing identities across both runs passed together in a narrower
--exact selection on the same workspace-built binary: 14 passed; 0 failed; 13807 filtered out, 4.01s.
- The entire remote-control module passed on that same binary under syscall tracing:
82 passed; 0 failed; 13739 filtered out, 2.01s. The failing journal mechanism did not reproduce in this narrower selection.
- These narrower green runs do not invalidate either full-suite failure. No retry-until-green, test removal, assertion weakening, or gate change was performed.
Why the CI test result differs
The CI test step runs:
sh scripts/with-hermetic-test-home.sh cargo nextest run --workspace --all-features --locked --profile ci
It separately runs workspace doctests. Nextest gives each test its own process; the requested libtest gate shares process globals across concurrent tests. The configured nextest profiles disable retries and have explicit integration concurrency groups. The difference is runner isolation/scheduling, not an absent #6666 patch.
The nextest configuration still declares the shared-process Cargo command authoritative for release. No workflow step runs that complete shared-process workspace gate. This discrepancy deserves an explicit maintainer decision about CI coverage; a nextest PASS does not establish a libtest PASS.
Integration CI run 36393467574 completed successfully, including Linux/macOS/Windows tests. Its head 2ea48583cde64ee84ab4b1bbb7e8357cbecaebcc and merged main have the same source tree 0074786dfb86ffcc01ab13ab6497173f3615cf7f. The Linux Run tests step is recorded SUCCESS. The GitHub CLI log request returned an empty log, so this report does not claim individual nextest PASS lines for the nine names. GitGuardian is a separate failed check on #6672 and is not the explanation for these Rust outcomes.
Diagnostic leads, not established causes
Intended outcome and acceptance criteria
- Establish the cause of each remaining failed identity using controlled selection/order/interference or deterministic reproduction; preserve failing and passing outcomes and distinguish hypotheses from causal evidence.
- Repair the identified product or fixture causes with affected tests and safety/data-integrity assertions retained. Do not ignore tests, reduce the selected gate, cap threads to hide a race, or substitute isolated/nextest passes for the shared-process result.
- Run the exact shared-process full-workspace command on one fixed source revision with equivalent documented conditions; require two complete passing runs and explicit executed results for all nine identities. Any new failure stays recorded and unresolved.
- Pass the existing CI runner/profile and resolve actionable findings; make the authoritative-gate/CI coverage difference explicit without silently changing acceptance or weakening gates.
Local evidence is retained separately from earlier diagnostics at evidence/postmerge-2026-09-28/ in the FEAT-029 MemoryBank: main-full-workspace-01.log/.json, main-full-workspace-02.log/.json, main-failure-outcomes.json, main-failing-identities-isolated-01.log/.json, main-remote-control-trace-02.log/.json/.strace, linux-flock-inheritance-experiment.json, and merged-main-validation-report.md. These files are local evidence, not links falsely represented as public artifacts. The available native failure details and exact commands are included above so the report is reviewable without access to that directory. FEAT-029 remains unchanged pending a verified baseline.
Paulo Aboim Pinto
Observed on clean upstream main
The shared-process full-workspace gate is red on clean
main@0bfe04e163e16ecd583a7f771ee2cd787cd6ff27, after integration PR #6672. No FEAT-029 commits or edits are in this checkout. This is an observed failure on the merged source, not a claim that #6672 caused every failure.The fixes from #6666 are present:
45017fa532bf6bb45b11b0b4d77aa4ad1b2d6258is an ancestor of main, its four changes remain in source, and all corresponding affected tests passed in this run. The current nine failing identities have zero overlap with the original nine FEAT-029 Phase 8 failures. Earlier baseline repairs #6581/#6666 are not missing or reverted.Exact reproduction and result
Environment: Linux 7.0.0-31-generic x86_64, Rust 1.98.0 (88d9e12ae 2026-08-18), 28 logical CPUs, default libtest concurrency, no custom RUSTFLAGS/CARGO_TARGET_DIR/test-thread override, fresh hermetic HOME. Cargo commands ran strictly sequentially. The wrapper sets
RUST_MIN_STACK=16777216. TMPDIR is on the disk filesystem with ample space and inodes.Native result:
Cargo exited 101 and stopped at the TUI library; later integration targets and doctests were not reached.
Failing identities:
remote_control::tests::classic_recovery_uses_persisted_seq_floor_and_ignores_older_terminalremote_control::tests::actual_start_reclaims_runtime_chat_writer_before_worker_spawnremote_control::tests::separate_predispatch_crashes_on_one_run_get_distinct_recovery_turn_idsruntime_api::tests::events_endpoint_respects_since_seq_cursorruntime_threads::tests::approval_required_awaits_external_decision_allowruntime_threads::tests::approval_remember_grants_tool_class_without_changing_posturetools::subagent::budget_handback_tests::budget_handback_inflight_wall_timeout_persists_unreported_usagetools::subagent::tests::child_permission_gate::wall_deadline_ends_pending_wait_with_receipttools::subagent::tests::resume_keeps_recorded_reasoning_in_manifest_and_request_after_parent_changesThe captured failure stdout in this run includes
timed out waiting for terminal turn statusforevents_endpoint_respects_since_seq_cursor. The native output did not retain stdout blocks for the other eight failures; those errors are not invented.Controlled comparisons, with all outcomes retained
13,783 passed; 18 failed; 20 ignored. Eight failures were a diagnostic runner mistake: a long TMPDIR made Unix socket paths exceed SUN_LEN. A controlled repeat changed only TMPDIR to a short path. All eight socket cases then passed; the nine failures above remained. That first result is retained, not counted as a valid socket regression.This saved session still owns an unfinished account turn in its original workspace; reconnect that workspace or create a new local session.The affected journal identities differed across runs.Elapsed(()); budget-handback result missing the expected in-flight wall-time exhaustion reason; a crash-dump fixture observed two files instead of one; an active-workspace restore fixture received 500 instead of 409. These describe that preceding run, not missing stdout invented for the second.--exactselection on the same workspace-built binary:14 passed; 0 failed; 13807 filtered out, 4.01s.82 passed; 0 failed; 13739 filtered out, 2.01s. The failing journal mechanism did not reproduce in this narrower selection.Why the CI test result differs
The CI test step runs:
It separately runs workspace doctests. Nextest gives each test its own process; the requested libtest gate shares process globals across concurrent tests. The configured nextest profiles disable retries and have explicit integration concurrency groups. The difference is runner isolation/scheduling, not an absent #6666 patch.
The nextest configuration still declares the shared-process Cargo command authoritative for release. No workflow step runs that complete shared-process workspace gate. This discrepancy deserves an explicit maintainer decision about CI coverage; a nextest PASS does not establish a libtest PASS.
Integration CI run 36393467574 completed successfully, including Linux/macOS/Windows tests. Its head
2ea48583cde64ee84ab4b1bbb7e8357cbecaebccand merged main have the same source tree0074786dfb86ffcc01ab13ab6497173f3615cf7f. The Linux Run tests step is recorded SUCCESS. The GitHub CLI log request returned an empty log, so this report does not claim individual nextest PASS lines for the nine names. GitGuardian is a separate failed check on #6672 and is not the explanation for these Rust outcomes.Diagnostic leads, not established causes
ClassicSessionOwnerLockcloses its descriptor on Drop but does not explicitly unlock, whereas the relay/runtime store owner locks do. A controlled Linux experiment demonstrates a forked child can retain the flock through a duplicated pre-exec descriptor after the parent closes it. This is only a mechanism consistent with the journal errors; an actual product failure trace or a faithful deterministic reproducer is still needed.Intended outcome and acceptance criteria
Local evidence is retained separately from earlier diagnostics at
evidence/postmerge-2026-09-28/in the FEAT-029 MemoryBank:main-full-workspace-01.log/.json,main-full-workspace-02.log/.json,main-failure-outcomes.json,main-failing-identities-isolated-01.log/.json,main-remote-control-trace-02.log/.json/.strace,linux-flock-inheritance-experiment.json, andmerged-main-validation-report.md. These files are local evidence, not links falsely represented as public artifacts. The available native failure details and exact commands are included above so the report is reviewable without access to that directory. FEAT-029 remains unchanged pending a verified baseline.Paulo Aboim Pinto