Skip to content

Track upstream agent CLI versions - #1620

Open
a5c-ai[bot] wants to merge 2 commits into
stagingfrom
agent-versions/daily-2026-08-04
Open

Track upstream agent CLI versions#1620
a5c-ai[bot] wants to merge 2 commits into
stagingfrom
agent-versions/daily-2026-08-04

Conversation

@a5c-ai

@a5c-ai a5c-ai Bot commented Aug 4, 2026

Copy link
Copy Markdown
Contributor

Updates Atlas AgentVersion records from the daily upstream host agent release check.

Artifacts:

  • artifacts/agent-version-tracker/upstream-targets-and-latest.json
  • artifacts/agent-version-tracker/summary.json

Verification:

  • npm run verify:metadata
  • npm run build --workspace=@a5c-ai/atlas

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out waiting for workflow completion.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30964834356

The workflow was dispatched against agent-versions/daily-2026-08-04 with checkout input ref=agent-versions/daily-2026-08-04. The Babysitter QA wait step timed out after 20 minutes before the live-stack matrix reached a terminal conclusion, so this is not a passing QA verdict yet.

Job Result
Build All pass
Compute Matrix pass
Live Stack (ubuntu-latest-l, vanilla, hermes/gpt-5.5, non-interactive) in progress
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, bridged-hooks) in progress
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, interactive) in progress
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) in progress
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) queued

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Overall verdict: pending / incomplete. No live-stack job had failed at the time of this comment, but the matrix had not completed, so catalog/evidence consistency still needs final CI confirmation from the linked run.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Run: https://github.com/a5c-ai/babysitter/actions/runs/30964850568

Overall verdict: not complete within the 20-minute QA wait window. No live-stack failures were observed before timeout; build, matrix computation, and all vanilla NI scenarios had passed. The two Codex BP interactive scenarios were still in progress when the wait window expired.

Job Result
Build All pass
Compute Matrix pass
Live Stack (ubuntu-latest-l, vanilla, pi/gpt-5.4-mini, non-interactive) pass
Live Stack (ubuntu-latest-l, vanilla, codex/gemini-3.5-flash, non-interactive) pass
Live Stack (ubuntu-latest-l, vanilla, claude-code/gpt-5.5, non-interactive) pass
Live Stack (ubuntu-latest-l, bp/create, codex/gpt-5.5, interactive) in progress at timeout
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) in progress at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"pi","model":"foundry-gpt54mini","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"}
]

Scope note: this matrix targets Atlas agent-version catalog/evidence consumption paths for PR #1620 rather than transport or hook changes.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: not final - timed out waiting for selected live-stack jobs.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30964913235

The workflow was dispatched for adversarial QA of PR #1620. The process wait window expired after Build All completed and while the selected scenario jobs were still in progress.

Tested matrix

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp predefined
gemini foundry-gpt55 ni vanilla -

Current job status

Job Status Conclusion
Build All completed success
Compute Matrix completed success
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) in_progress -
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, bridged-hooks) in_progress -
Live Stack (ubuntu-latest-l, vanilla, gemini-cli/gpt-5.5, non-interactive) in_progress -
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) in_progress -

Overall verdict: not passed yet. No scenario failures were observed before timeout, but the live-stack jobs had not completed.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete. The live-stack workflow was dispatched and is still running, but the QA polling process timed out after 20 minutes before the scenario jobs completed.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30964922275

Matrix tested:

Agent Model Mode Install Process mode
codex foundry-gpt55 ni vanilla predefined
claude anthropic-sonnet46 bridged-interactive vanilla predefined
gemini google-gemini31 ni vanilla predefined
pi foundry-gpt55 bridged-interactive vanilla predefined
copilot foundry-gpt55 ni vanilla predefined
hermes foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp create

Current job status:

Job Result
Build All pass
Compute Matrix pass
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, bridged-hooks) in_progress
Live Stack (ubuntu-latest-l, bp/create, hermes/gpt-5.5, interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, claude-code/claude-sonnet-4-6, bridged-interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, copilot-cli/gpt-5.5, non-interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, codex/gpt-5.5, non-interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, pi/gpt-5.5, bridged-interactive) in_progress
Live Stack (ubuntu-latest-l, vanilla, gemini-cli/gemini-3.5-flash, non-interactive) in_progress

Overall verdict: not passed yet. Follow the linked run for final scenario conclusions.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review result: changes requested. GitHub would not let this actor submit a formal request-changes review because the PR is bot-authored by the same app, so I’m posting the blocking review as a comment.

Blockers

  • packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 adds a new batch of AgentVersion records, but the PR does not add the matching packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml shard. The AgentVersion schema marks versionRange, cliCommand, and release dates as evidence-bound, and prior upstream tracker drops include EvidenceSource records that reference each added AgentVersion. Without that shard, these new version claims land without first-class Atlas provenance and evidence-manifest consumers cannot audit them. Please add the 2026-08-04 EvidenceSource file mirroring the existing 2026-07-17 tracker format and reference the exact new AgentVersion IDs.

Majors

  • artifacts/agent-version-tracker/summary.json:329 says release note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but those directories/files are not present in the PR; the PR only changes summary.json and upstream-targets-and-latest.json under that artifact directory. Either include the referenced artifacts or update the summary so it only describes files that actually exist.

QA

I dispatched QA via qa-dispatch.yml for PR #1620. The dispatch workflow completed, but the underlying Live Stack run did not fully pass within the QA wait window: build, matrix, and three vanilla non-interactive jobs passed; two Codex BP interactive jobs were still in progress at timeout. The QA process posted its comment at #1620 (comment). I am treating that as inconclusive/not passed for this review decision.

Risk Assessment

Risk level: risk:medium.

  • Catalog provenance regression: Atlas consumers may ingest new AgentVersion records without linked EvidenceSource nodes. Mitigation: add the missing 2026-08-04 EvidenceSource shard before merge and rerun Atlas/catalog validation.
  • Generated artifact audit gap: the summary points maintainers to non-existent local evidence files. Mitigation: commit the referenced artifact bodies or correct the generated summary note.
  • QA uncertainty: two selected live-stack jobs were still running when the QA process timed out. Mitigation: wait for the Live Stack run to finish or rerun QA after the data fixes, and do not merge while QA remains inconclusive.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Requesting changes based on the adversarial review process.

Major finding:

packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:3 adds the new 2026-08-04 AgentVersion batch, but the PR does not add the matching packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml provenance file. Recent upstream-current batches include a same-date evidence-source file under catalog-meta/evidence-sources/; for example the 2026-07-17 batch links each release/package source back to the corresponding AgentVersion and agent IDs. A grep of the PR head finds representative new IDs such as agentVersion:amp:0-0-1785819659-g30d128, agentVersion:antigravity:1-1-10, agent-version:codex@0.146.0, and agentVersion:qwen:0-21-5 only in the AgentVersion YAML, not in catalog-meta evidence. This leaves the new catalog records without graph evidence/provenance for downstream evidence APIs and review workflows.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml with EvidenceSource records for each new version, following packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-07-17.yaml, and rerun the Atlas build/index generation.

QA status:

The dispatched qa-dispatch.yml run completed, but the nested QA verdict was incomplete. It dispatched live-stack workflow 30964922275; Build All and Compute Matrix passed, while seven live-stack scenario jobs were still in_progress when the QA process hit its 20-minute polling timeout. The PR also currently has a failed Docs QA check due to stale generated docs.

Risk Assessment

Risk level: risk:medium

  • Catalog provenance regression: new AgentVersion records can appear without corresponding evidence/source records. Mitigation: add the missing evidence-source YAML and verify evidence lookup for the new IDs.
  • Validation risk: QA is incomplete and Docs QA is red. Mitigation: wait for live-stack completion and get CI green before merge, or explicitly resolve any known unrelated docs freshness failure.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete. The live-stack workflow was dispatched and was still running when the 20-minute QA polling window expired.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30965049909

Focused matrix:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 ni vanilla -

Current job status at timeout:

Job Result
Build All pass
Compute Matrix pass
Live Stack (ubuntu-latest-l, vanilla, codex/gemini-3.5-flash, non-interactive) in progress
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) in progress
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, interactive) in progress
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) in progress
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, interactive) in progress

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix focuses on BP predefined/create graph/catalog consumers plus one vanilla adapter baseline. The workflow result should be checked once the remaining jobs complete.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: in progress after Babysitter polling timeout. The QA process waited 20 minutes; the GitHub Actions run is still active.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30965059325

Focus: adversarial review of Atlas graph agent-version metadata update, including affected Atlas graph/catalog surfaces, provenance/evidence-source coverage, and generated tracker artifact consistency.

Tested matrix

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp create
claude foundry-gpt55 bridged-hooks bp predefined
hermes foundry-gpt55 ni vanilla predefined
gemini google-gemini31 bridged-interactive vanilla predefined
claude anthropic-sonnet46 ni vanilla predefined

Current job status

Job Result
Compute Matrix pass
Build All pass
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, bridged-hooks) in progress
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, interactive) in progress
Live Stack (ubuntu-latest-l, vanilla, gemini-cli/gemini-3.5-flash, bridged-interactive) in progress
Live Stack (ubuntu-latest-l, vanilla, hermes/gpt-5.5, non-interactive) in progress
Live Stack (ubuntu-latest-l, vanilla, claude-code/claude-sonnet-4-6, non-interactive) in progress

Overall verdict: not yet complete; no scenario failure has been reported, but the live-stack scenario jobs have not reached terminal conclusions yet.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review result: changes requested.

GitHub would not let this actor submit a formal request-changes review because the PR is bot-authored by the same app, so I am posting the review decision as a comment.

Major Finding

  • packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 adds 16 new AgentVersion records but does not add the matching packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml provenance file. Prior upstream-current tracker runs, including upstream-current-2026-07-17.yaml, add EvidenceSource nodes that reference each added AgentVersion and product. Without that same-date evidence shard, the new version claims are less traceable for catalog consumers and review workflows.

Fix: add the 2026-08-04 EvidenceSource records using the established 2026-07-17 pattern, reference the exact new AgentVersion IDs, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Additional Issues

  • artifacts/agent-version-tracker/summary.json:322 lists only the new AgentVersion YAML under changedFiles, even though this PR also changes the tracker summary and upstream-targets artifact. Include all committed tracker files or narrow the field name/meaning.
  • artifacts/agent-version-tracker/summary.json:329 says release-note bodies and issue bodies are stored under artifacts/agent-version-tracker/release-notes/ and artifacts/agent-version-tracker/issues/, but those directories are not present in the PR. Commit the referenced artifacts or remove/update the note.

QA

Local verification passed for npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas in an isolated PR worktree after installing Atlas workspace dependencies. The Atlas build still emitted existing library bridge quality failures and BAD_ALIAS warnings but exited 0.

Dispatched QA via qa-dispatch.yml. The dispatcher run 30964910121 completed, but nested Live Stack run 30965059325 did not produce a passing verdict. The QA process comment reported Build All and Compute Matrix passed while live-stack jobs were still in progress at its polling timeout, and a follow-up status check showed Live Stack (ubuntu-latest-l, vanilla, claude-code/claude-sonnet-4-6, non-interactive) failed in Run selected live stack E2E.

Risk Assessment

Risk level: risk:medium.

  • Catalog provenance regression: new AgentVersion records can appear without corresponding evidence/source records. Mitigation: add the missing 2026-08-04 evidence-source YAML and verify evidence lookup for the new IDs.
  • Tracker artifact contract drift: summary.json changed shape and underreports changed files. Mitigation: stabilize or version the summary schema and add a consistency check for referenced artifact paths.
  • QA risk: nested live-stack QA has a failed scenario. Mitigation: inspect run 30965059325, fix the failure or prove it unrelated, then rerun QA to a terminal passing verdict.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: pending / not yet passed. The live-stack workflow was dispatched for adversarial QA, but the process wait window timed out while the GitHub Actions run was still in progress.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30965297624

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Matrix tested:

[
  {"agent":"codex","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"claude","model":"anthropic-sonnet46","mode":"ni","install":"vanilla","live":true},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true},
  {"agent":"pi","model":"foundry-deepseek","mode":"ni","install":"vanilla","live":true},
  {"agent":"hermes","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"antigravity","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"create"}
]

Overall verdict: not passed yet because final job conclusions are unavailable. Re-check the run link for the final result.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: timeout / still running.

Run: https://github.com/a5c-ai/babysitter/actions/runs/30965295899

The predefined QA process waited 20 minutes, but the run was still in_progress. Current job status at timeout:

Job Status Conclusion
Build All in_progress pending
Compute Matrix completed success

Matrix tested:

Agent Model Mode Install Process mode
codex google-gemini31 ni vanilla predefined
claude foundry-gpt55 ni vanilla predefined
hermes foundry-gpt55 ni vanilla predefined
codex google-gemini31 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 bridged-hooks bp predefined

Reasoning: PR #1620 changes Atlas agent-version graph metadata and generated tracker artifacts. This matrix covers representative graph consumers across raw adapter install and babysitter-plugin install paths without running the full cross-product.

@a5c-ai

a5c-ai Bot commented Aug 5, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review decision: changes requested

GitHub would not allow this bot account to submit a formal request-changes review because it is the PR author, so I am posting the decision as a PR comment instead.

I cannot approve this as a complete Atlas graph update in its current form. The review found no security blockers, but it did find major correctness/provenance problems and the higher-level approach check failed.

Major findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:25 - New AgentVersion facts are not linked to evidence.

All 16 new AgentVersion records add release dates, upstream tags, release-note summaries, and behavioral/security claims, but their edges only contain version_of. The prior upstream-current-2026-07-17 sweep added matching EvidenceSource records and sourced_from links. Without equivalent provenance here, catalog consumers cannot verify trust, freshness, or source evidence for these facts.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml with EvidenceSource records for the upstream/package sources, then link each new AgentVersion to the matching evidence with sourced_from edges or direct evidence source references.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - New AgentVersion IDs do not match the SDK-generated stable ID shape.

The Codex record uses agent-version:codex@0.146.0; the expected canonical form for agentId=codex and versionRange=0.146.0 is agentVersion:codex:0-146-0. The Cursor record has the same class of issue. ID-based lookups, generated references, evidence links, and future relationships can miss these nodes or create duplicate logical versions.

Fix: rename the affected IDs to the canonical agentVersion:<agentId>:<slugified versionRange> form and update any references or evidence records accordingly.

Additional issues

  • artifacts/agent-version-tracker/summary.json:329 says release-note bodies live under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but those directories are not present in the PR. Commit the referenced artifacts or adjust the summary to match what is actually included.
  • The existing catalog ID-alignment test only checks Copilot, so it does not catch the Codex/Cursor ID regressions. Please generalize this across AgentVersion nodes and add a provenance/evidence coverage check for upstream-current records.

Missing scope

The PR description says this updates Atlas AgentVersion records from the daily upstream release check, but the diff only adds raw AgentVersion graph records and tracker artifacts. For this to be mergeable as a complete graph update, it should also include EvidenceSource records, sourced_from links, stable IDs, current-version/current-product pointer updates where applicable, expanded tests, and any catalog/adapter behavior updates implied by the release notes. Alternatively, explicitly narrow the PR scope to raw tracker output rather than a complete Atlas assimilation.

QA

QA Dispatch was triggered for PR #1620 on branch agent-versions/daily-2026-08-04 as run 30965054962. It was still in progress when the review process hit its polling timeout, but it completed successfully shortly afterward. The decision remains changes requested because of the major graph correctness/provenance issues above.

Risk Assessment

Risk level: risk:high

  • Unbacked release claims can enter the graph and later be treated as trusted catalog facts. Mitigation: require EvidenceSource coverage and sourced_from/evidence links for every upstream version record before merge.
  • Consumers that use canonical AgentVersion IDs may fail to resolve the new Codex and Cursor versions. Mitigation: normalize IDs to the SDK-generated form and add a catalog-wide ID consistency test.
  • Partial graph updates can be interpreted as complete upstream coverage by dashboards or automation. Mitigation: either complete the assimilation in this PR or narrow the declared scope so downstream consumers do not treat it as authoritative coverage.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: timeout / still running.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061195839

The predefined QA process dispatched live-stack QA for adversarial review of PR #1620 and waited 20 minutes. The GitHub Actions run was still in_progress when the wait window expired, so this is not a passing QA verdict yet.

Job Status Result
Build All in_progress pending
Compute Matrix completed pass

Matrix tested:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 interactive bp predefined
claude foundry-gpt55 interactive bp create
codex google-gemini31 ni vanilla -

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this adversarial QA matrix focuses on BP predefined/create graph/catalog consumer paths across Codex and Claude, plus one raw Codex adapter baseline.

Overall verdict: not passed yet because final job conclusions are unavailable. Re-check the linked run for the terminal result.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out waiting for workflow completion.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061208047

The predefined QA process was executed for adversarial review of PR #1620 against branch agent-versions/daily-2026-08-04. The process waited 20 minutes, but the Live Stack workflow had not reached a terminal conclusion.

Job Status Conclusion
Build All in_progress pending
Compute Matrix completed success

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix focuses on BP predefined/create graph/catalog consumers, bridged-hooks plugin coverage, and vanilla Hermes/Gemini adapter baselines across Foundry and Gemini providers.

Overall verdict: not passed yet. Follow the linked run for final Live Stack conclusions.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out waiting for workflow completion.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061241026

The predefined QA process dispatched live-stack.yml against agent-versions/daily-2026-08-04 for adversarial review. The 20-minute polling window expired while the Actions run was still in_progress, so this is not a passing QA verdict yet.

Job Status Result
Compute Matrix completed pass
Build All in_progress pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this matrix targets BP predefined graph/catalog consumption, BP create process-generation paths, bridged-hooks plugin behavior, and vanilla adapter baselines across Foundry and Gemini providers.

Overall verdict: not passed yet. Re-check the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / queued at timeout.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061264182

The predefined QA process dispatched a focused live-stack matrix for adversarial review of the Atlas AgentVersion graph metadata update, with attention to catalog/evidence/source coverage and generated artifact consistency. The process waited 20 minutes, but the GitHub Actions run did not reach a terminal result. At timeout, the workflow was still queued overall.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Focused matrix:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 bridged-hooks bp predefined
gemini google-gemini31 bridged-interactive vanilla -
hermes foundry-gpt55 ni vanilla -

Overall verdict: not passed yet. No live-stack scenario failure was observed, but final job conclusions are unavailable because the run remained queued through the QA wait window.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas agent-version graph update. The review found no security blocker, but it found major graph correctness/provenance issues, generated artifact inconsistencies, incomplete QA, and a failed required check.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:25 - New AgentVersion facts have no evidence/provenance links.

The added records carry release dates, upstream tags, package names, CLI commands, summaries, and release-note claims, but their edges only contain version_of. The PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior same-date tracker drops, including upstream-current-2026-07-17.yaml, add EvidenceSource nodes that reference each new AgentVersion and product. Without equivalent provenance, Atlas evidence/catalog consumers cannot audit these claims.

Fix: add the 2026-08-04 EvidenceSource shard following the 2026-07-17 pattern, and reference every new AgentVersion plus its product/source.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - Codex uses a divergent ID shape.

The Codex record is agent-version:codex@0.146.0, while the surrounding generated records use agentVersion:<agent>:<slugified-version>. ID-based lookup, generated references, and future evidence links can miss this logical version or create duplicate records.

Fix: normalize the ID to the catalog stable ID convention and update any references/evidence records.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor ID does not match its versionRange slug.

The Cursor record uses agentVersion:cursor:changelog-2026-08-03, but versionRange is 2026-08-03-changelog. That breaks the common agentVersion:<agentId>:<slugified versionRange> shape unless this exception is intentional, documented, and tested.

Fix: use the canonical slug for versionRange, or add an explicit tested exception.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed artifact set.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs, or rename/narrow the field so downstream consumers do not misinterpret it.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are not committed.

The notes say release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but neither directory exists in the PR head.

Fix: commit those referenced body artifacts or remove/update the note.

  1. CI/QA is not merge-ready.

gh pr checks currently shows Docs QA as failed, and the newly dispatched QA run did not produce a terminal passing live-stack verdict. QA dispatch 31061059080 completed, but the nested Live Stack run 31061208047 timed out with Compute Matrix passed and Build All still in progress. The QA comment is #1620 (comment).

Fix: get required PR checks green and obtain a terminal passing Live Stack verdict after the graph fixes.

Missing Guardrails

The current checks did not catch the missing evidence-source shard or the ID-shape issues. Please add metadata verification that every upstream-current AgentVersion has matching evidence coverage and that AgentVersion IDs follow the stable generated form, with explicit exceptions only where intentional.

Risk Assessment

Risk level: risk:high.

  • Unbacked release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require EvidenceSource coverage and reference/evidence links for every upstream version record before merge.
  • Canonical-ID consumers may fail to resolve Codex/Cursor versions or produce duplicate logical nodes. Mitigation: normalize IDs and add a catalog-wide ID consistency test.
  • Generated artifact contract drift can mislead downstream automation. Mitigation: make summary.json internally consistent and add a tracker artifact path consistency check.
  • QA remains inconclusive and Docs QA is red. Mitigation: wait for all PR checks and focused live-stack QA to reach terminal success before merge.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: in progress after Babysitter polling timeout. The predefined QA process waited 20 minutes, but the GitHub Actions run had not reached a terminal conclusion.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061260206

Focused matrix:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 bridged-hooks bp predefined
claude foundry-gpt55 interactive bp create
hermes foundry-gpt55 ni vanilla predefined
gemini google-gemini31 bridged-interactive vanilla predefined

Current job status at timeout:

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts. This matrix focuses on babysitter-plugin predefined/create catalog consumers, bridged-hooks transport coverage, and vanilla adapter baselines for adversarial review coverage.

Overall verdict: not passed yet. No failing live-stack job had been reported when the QA process timed out, but the run is still active and must be checked for final conclusions.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / timed out while queued.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061275508

The workflow was dispatched for adversarial QA of PR #1620 against agent-versions/daily-2026-08-04, but the Babysitter QA wait window expired after 20 minutes before the live-stack run reached a terminal verdict. This is not a passing QA result yet.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Matrix tested:

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 bridged-hooks bp predefined
claude foundry-gpt55 interactive bp create
hermes foundry-gpt55 ni vanilla predefined
gemini google-gemini31 bridged-interactive vanilla predefined

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined/create graph/catalog consumers, bridged-hooks plugin integration, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Re-check the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061264056

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA. The Babysitter QA process waited through its polling window; the run moved from queued to in_progress, but did not reach a terminal conclusion.

Job Result
Compute Matrix pass
Build All in progress

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Overall verdict: not passed yet. No scenario failures were available at the time of this comment, but final live-stack conclusions are still pending.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / still running after the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31061260206

Focus: adversarial QA for the Atlas agent-version graph metadata update, with attention to catalog/evidence/source coverage and generated artifact consistency.

Tested matrix

Agent Model Mode Install Process mode
codex google-gemini31 interactive bp predefined
codex google-gemini31 interactive bp create
claude foundry-gpt55 bridged-hooks bp predefined
codex google-gemini31 ni vanilla predefined
hermes foundry-gpt55 ni vanilla predefined

Current job status at timeout

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Overall verdict: not passed yet. The workflow did not reach a terminal result within the QA process timeout; scenario jobs had not started by the final poll. Follow the linked run for final job conclusions before treating this QA as green.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion update. The review found two blockers, two major issues, and no terminal passing QA verdict.

Blockers

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:25 - New AgentVersion facts are not linked to evidence.

The new AgentVersion records only add edges.version_of; there is no sourced_from or evidence/reference edge for the release-date, version, CLI-command, source-package, release-note, and summary claims. The matching packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml shard is also absent at the PR head. Prior upstream tracker drops, for example the 2026-07-17 shard, include EvidenceSource nodes that reference each added AgentVersion and agent.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding product/agent, and rerun Atlas metadata/build validation.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - The Codex AgentVersion ID uses a different namespace and shape.

The Codex record uses agent-version:codex@0.146.0 while the surrounding records use the agentVersion:<agent>:<slugified-version> shape, for example agentVersion:amp:0-0-1785819659-g30d128 and agentVersion:codex-sdk:7-4-0. ID-based lookups, evidence references, generated indexes, and future relationships can miss this node or create duplicate logical versions.

Fix: normalize the Codex ID to the canonical AgentVersion shape expected by the catalog generator and update any evidence or references added for this version.

Major Issues

  • artifacts/agent-version-tracker/summary.json:322 underreports generated changed files. It lists only the new AgentVersion YAML, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json. Include all committed tracker artifacts or narrow the field name/meaning.

  • artifacts/agent-version-tracker/summary.json:329 points to artifacts/agent-version-tracker/release-notes/ and artifacts/agent-version-tracker/issues/, but those directories are not present in the PR. Commit the referenced artifacts or update the summary note to match the committed output.

QA

I dispatched qa-dispatch.yml for PR #1620 as run 31061064605. The dispatcher completed successfully and launched nested Live Stack run 31061260206, but the QA result posted to the PR is not passed yet / in progress after timeout. At timeout, Compute Matrix had passed and Build All was still running. This is not a terminal passing QA verdict.

Risk Assessment

Risk level: risk:high.

  • Unbacked release claims can enter Atlas and later be treated as trusted catalog facts. Mitigation: require EvidenceSource records and evidence references for every new AgentVersion before merge.
  • Consumers using canonical AgentVersion IDs may fail to resolve the new Codex release or may create duplicate references. Mitigation: normalize IDs and add an ID consistency check derived from agentId and versionRange.
  • Generated tracker artifact metadata can drift from the actual committed artifacts. Mitigation: fix changedFiles and artifact-path notes, then add a generated-artifact consistency check.
  • QA remains inconclusive. Mitigation: rerun QA after the data/provenance fixes and wait for a terminal passing Live Stack verdict.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas agent-version graph update. The review found no command-injection/secret/security blocker in the changed data files, but it found a blocking graph provenance issue, multiple major correctness/artifact issues, and QA is not green.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:3 - New AgentVersion facts are added without evidence coverage.

The PR adds 16 AgentVersion records with release dates, version ranges, CLI commands, release notes URLs, summaries, and assimilation notes. Each record only has a version_of edge. The PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml, and the new nodes do not carry evidenceSourceIds or matching Claim records. Prior tracker drops such as upstream-current-2026-07-17.yaml add same-date EvidenceSource records that reference each new version and product. Without equivalent provenance, Atlas catalog/evidence consumers cannot audit these release claims.

Fix: add the 2026-08-04 EvidenceSource shard following the 2026-07-17 pattern, reference every new AgentVersion plus its product/source, and wire the claims through the graph's established evidence model.

Major Findings

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs, or rename/narrow the field so downstream consumers do not treat it as a complete PR artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories that are not present.

The notes say release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but those directories are not in the PR head.

Fix: commit the referenced body artifacts or update the summary so it only describes files that actually exist.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - The Codex record uses an ID shape that SDK-generated lookup paths do not generate.

The new Codex node is agent-version:codex@0.146.0, while packages/atlas/src/catalog/sdk.ts generates AgentVersion IDs as agentVersion:<agentId>:<slugified versionRange>, which would be agentVersion:codex:0-146-0. ID-based lookup, generated references, and future evidence links can miss this version unless the exception is intentional and explicitly handled.

Fix: normalize the generated record to the SDK-generated shape or add a documented alias/migration path, then test it.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - The Cursor record ID does not match its versionRange slug.

The Cursor node uses agentVersion:cursor:changelog-2026-08-03, but versionRange is 2026-08-03-changelog; the SDK slugifier would generate agentVersion:cursor:2026-08-03-changelog.

Fix: rename the node to the generated slug form or add an explicit tested exception.

Missing Guardrails

  • packages/atlas/src/catalog/catalog.test.ts:101 only checks AgentVersion ID alignment for Copilot. Generalize it across all AgentVersion nodes so generated upstream drops cannot drift by agent.
  • Add metadata verification that every upstream-current-* AgentVersion record has matching evidence coverage.
  • Add a tracker summary consistency check that validates every path mentioned in summary.json exists in the committed artifact tree.

QA

QA Dispatch run 31061110695 completed successfully, but it launched Live Stack run 31061260206, which did not reach a terminal passing verdict within the QA process timeout. At timeout, Compute Matrix had passed, Build All was still in progress, and scenario jobs had not started. The QA result was posted at #1620 (comment). This is not a passing QA result.

Risk Assessment

Risk level: risk:high

  • Unprovenanced release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require EvidenceSource or Claim coverage for every new upstream version record before merge.
  • SDK and generated-reference consumers may fail to resolve Codex/Cursor versions or create duplicate logical nodes. Mitigation: normalize IDs or add explicit aliases, and enforce catalog-wide ID consistency in tests.
  • Generated tracker artifacts can mislead downstream automation and reviewers. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA remains inconclusive. Mitigation: rerun focused Live Stack QA after the graph fixes and wait for terminal success before merge.

@a5c-ai

a5c-ai Bot commented Aug 6, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas agent-version graph update. The review found a blocker, several major correctness/provenance issues, a failed approach check, and QA did not produce a terminal passing verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:3 - New AgentVersion facts are not backed by same-date evidence.

The PR adds 16 AgentVersion records with evidence-bound facts such as versionRange, releasedAt, releaseNotesUrl, cliCommand, summaries, and release highlights, but each record only has a version_of edge. The PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. The prior upstream-current-2026-07-17 batch includes EvidenceSource records that reference each added AgentVersion and product. Without equivalent provenance, Atlas catalog/evidence consumers can ingest these release claims without an auditable source.

Fix: add a 2026-08-04 EvidenceSource shard for every upstream/package source and connect/reference the exact new AgentVersion IDs before merge.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:129 - Codex uses an ID shape that does not match the SDK-generated form.

packages/atlas/src/catalog/sdk.ts builds agent-version IDs as agentVersion:<agentId>:<slugified versionRange>. For agentId=codex and versionRange=0.146.0, that produces agentVersion:codex:0-146-0, but this PR adds agent-version:codex@0.146.0. ID-based lookups, future evidence links, and generated references can miss this version or create a duplicate logical node.

Fix: normalize the ID to the SDK-generated form and update references/evidence records accordingly.

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor ID does not match its versionRange slug.

The Cursor record uses agentVersion:cursor:changelog-2026-08-03, but versionRange is 2026-08-03-changelog. The SDK-generated form would be agentVersion:cursor:2026-08-03-changelog.

Fix: use the canonical slug for the versionRange, or add an explicit tested exception if this reversal is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed artifact set.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but the PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs, or rename/narrow the field so downstream consumers do not treat it as a full changed-file manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are not committed.

The notes say release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR head only contains summary.json and upstream-targets-and-latest.json in that artifact directory.

Fix: commit the referenced release-note and issue body artifacts, or remove/update the note to match the artifact set actually included in the PR.

  1. Missing validation guardrails let this class of issue pass the main checks.

The PR adds a generated graph shard but does not add or extend checks that would fail on missing evidence-source coverage, missing direct evidence/claims, summary artifact self-inconsistency, or SDK ID drift for AgentVersion records.

Fix: add validation for upstream-current AgentVersion evidence coverage, ID alignment with the SDK slug convention, and tracker summary path consistency.

QA

I dispatched qa-dispatch.yml for PR #1620 as run 31061102769. The dispatcher completed successfully, but the nested Live Stack QA verdict was incomplete/not passed within the QA wait window. Comments were posted with the observed nested run states, including https://github.com/a5c-ai/babysitter/pull/1620#issuecomment-5199289592 and https://github.com/a5c-ai/babysitter/pull/1620#issuecomment-5199289355.

The PR branch also had failing CI history: the latest PR CI run I inspected had Docs QA failed while the main lint/test/package job passed.

Risk Assessment

Risk level: risk:high.

  • Unbacked release/version claims can enter Atlas and be treated as trusted catalog facts. Mitigation: require EvidenceSource coverage and reference/evidence links for every upstream version record before merge.
  • Canonical-ID consumers may fail to resolve Codex/Cursor versions or produce duplicate logical nodes. Mitigation: normalize IDs and add a catalog-wide ID consistency test.
  • Generated artifact contract drift can mislead downstream automation. Mitigation: make summary.json internally consistent and add a tracker artifact path consistency check.
  • QA remains inconclusive/not passed. Mitigation: rerun focused Live Stack QA to a terminal passing verdict after the graph fixes.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138600758

The workflow was dispatched against agent-versions/daily-2026-08-04 with checkout input ref=agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata, provenance/evidence-source coverage, SDK AgentVersion ID consistency, tracker artifact summary consistency, and affected catalog/evidence consumers. The Babysitter QA process waited through its polling window; the run was still in_progress at timeout.

Job Status Conclusion
Compute Matrix completed success
Build All in_progress pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined/create catalog/evidence consumers, Codex ID-sensitive coverage, bridged-hooks plugin integration, and representative vanilla adapter baselines.

Overall verdict: not passed yet. No terminal live-stack conclusion was available at timeout; follow the linked run for final job results before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138615415

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620, focused on Atlas AgentVersion graph metadata, evidence-source/provenance coverage, SDK AgentVersion ID consistency, tracker artifact summary consistency, and affected catalog/evidence consumers. The Babysitter QA process waited through its polling window, but the workflow remained queued and did not reach a terminal conclusion.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"gemini","model":"google-gemini31","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating this QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138679513

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA. The Babysitter QA process waited through its polling window; the workflow remained queued overall. Compute Matrix completed successfully, but Build All was still queued at timeout, so no live-stack scenario jobs reached terminal conclusions.

Job Status Conclusion
Compute Matrix completed success
Build All queued pending

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined/create graph/catalog consumers, Codex bridged-hooks plugin integration, and representative vanilla adapter baselines across Gemini/Foundry providers.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 7, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31138694128

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA with checkout input ref=agent-versions/daily-2026-08-04. The Babysitter QA process waited through its polling window, but the run remained queued and did not reach a terminal conclusion.

Job Result
Compute Matrix pass
Build All queued

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so this focused matrix targets BP predefined/create graph/catalog consumers, bridged-hooks plugin integration, and representative vanilla adapter baselines.

Overall verdict: not passed yet. No scenario failures were available because the live-stack jobs did not start within the QA polling window. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230747158

The workflow was dispatched against agent-versions/daily-2026-08-04 with checkout input ref=agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata, evidence-source coverage, generated tracker artifact consistency, and affected Atlas catalog consumers. The Babysitter QA process waited 20 minutes, but the GitHub Actions run was still in_progress.

Job Result
Compute Matrix pass
Build All in progress at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create catalog-consumer paths, Codex graph-read paths with Google provider coverage, Claude create/predefined coverage on Foundry, bridged-hooks plugin integration, and representative vanilla Hermes/Gemini adapter baselines without running the full cross-product.

Overall verdict: not passed yet. No live-stack scenario job conclusions were available by timeout; follow the linked run for final job results before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230756861

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited through its polling window. Build All and Compute Matrix completed successfully, but all selected live-stack scenario jobs were still queued when the wait window expired.

Job Result
Compute Matrix pass
Build All pass
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, bridged-hooks) queued
Live Stack (ubuntu-latest-l, vanilla, hermes/gpt-5.5, non-interactive) queued
Live Stack (ubuntu-latest-l, vanilla, gemini-cli/gemini-3.5-flash, bridged-interactive) queued
Live Stack (ubuntu-latest-l, bp/create, claude-code/gpt-5.5, interactive) queued
Live Stack (ubuntu-latest-l, bp/predefined, claude-code/gpt-5.5, interactive) queued
Live Stack (ubuntu-latest-l, bp/predefined, codex/gemini-3.5-flash, interactive) queued
Live Stack (ubuntu-latest-l, bp/create, codex/gemini-3.5-flash, interactive) queued

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets Atlas AgentVersion graph/catalog consumers through BP predefined/create paths, Codex bridged-hooks plugin integration, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found a blocking provenance gap, multiple generated-artifact consistency issues, a failed approach check, and no passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:2 - New AgentVersion facts are added without same-date evidence-source coverage.

The PR adds 16 AgentVersion records with release claims such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes. The PR head does not include packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml; a contents lookup for that path returned 404. Prior upstream-current drops include same-date EvidenceSource records that reference each added AgentVersion and agent/product. Without this shard, Atlas can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the versionRange slug convention.

The Cursor record ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. SDK-generated IDs, future evidence references, and claim subjects can reasonably derive agentVersion:cursor:2026-08-03-changelog, causing missed lookups or duplicate logical nodes.

Fix: rename the Cursor node to agentVersion:cursor:2026-08-03-changelog, or add an explicit documented and tested exception.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not interpret it as a complete generated-artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are absent from the PR.

The note says release-note bodies are under artifacts/agent-version-tracker/release-notes/ and issue bodies are under artifacts/agent-version-tracker/issues/, but the PR-head directory contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note to describe only artifacts actually present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The existing test filters only Copilot nodes, so generated upstream-current AgentVersion records for other agents can drift from the ID/versionRange convention and still pass local verification. This PR demonstrates the gap with Cursor's ID/versionRange mismatch.

Fix: generalize ID-alignment validation for generated upstream-current AgentVersion records, with explicit allowlisted historical exceptions where intentional. Add evidence coverage validation for upstream-current shards.

QA

I dispatched fresh QA via qa-dispatch.yml as run 31230605517. The dispatcher completed, but the resulting Live Stack QA did not produce a terminal passing verdict. The latest PR QA comment reports nested Live Stack run 31230756861 as incomplete / not passed within the 20-minute wait window: Compute Matrix and Build All passed, but selected live-stack scenario jobs were still queued at timeout. Current PR checks also show Docs QA failed.

Risk Assessment

Risk level: risk:high.

  • Unprovenanced release/version facts can enter Atlas and be treated as trusted catalog data. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion before merge and rerun Atlas metadata/build verification.
  • ID drift can split one logical AgentVersion across graph-node IDs, SDK-generated IDs, future evidence references, and claim subjects. Mitigation: normalize Cursor's ID or document/test an exception, then add generated-record ID consistency checks.
  • Generated tracker metadata can mislead reviewers or downstream automation about produced artifacts. Mitigation: fix changedFiles and absent artifact-path notes, then add a tracker summary consistency check.
  • QA is not green. Mitigation: resolve failed Docs QA and rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230804232

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window. The run moved from queued to in progress near the end, but did not reach a terminal verdict before timeout.

Job Result
Compute Matrix pass
Build All in progress at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths, a bridged-hooks plugin lane, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

GitHub did not accept a formal request-changes review from this actor, so I am posting the decision as a comment.

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found a blocking provenance gap, multiple generated-artifact consistency issues, a failed approach check, and no passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:2 - New AgentVersion facts are added without same-date evidence-source coverage.

The PR adds 16 AgentVersion records with release claims such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes. The PR head does not include packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior upstream-current drops include same-date EvidenceSource records that reference each added AgentVersion and agent/product. Without this shard, Atlas can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the versionRange slug convention.

The Cursor record ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts synthesizes AgentVersion IDs as agentVersion:<agentId>:<slugified versionRange> for agent references, capability support, claims, topology, and UI lookups. The generated ID would be agentVersion:cursor:2026-08-03-changelog, so future evidence references and SDK ID-based surfaces can miss or duplicate this logical version.

Fix: rename the Cursor node to agentVersion:cursor:2026-08-03-changelog, or add an explicit documented and tested exception.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not interpret it as a complete generated-artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories that are absent from the PR.

The note says release-note bodies are under artifacts/agent-version-tracker/release-notes/ and issue bodies are under artifacts/agent-version-tracker/issues/, but the PR-head directory contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note to describe only artifacts actually present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The existing test filters only Copilot nodes, so generated upstream-current AgentVersion records for other agents can drift from the ID/versionRange convention and still pass local verification. This PR demonstrates the gap with Cursor's ID/versionRange mismatch.

Fix: generalize ID-alignment validation for generated upstream-current AgentVersion records, with explicit allowlisted historical exceptions where intentional. Add evidence coverage validation for upstream-current shards.

QA

I dispatched fresh QA via qa-dispatch.yml as run 31230604935. The dispatcher wrapper completed successfully, but it did not establish a passing live-stack verdict. The latest PR QA comment reports nested Live Stack run 31230756861 as incomplete / not passed within the 20-minute wait window: Compute Matrix and Build All passed, but selected live-stack scenario jobs were still queued at timeout. Current PR checks also show Docs QA failed.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or artifact directory referenced by generated notes exists in the committed artifact tree.
  • Add catalog-wide/generated-record ID consistency checks with documented exceptions.

Risk Assessment

Risk level: risk:high.

  • Unprovenanced release/version facts can enter Atlas and be treated as trusted catalog data. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion before merge and rerun Atlas metadata/build verification.
  • ID drift can split one logical AgentVersion across graph-node IDs, SDK-generated IDs, future evidence references, and claim subjects. Mitigation: normalize Cursor's ID or document/test an exception, then add generated-record ID consistency checks.
  • Generated tracker metadata can mislead reviewers or downstream automation about produced artifacts. Mitigation: fix changedFiles and absent artifact-path notes, then add a tracker summary consistency check.
  • QA is not green. Mitigation: resolve failed Docs QA and rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31230832259

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. At the Babysitter polling cutoff, the workflow had moved from queued to in progress: Compute Matrix had passed and Build All was still running. No live-stack scenario jobs had reported terminal conclusions yet.

Job Status Conclusion
Compute Matrix completed success
Build All in_progress -

Matrix tested:

[{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},{"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts used by catalog/evidence workflows. The matrix targets BP predefined/create paths for catalog-consuming process flows across Codex and Claude, includes a bridged-hooks BP lane for plugin/hook integration, and adds a vanilla Hermes non-interactive baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 8, 2026

Copy link
Copy Markdown
Contributor Author

GitHub did not accept a formal request-changes review from this actor, so I am posting the decision as a comment.

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found one blocker, four major issues, a failed approach check, and QA did not produce a terminal passing verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New AgentVersion facts are not backed by same-date evidence coverage.

The PR adds 16 AgentVersion records with release/version facts such as versionRange, currentVersion, releasedAt, sourcePackage / upstreamReleaseTag, releaseNotesUrl, cliCommand, summaries, release highlights, and assimilation notes. The new records only have version_of edges, and the PR does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml.

Prior upstream-current drops, including packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-07-17.yaml, include EvidenceSource records that reference each added AgentVersion and agent/product. Without the same 2026-08-04 evidence shard, Atlas catalog/evidence consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID plus its product/agent, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the SDK-generated ID convention.

The Cursor node ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts synthesizes AgentVersion subject IDs as agentVersion:<agentId>:<slugified versionRange> for capability/evidence lookup paths, so this record's generated ID would be agentVersion:cursor:2026-08-03-changelog. That can split references between the graph node and SDK-generated subject IDs.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add an explicit documented/tested exception if the changelog-prefix form is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not treat it as a complete artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories absent from the PR.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR-head artifact directory contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note so it only describes artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The current ID-alignment test filters only Copilot nodes. This PR adds generated upstream AgentVersion records but does not generalize the guardrail to generated upstream records or evidence coverage. Existing historical exceptions may need an allowlist, but generated drops need validation that catches ID/evidence drift before review.

Fix: add validation for generated upstream-current AgentVersion ID and evidence coverage, with documented allowlisted exceptions where intentional.

QA

I dispatched qa-dispatch.yml for PR #1620 against agent-versions/daily-2026-08-04 as run 31230629569. After 25 one-minute polls, the dispatcher was still in_progress with no conclusion, and its qa job was still in the trigger step. Treating this as inconclusive/not passed. The PR checks also currently show Docs QA failed while Lint, Tests, Package, Workspace Coverage, and Observer Dashboard passed.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or directory referenced by generated summary notes exists in the committed artifact tree.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Atlas can ingest release/version claims without first-class provenance. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion before merge and verify evidence lookup for the new IDs.
  • AgentVersion ID drift can split graph-node references from SDK-generated subject IDs. Mitigation: normalize Cursor's ID or document/test an exception before merge.
  • Generated tracker metadata can mislead reviewers or automation. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA is not green. Mitigation: fix Docs QA and rerun focused live-stack QA to terminal success after the graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286708292

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process polled 20 times at roughly one-minute intervals. The workflow did not reach a terminal conclusion before timeout.

Job Result
Compute Matrix pass
Build All queued at timeout

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts consumed by catalog/evidence workflows. This matrix targets BP predefined/create catalog-consuming flows across Codex and Claude, includes a bridged-hooks BP lane for plugin/hook integration, and adds vanilla Hermes and Gemini baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked workflow for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286708602

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited for 20 minutes; the run remained queued overall. Compute Matrix completed successfully, but Build All was still queued and no live-stack scenario jobs had started by timeout.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create catalog-consuming process flows, a Codex bridged-hooks BP lane for plugin/hook integration, and one vanilla Hermes non-interactive adapter baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286714260

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the workflow did not reach a terminal verdict. Compute Matrix passed; Build All was still queued, and no live-stack scenario jobs had started.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: PR #1620 changes Atlas AgentVersion graph metadata and generated tracker artifacts, so the matrix targets BP predefined/create catalog-consuming process paths across Codex and Claude, adds a Codex bridged-hooks BP lane for plugin/hook integration, and includes a Hermes vanilla non-interactive adapter/provider baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286715705

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited through its polling window, but the live-stack workflow did not reach a terminal verdict. Compute Matrix completed successfully; Build All was still queued at timeout, so no live-stack scenario jobs had terminal conclusions yet.

Job Result
Compute Matrix pass
Build All queued at timeout

Matrix tested:

[{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},{"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},{"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},{"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},{"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths for Codex and Claude, a BP bridged-hooks lane, and representative vanilla adapter baselines for Hermes and Gemini.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286733702

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the run remained queued. Compute Matrix completed successfully; Build All had not started before timeout.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths, a bridged-hooks BP lane for plugin/hook integration, and representative vanilla adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286725203

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. The Babysitter QA process waited through its polling window, but the workflow remained queued and did not reach a terminal verdict.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a bridged-hooks BP lane for plugin/hook integration, and adds a vanilla Hermes non-interactive baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286734794

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of the Atlas AgentVersion graph metadata and generated tracker artifacts. At the polling cutoff, the workflow had not reached a terminal verdict and was still queued.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create catalog-consuming paths across Codex and Claude, a BP bridged-hooks plugin lane, and representative vanilla Hermes/Gemini adapter baselines without running the full cross-product.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31286733222

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 minutes. The run remained queued at timeout; Compute Matrix had completed successfully and Build All was still queued, so no live-stack scenario jobs had reached terminal conclusions.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds representative Hermes/Gemini vanilla adapter baselines.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

GitHub did not accept a formal request-changes review from this actor (Review Can not request changes on your own pull request), so I am posting the adversarial review decision as a comment.

Adversarial Review Decision: Changes Requested

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found one blocker, four major issues, a failed approach check, and no passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New AgentVersion facts are not backed by same-date evidence coverage.

The PR adds a 2026-08-04 batch of AgentVersion records with release/version claims, source packages/tags, release notes URLs, CLI commands, summaries, release highlights, and assimilation notes, but PR head does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior upstream tracker batches, including upstream-current-2026-07-17.yaml, add EvidenceSource nodes that reference the corresponding AgentVersion IDs and agent/product IDs. Without the new evidence shard, Atlas catalog/evidence consumers can ingest unprovenanced release facts.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and related agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the SDK-derived ID convention. The graph node is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog; SDK-generated subject IDs use the slugified versionRange, so references can split. Fix by renaming it to agentVersion:cursor:2026-08-03-changelog or adding a documented/tested exception.

  2. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports the committed tracker output. It lists only the AgentVersion YAML, but this PR also modifies summary.json and upstream-targets-and-latest.json. Include all generated tracker outputs or narrow/rename the field.

  3. artifacts/agent-version-tracker/summary.json:329 - The summary points to absent artifact directories. PR head contains only summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker, not release-notes/ or issues/. Commit those artifacts or update the note.

  4. packages/atlas/src/catalog/catalog.test.ts:103 - Existing validation only protects Copilot ID alignment. Generalize generated/upstream AgentVersion ID alignment validation and add evidence coverage validation for upstream-current shards.

QA

I dispatched qa-dispatch.yml as run 31286613202. Polls 1-24 showed the qa job still in_progress, stuck in Run a5c-ai/babysitter/packages/adapters/triggers@staging; poll 25 hit the GitHub installation API rate limit at 2026-08-09T01:03:07Z. No terminal passing QA verdict was obtained, so QA is inconclusive/not passed. The PR CI rollup observed before rate limit also showed Docs QA failed.

Missing Guardrails

  • Validate that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Validate that every path or artifact directory referenced by generated tracker notes exists.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Catalog provenance regression: release/version claims can enter Atlas without first-class EvidenceSource records. Mitigation: require same-date EvidenceSource coverage and rerun Atlas metadata/build verification before merge.
  • AgentVersion ID drift can split graph node IDs from SDK-generated subject IDs. Mitigation: normalize Cursor's ID or document/test an exception, then add generated-record ID checks.
  • Generated tracker metadata can mislead reviewers or automation. Mitigation: make summary.json internally consistent and add path-existence validation.
  • QA is not green. Mitigation: resolve failed Docs QA and rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial review decision: changes requested.

I cannot approve this PR as a complete Atlas AgentVersion tracker update. The review found one blocker, four major issues, a failed approach check, and no passing QA verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:2 - New AgentVersion facts lack same-date EvidenceSource coverage.

The PR adds 16 AgentVersion records with release/version facts, but does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. The PR-head contents lookup returns 404 for that expected shard, while the prior upstream-current-2026-07-17.yaml evidence shard includes EvidenceSource records that reference each added AgentVersion and agent/product. Without the same 2026-08-04 evidence shard, Atlas can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID plus its agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the versionRange slug convention.

The Cursor node ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog at line 240. SDK/catalog surfaces synthesize AgentVersion IDs from agentId plus slugified versionRange, so this can split graph-node references from generated subject IDs, future evidence links, and UI lookups.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add an explicit documented and tested exception if the reverse ordering is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports committed tracker output.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker outputs in changedFiles, or rename/narrow the field so downstream consumers do not treat it as a complete generated-artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories absent from the PR.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR-head artifact directory contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note so it only describes artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Validation only guards Copilot AgentVersion ID alignment.

The current ID-alignment test filters only Copilot nodes. This PR adds generated upstream AgentVersion records but does not generalize the guardrail to generated upstream records or evidence coverage, so the Cursor mismatch can pass existing validation.

Fix: add validation for generated upstream-current AgentVersion ID and evidence coverage, with documented allowlisted exceptions where intentional.

QA

I dispatched qa-dispatch.yml for PR #1620 against agent-versions/daily-2026-08-04 as run 31286610082. Polls 1-23 showed the dispatcher still in_progress with no conclusion; the qa job was stuck in the trigger step and updatedAt stayed at 2026-08-09T00:38:31Z. Polls 24-25 hit GitHub installation API rate limits, so no terminal QA verdict was available within the 25-minute process window. Treating QA as inconclusive/not passed. Current PR checks also show Docs QA failed.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or directory referenced by generated summary notes exists in the committed artifact tree.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Atlas can ingest release/version claims without first-class provenance. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion before merge and verify evidence lookup for the new IDs.
  • AgentVersion ID drift can split graph-node references from SDK-generated subject IDs. Mitigation: normalize Cursor's ID or document/test an exception before merge.
  • Generated tracker metadata can mislead reviewers or automation. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA is not green. Mitigation: fix Docs QA and rerun focused live-stack QA to terminal success after the graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR. It adds Atlas AgentVersion facts without the matching provenance shard, leaves a generated ID drift that can break SDK subject lookups, has generated artifact metadata inconsistencies, and QA did not produce a passing terminal verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New upstream AgentVersion facts are not backed by same-date EvidenceSource records.

The PR adds 16 AgentVersion records with version/release facts, but it does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior upstream-current drops, including upstream-current-2026-07-17.yaml, include same-date EvidenceSource nodes that reference each added AgentVersion and product. Without the 2026-08-04 evidence shard, Atlas catalog/evidence consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding product/agent, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the SDK slug convention.

The Cursor node ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts synthesizes IDs as agentVersion:${agentId}:${slugify(versionRange)} for references, capability support, claims, topology, and UI lookups, so the generated subject ID would be agentVersion:cursor:2026-08-03-changelog.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add a documented and tested exception if the prefix form is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports committed tracker outputs.

changedFiles lists only the new AgentVersion YAML, but the PR also changes summary.json itself and upstream-targets-and-latest.json. Downstream automation or reviewers treating changedFiles as a generated-artifact manifest will miss committed tracker outputs.

Fix: include all committed tracker output paths in changedFiles, or rename/narrow the field so consumers do not interpret it as complete.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories absent from the PR.

The note says release-note bodies are under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR head contains only summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note to describe only artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Existing validation only guards Copilot AgentVersion ID alignment.

The current ID-alignment test filters only Copilot nodes, so generated upstream-current AgentVersion records for other agents can drift from the ID/versionRange convention and still pass local verification. This PR's Cursor record demonstrates the gap.

Fix: generalize generated upstream AgentVersion ID-alignment validation, add explicit allowlisted historical exceptions where intentional, and add evidence coverage validation for upstream-current shards.

QA

I dispatched qa-dispatch.yml against agent-versions/daily-2026-08-04 for PR #1620 as run 31286613166. Polls 1-24 showed the dispatcher still in_progress in the triggers action after checkout. The final poll hit the GitHub installation API rate limit before a terminal conclusion could be read. Treating this as inconclusive/not passed.

Missing Guardrails

  • Add metadata validation that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Add tracker summary validation that every path or directory referenced by generated summary notes exists in the committed artifact tree.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Atlas catalog provenance regression: release/version facts can be consumed without linked EvidenceSource nodes. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion and rerun Atlas metadata/build verification before merge.
  • AgentVersion ID drift: graph nodes can split from SDK-generated subject IDs, capability support, claims, topology, and UI lookups. Mitigation: normalize Cursor's ID or add a documented/tested exception before merge.
  • Generated tracker metadata inconsistency: downstream automation and reviewers can be misled about which artifacts were produced and where evidence bodies live. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA uncertainty: no terminal passing QA verdict was obtained. Mitigation: rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes and resolve any red checks before merge.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31345217920

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 one-minute polls. The run remained queued at timeout; Compute Matrix completed successfully and Build All was still queued, so no live-stack scenario jobs reached terminal conclusions.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds a Hermes vanilla adapter baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31345251972

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 minutes. The workflow remained queued at timeout; Compute Matrix completed successfully, Build All was still queued, and no live-stack scenario jobs had reached terminal conclusions.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds Hermes/Gemini vanilla adapter baselines.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31345210522

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 minutes. The run remained queued at timeout; Compute Matrix completed successfully, Build All was still queued, and no live-stack scenario jobs had started.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds representative Hermes/Gemini vanilla adapter baselines.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31345263549

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 minutes. The run remained queued at timeout; Compute Matrix completed successfully, Build All was still queued, and no live-stack scenario jobs had started.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds a Hermes vanilla adapter baseline.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31345254060

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 minutes. The run remained queued at timeout; Compute Matrix had completed successfully and Build All was still queued, so no live-stack scenario jobs reached terminal conclusions.

Job Status Conclusion
Compute Matrix completed success
Build All queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds representative Hermes/Gemini vanilla adapter baselines.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31345256835

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process polled for 20 iterations; every successful poll showed the run still queued. The final status snapshot was blocked by the GitHub installation API rate limit, so no terminal live-stack verdict was obtained.

Job Status Conclusion
Live Stack workflow run queued throughout 20-poll wait window incomplete
Final status snapshot blocked by GitHub installation API rate limit rate_limited

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds representative Hermes/Gemini vanilla adapter baselines.

Overall verdict: not passed yet. Follow the linked workflow run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31345256492

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 one-minute polls. The run remained queued for every poll, so no live-stack scenario jobs reached a terminal conclusion within the wait window. The final job-detail lookup hit the GitHub installation API rate limit at 2026-08-10T01:04:57Z.

Job Status Conclusion
Live Stack workflow 31345256492 queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds representative Hermes/Gemini vanilla adapter baselines for the Atlas AgentVersion metadata and generated tracker artifact changes.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Live-stack QA

Result: incomplete / not passed within the 20-minute QA wait window.

Run: https://github.com/a5c-ai/babysitter/actions/runs/31345268029

The workflow was dispatched against agent-versions/daily-2026-08-04 for adversarial QA of PR #1620. The Babysitter QA process waited 20 minutes. The run remained queued through the last successful poll at 2026-08-10T01:04:17Z; the final status snapshot hit the GitHub installation API rate limit at 2026-08-10T01:05:18Z, so no live-stack scenario jobs reached terminal conclusions during this process.

Job Status Conclusion
Live Stack workflow queued -

Matrix tested:

[
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"codex","model":"google-gemini31","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"claude","model":"foundry-gpt55","mode":"interactive","install":"bp","live":true,"process_mode":"create"},
  {"agent":"codex","model":"google-gemini31","mode":"bridged-hooks","install":"bp","live":true,"process_mode":"predefined"},
  {"agent":"hermes","model":"foundry-gpt55","mode":"ni","install":"vanilla","live":true,"process_mode":"predefined"},
  {"agent":"gemini","model":"google-gemini31","mode":"bridged-interactive","install":"vanilla","live":true,"process_mode":"predefined"}
]

Coverage rationale: this matrix targets BP predefined/create graph/catalog consumer paths across Codex and Claude, includes a Codex bridged-hooks BP lane, and adds representative Hermes/Gemini vanilla adapter baselines.

Overall verdict: not passed yet. Follow the linked run for final job conclusions before treating QA as green.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR. It adds Atlas AgentVersion facts without the matching same-date provenance shard, leaves an ID convention mismatch that can split SDK subject lookups, has generated artifact metadata inconsistencies, and QA did not produce a terminal passing verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New upstream AgentVersion facts are not backed by same-date EvidenceSource records.

The PR adds 16 AgentVersion records with version/release/source facts, but it does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. The previous upstream tracker batch includes a same-date EvidenceSource shard that references each added AgentVersion and product/agent. Without this shard, Atlas catalog/evidence consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the upstream-current-2026-07-17.yaml pattern, reference every new AgentVersion plus its related agent/product, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the SDK slug convention.

The node ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog at line 240. packages/atlas/src/catalog/sdk.ts builds AgentVersion subject IDs as agentVersion:${agentId}:${slugify(versionRange)} for capability support, claims, topology, and UI references, so the SDK-convention ID would be agentVersion:cursor:2026-08-03-changelog.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add a documented and tested exception if the reverse-order ID is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports committed tracker outputs.

changedFiles lists only the new AgentVersion YAML, but the PR also modifies artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker output paths in changedFiles, or rename/narrow the field so consumers do not interpret it as a complete generated-artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories absent from the PR.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR-head artifact tree contains only summary.json and upstream-targets-and-latest.json.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note so it only describes artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Existing ID validation only guards Copilot.

The current test filters only AgentVersion nodes with agentId === "copilot", so generated upstream-current records for other agents can drift from the SDK ID convention and still pass local verification. This PR's Cursor record demonstrates the gap.

Fix: generalize generated/upstream AgentVersion ID-alignment validation, with explicit allowlisted historical exceptions where intentional, and add validation that upstream-current AgentVersion records have matching evidence coverage.

QA

I dispatched qa-dispatch.yml against agent-versions/daily-2026-08-04 for PR #1620 as run 31345081836. The 25-sample polling window ended with the dispatcher still in_progress, stuck in Run a5c-ai/babysitter/packages/adapters/triggers@staging after checkout. No terminal passing QA verdict was available, so QA is inconclusive/not passed. Current PR checks also show Docs QA failed.

Missing Guardrails

  • Validate that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Validate that every generated upstream AgentVersion ID matches the SDK slug convention, with explicit documented exceptions only where intentional.
  • Validate that summary notes and artifact manifests only reference paths present in the committed artifact tree.

Risk Assessment

Risk level: risk:high.

  • Atlas catalog provenance regression: release/version facts can be consumed without linked EvidenceSource records. Mitigation: require same-date evidence coverage for every new upstream AgentVersion and verify evidence lookup for the new IDs before merge.
  • AgentVersion ID drift: graph nodes can split from SDK-generated subject IDs, capability support, claims, topology, and UI lookups. Mitigation: normalize Cursor's ID or add a documented/tested exception before merge.
  • Generated tracker metadata inconsistency: downstream automation and reviewers can be misled about which artifacts were produced and where evidence bodies live. Mitigation: make summary.json internally consistent and add path-existence validation.
  • QA uncertainty: no terminal passing QA verdict was obtained, and Docs QA is red. Mitigation: fix Docs QA and rerun focused live-stack QA to terminal success after the graph/provenance/artifact fixes.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

GitHub rejected the formal request-changes review from this actor: Review Can not request changes on your own pull request. Posting the adversarial review decision as a comment.

Adversarial Review Decision: Changes Requested

I cannot approve this PR. It adds Atlas AgentVersion facts without the matching provenance shard, leaves generated metadata inconsistencies, has an ID-alignment risk, and QA did not produce a passing terminal verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New AgentVersion facts lack same-date EvidenceSource coverage.

The PR adds the 2026-08-04 AgentVersion batch, but PR head does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. Prior upstream-current batches, including upstream-current-2026-07-17.yaml, include EvidenceSource nodes that reference each added AgentVersion and product. Without this shard, Atlas catalog/evidence consumers can ingest release/version facts without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion and corresponding product/agent, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion id does not align with versionRange. The node id is agentVersion:cursor:changelog-2026-08-03 while versionRange is 2026-08-03-changelog at line 240. Rename it to the slugified convention or document and test an explicit exception.

  2. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports committed tracker outputs. It lists only the AgentVersion YAML, but this PR also changes summary.json and upstream-targets-and-latest.json. Include all committed tracker outputs or narrow the field semantics.

  3. artifacts/agent-version-tracker/summary.json:329 - The summary points to artifact directories absent from PR head. It says release-note bodies live under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but PR head contains only summary.json and upstream-targets-and-latest.json there. Commit the referenced artifacts or update the note.

  4. packages/atlas/src/catalog/catalog.test.ts:103 - Existing validation only guards Copilot ID alignment. Generalize generated upstream AgentVersion ID-alignment validation and add evidence coverage validation for upstream-current shards.

QA

I dispatched qa-dispatch.yml as run 31345106255. Poll 1 was queued; polls 2 through 24 remained in progress in Run a5c-ai/babysitter/packages/adapters/triggers@staging; poll 25 hit the GitHub installation API rate limit. No terminal passing QA verdict was obtained, so QA is inconclusive/not passed. Current PR checks also show Docs QA failed.

Risk Assessment

Risk level: risk:high.

  • Catalog provenance regression: new release/version facts can be consumed without linked EvidenceSource nodes. Mitigation: require same-date evidence coverage and verify evidence lookup before merge.
  • AgentVersion ID drift: graph nodes can split from SDK/generated subject IDs and future evidence links. Mitigation: normalize the Cursor ID or add a documented/tested exception before merge.
  • Generated tracker metadata inconsistency: reviewers and automation can be misled about produced artifacts. Mitigation: make summary.json internally consistent and validate referenced paths.
  • QA uncertainty: no terminal passing QA verdict was obtained and Docs QA is red. Mitigation: resolve CI and rerun focused live-stack QA to terminal success after the data fixes.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Adversarial Review Decision: Changes Requested

I cannot approve this PR. It adds Atlas AgentVersion facts without the matching provenance shard, leaves generated metadata inconsistencies, lacks guardrails for the new generated records, and QA did not produce a terminal passing verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New upstream AgentVersion facts are not backed by same-date EvidenceSource records.

The PR adds a 2026-08-04 batch of AgentVersion records, but it does not add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. A PR-head contents lookup for that expected shard returned 404, while prior upstream-current drops include same-date evidence-source shards, including upstream-current-2026-07-17.yaml. Without the 2026-08-04 evidence shard, Atlas catalog/evidence consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the established upstream-current pattern, reference every new AgentVersion ID and corresponding product/agent, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the SDK-derived versionRange slug.

The new Cursor node is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog. packages/atlas/src/catalog/sdk.ts derives AgentVersion subject IDs as agentVersion:${agentId}:${slugify(versionRange)} for claims, capability, topology, and lookup surfaces, so consumers can synthesize agentVersion:cursor:2026-08-03-changelog unless this exception is explicitly handled.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add a documented and tested exception if the reverse-order changelog ID is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports committed tracker outputs.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json. Downstream automation or reviewers treating changedFiles as a generated-artifact manifest will miss committed outputs.

Fix: include all committed tracker output paths in changedFiles, or rename/narrow the field so it is not interpreted as complete.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories absent from the PR.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but the PR only includes summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note to describe only artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Existing validation only guards Copilot AgentVersion ID alignment.

The current ID-alignment test filters only Copilot nodes, so generated upstream AgentVersion records for Cursor and similar agents can drift from the SDK ID convention and still pass verification. This PR demonstrates that gap.

Fix: generalize generated upstream AgentVersion ID-alignment validation and add evidence coverage validation for upstream-current shards, with explicit allowlisted exceptions where intentional.

QA

I dispatched qa-dispatch.yml for PR #1620 against agent-versions/daily-2026-08-04 as run 31345099816. Polling showed the qa job queued, then in_progress, with checkout completed and Run a5c-ai/babysitter/packages/adapters/triggers@staging still active from 2026-08-10T00:42:04Z through the last successful poll at 2026-08-10T01:04:18Z. The final poll at 2026-08-10T01:05:18Z hit the GitHub installation API rate limit, so no terminal passing QA verdict was obtained. Treating QA as inconclusive/not passed.

Missing Guardrails

  • Validate that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Validate that generated tracker summaries only reference committed artifact paths/directories.
  • Generalize generated-record ID consistency checks, with explicit documented exceptions where intentional.

Risk Assessment

Risk level: risk:high.

  • Atlas catalog provenance regression: release/version facts can be consumed without linked EvidenceSource nodes. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion and rerun Atlas metadata/build verification before merge.
  • AgentVersion ID drift: graph nodes can split from SDK-generated subject IDs, claims, topology, capability support, and UI lookups. Mitigation: normalize Cursor's ID or add a documented/tested exception before merge.
  • Generated tracker metadata inconsistency: downstream automation and reviewers can be misled about which artifacts were produced and where evidence bodies live. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA uncertainty: no terminal passing QA verdict was obtained. Mitigation: rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes and resolve any red checks before merge.

@a5c-ai

a5c-ai Bot commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

GitHub refused a formal request-changes review from this actor (Review Can not request changes on your own pull request), so I am posting the adversarial review decision as a comment.

Adversarial review decision: changes requested.

I cannot approve this PR as a complete Atlas AgentVersion tracker update. It adds new Atlas release/version facts without matching same-date provenance, leaves a generated ID drift that can split SDK/catalog lookups, has generated artifact metadata inconsistencies, and QA did not produce a passing terminal verdict.

Blocker

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:1 - New upstream AgentVersion facts are not backed by same-date EvidenceSource records.

The PR adds the 2026-08-04 AgentVersion shard, starting at line 2, but PR head does not include packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml. The prior 2026-07-17 upstream-current batch includes same-date EvidenceSource records that reference the corresponding AgentVersion IDs and products. Without the 2026-08-04 evidence shard, Atlas catalog/evidence consumers can ingest release/version claims without first-class provenance.

Fix: add packages/atlas/graph/catalog-meta/evidence-sources/upstream-current-2026-08-04.yaml following the 2026-07-17 pattern, reference every new AgentVersion ID and corresponding product/agent, then rerun npm run verify:metadata and npm run build --workspace=@a5c-ai/atlas.

Major Findings

  1. packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml:233 - Cursor AgentVersion ID does not match the versionRange slug convention.

The Cursor node ID is agentVersion:cursor:changelog-2026-08-03, while versionRange is 2026-08-03-changelog at line 240. packages/atlas/src/catalog/sdk.ts and packages/atlas/src/catalog/data.ts synthesize IDs as agentVersion:${agentId}:${slugify(versionRange)} for lookups and capability joins, so this can split graph-node references from generated subject IDs.

Fix: rename the node to agentVersion:cursor:2026-08-03-changelog, or add a documented and tested exception if the reverse-order ID is intentional.

  1. artifacts/agent-version-tracker/summary.json:322 - changedFiles underreports committed tracker outputs.

changedFiles lists only packages/atlas/graph/agent-stack/agent-versions/upstream-current-2026-08-04.yaml, but this PR also changes artifacts/agent-version-tracker/summary.json and artifacts/agent-version-tracker/upstream-targets-and-latest.json.

Fix: include all committed tracker output paths in changedFiles, or rename/narrow the field so downstream consumers do not treat it as a complete generated-artifact manifest.

  1. artifacts/agent-version-tracker/summary.json:329 - The summary references artifact directories absent from PR head.

The note says release-note bodies are stored under artifacts/agent-version-tracker/release-notes/ and issue bodies under artifacts/agent-version-tracker/issues/, but PR head contains only summary.json and upstream-targets-and-latest.json under artifacts/agent-version-tracker.

Fix: commit the referenced release-note and issue body artifacts, or update the generated note so it only describes artifacts present in the PR.

  1. packages/atlas/src/catalog/catalog.test.ts:103 - Existing validation only guards Copilot AgentVersion ID alignment.

The current ID-alignment test filters only Copilot nodes. This PR adds generated upstream AgentVersion records for many agents but does not generalize the guardrail to generated upstream records or evidence coverage, so the Cursor mismatch and missing evidence shard can pass local verification.

Fix: generalize upstream-current AgentVersion ID-alignment validation, add documented allowlisted exceptions where intentional, and add validation that each generated upstream-current AgentVersion has matching EvidenceSource coverage.

QA

I dispatched qa-dispatch.yml for PR #1620 against agent-versions/daily-2026-08-04. Dispatch returned run 31345109521. Polls 1-24 did not produce a terminal passing verdict for that run; nearby same-workflow runs were mostly queued or completed skipped. Poll 25 hit the GitHub installation API rate limit at 2026-08-10T01:05:31Z, and the final run view also failed with HTTP 403 rate limit at 2026-08-10T01:06:31Z. Treating QA as inconclusive/not passed.

Missing Guardrails

  • Validate that every generated upstream-current-* AgentVersion record has matching evidence coverage.
  • Validate that every path or directory referenced by generated tracker notes exists in the committed artifact tree.
  • Add generated-record ID consistency checks, with explicit documented exceptions where necessary.

Risk Assessment

Risk level: risk:high.

  • Atlas catalog provenance regression: release/version facts can be consumed without linked EvidenceSource nodes. Mitigation: require same-date EvidenceSource coverage for every new upstream AgentVersion and rerun Atlas metadata/build verification before merge.
  • AgentVersion ID drift: graph nodes can split from SDK-generated subject IDs, capability support, claims, topology, and UI lookups. Mitigation: normalize Cursor's ID or add a documented/tested exception before merge.
  • Generated tracker metadata inconsistency: downstream automation and reviewers can be misled about which artifacts were produced and where evidence bodies live. Mitigation: make summary.json internally consistent and validate referenced artifact paths.
  • QA uncertainty: no terminal passing QA verdict was obtained. Mitigation: rerun focused live-stack QA to terminal success after graph/provenance/artifact fixes and resolve any red checks before merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants