This directory holds a deterministic, non-agentic end-to-end test of the
ado-aw execute (Stage 3) safe-output executor.
The agentic smoke suite in tests/safe-outputs/ exercises
the full Stage 1 → 2 → 3 pipeline shape, but it drives Stage 3 by having an
LLM agent (Stage 1) emit the safe output — so a failure can be the model's
fault and the suite is inherently flaky. It also only runs a small number of
omnibus pipelines rather than one per tool.
This suite removes the LLM from the loop. For every ADO-write safe output it:
- sets up preconditions deterministically via the ADO REST API,
- crafts the executor's
safe_outputs.ndjsoninput directly (fixed literal values), - runs the real
ado-aw executebinary (built from the checkout), - asserts the effect via the ADO REST API,
- cleans up every object it created,
and, on any failure, files a GitHub issue on the configured issue repository
and fails the build. AgentPlayground currently uses
jamesadevine/ado-aw-issues because a canonical-repository credential is not
available.
| File | Purpose |
|---|---|
azure-pipelines.yml |
Hand-authored ADO pipeline (daily schedule on main + path-filtered PR validation + manual). Builds ado-aw, builds the harness, runs it against AgentPlayground. |
README.md |
This file. |
The harness itself lives in
scripts/ado-script/src/executor-e2e/.
It is a test-only ado-script bundle: it is built by npm run build:executor-e2e to scripts/ado-script/test-bin/executor-e2e.js (a
gitignored, non-root path) and is deliberately excluded from the released
ado-script.zip (the release glob only packages ado-script/*.js, and the
executor-e2e dir is listed in NON_BUNDLE_DIRS in
src/__tests__/bundle-coverage.test.ts).
All deterministically-assertable ADO-write safe outputs plus the flagship
create-pull-request, and the four signal-only tools:
- Signals:
noop,missing-tool,missing-data,report-incomplete(no ADO write path; assert that the executor emits the expected status) - Work items:
create-work-item,update-work-item,comment-on-work-item,link-work-items,upload-workitem-attachment - Wiki:
create-wiki-page,update-wiki-page - PR:
add-pr-comment,reply-to-pr-comment,resolve-pr-thread,submit-pr-review,update-pr - Git:
create-branch,create-git-tag - Build:
add-build-tag,queue-build,upload-build-attachment,upload-pipeline-artifact - Flagship:
create-pull-requestcovers both a named additional checkout at<BUILD_SOURCESDIRECTORY>/<alias>andrepository: selfbeneath a non-Git multi-checkout root, with compiler-owned self repository identity taking precedence over trigger-scopedBUILD_REPOSITORY_*values. Theselfscenario supplies onlyADO_AW_SELF_REPOSITORY_NAME— matching what the compiler emits — so it also proves the executor resolves a repository from its name alone. - GitHub issues:
create-github-issue,set-github-issue-type, and the same-runtemporary_idhandoff between them. These are the only scenarios that assert against GitHub rather than ADO — see GitHub issue scenarios below.
Excluded (out of scope): none of the currently shipped safe outputs.
Coverage note. The signal scenarios (
noop,missing-tool,missing-data,report-incomplete) were previously exercised only by now-deleted per-tool agentic smoke pipelines. Adding them here closes the coverage gap while keeping the test deterministic.
create-github-issue and set-github-issue-type had zero runtime
coverage before these scenarios: their only proof was a wiremock unit test,
and the last thing exercising create-github-issue end to end
(smoke-failure-reporter) was removed by the smoke-suite rework.
| Scenario id | Tool | What it proves |
|---|---|---|
create-github-issue |
create-github-issue |
final title is title-prefix + the agent title; the body carries the agent text and the <!-- ado-aw --> traceability footer; config-injected static labels merge with allowed agent labels |
create-github-issue-label-denied |
create-github-issue |
allowed-labels is default-deny — an agent label outside the allowlist is rejected and no issue is filed |
set-github-issue-type |
set-github-issue-type |
a native issue type is applied to an existing issue |
set-github-issue-type-clear |
set-github-issue-type |
the documented issue_type: "" clear operation |
create-github-issue-temporary-id-handoff |
set-github-issue-type (with create-github-issue staged ahead of it) |
the same-run temporary_id handoff |
create-github-issue may mint a temporary_id, and
set-github-issue-type.issue_number accepts either a real number or that id.
The registry backing this (ExecutionContext::resolved_github_issues) is an
in-process Arc<Mutex<HashMap<…>>> that is never persisted, so the handoff is
only observable inside a single ado-aw execute invocation.
That matches production — a SafeOutputs job runs one ado-aw execute over the
whole safe_outputs.ndjson — but every other scenario here runs one entry per
invocation. The handoff scenario therefore uses the harness's priorEntries
hook (see Scenario.priorEntries in scripts/ado-script/src/executor-e2e/scenario.ts)
to stage create-github-issue as an extra NDJSON line ahead of its own, in the
same executor process. The assertion is that the set-github-issue-type record
reports the issue number create-github-issue actually filed.
Why this can't be split across jobs. Because the registry is per-process, putting
require-approvalon only one of the two tools would split Stage 3 into twoado-aw executeprocesses (SafeOutputsandSafeOutputs_Reviewed), and atemporary_idminted in one could not resolve in the other. The compiler rejects that configuration up front —validate_github_issue_outputs_configinsrc/compile/common.rsrequires both tools to have the same effectiverequire-approvalsetting, so the section-level default and a per-tool override are both accounted for.
Adding this coverage immediately found a real defect in Stage 3 config
handling, now fixed in ExecutionContext::get_tool_config
(src/safe_outputs/result.rs). It is recorded here because it is exactly the
class of bug the wiremock unit tests structurally could not see.
Stage 3 injects synthetic staged and require-approval keys into every
tool config (src/main.rs for --source; src/compile/custom_tools.rs for the
compiler-generated --resolved-config that production uses).
CreateGithubIssueConfig and SetGithubIssueTypeConfig are the only
safe-output configs declared #[serde(deny_unknown_fields)], and neither
declares a staged field — so deserialization failed, get_tool_config
swallowed the error via .ok().unwrap_or_default(), and the operator config
was silently replaced with Default::default().
Observable effects: target-repo ignored (Stage 3 failed outright on
non-GitHub-backed ADO builds), title-prefix never applied, static
labels/assignees dropped, allowed-labels emptied so default-deny rejected
every agent label, require-temporary-id unenforced, and
set-github-issue-type.allowed never gating anything — the last of which
failed open.
The unit tests missed it because they build an ExecutionContext directly with
a config map that has no staged key, i.e. a shape that never occurs in
production. The fix strips both orchestration keys and logs a warning instead of
silently defaulting; result.rs carries regression tests asserting an operator
config survives the injected keys.
Note that create-github-issue-label-denied deliberately matches only the
labels not in allowed-labels message. The alternative message
(no allowed-labels configured) is precisely what the executor emitted when the
config was dropped, so accepting both would have let the scenario pass either
way — the failure mode this suite exists to prevent.
GitHub has no REST endpoint to delete an issue. These scenarios are the only
ones in the suite that cannot tear down completely: cleanup() closes each
issue as not_planned instead.
Every scratch issue title embeds the standard ado-aw-det-$(Build.BuildId)-<id>
marker, so anything a cleanup misses is findable with a single search on the
scratch repository. Cleanup also does not depend solely on state captured in
assert() — when the executor filed an issue but the record came back
non-succeeded, assert() never runs, so cleanup falls back to an exact-title
search on that marker.
Because issues accumulate (closed, never deleted), point these scenarios at a scratch repository, not a canonical one.
| Variable | Meaning |
|---|---|
EXECUTOR_E2E_GITHUB_TOKEN |
Reused from failure-issue filing. It must now also carry Issues: write on the scratch repository, because these scenarios create, mutate, and close issues. |
EXECUTOR_E2E_SCENARIO_ISSUE_REPO |
Optional. owner/repo for scratch issues; falls back to EXECUTOR_E2E_ISSUE_REPO. Set it to keep scenario issues away from the failure-report repository. |
E2E_GITHUB_ISSUE_TYPE |
Optional. Forces a native issue-type name for environments where the token cannot read org metadata but the type is known to exist. |
There is deliberately no default repo for these scenarios: when neither
variable is set they skip rather than filing scratch issues onto
githubnext/ado-aw.
Some scenarios need optional infrastructure and skip (rather than fail) when it is not available:
queue-build— needs a target pipeline id inE2E_QUEUE_PIPELINE_ID.create-wiki-page/update-wiki-page— need a wiki in the project. The harness auto-discovers the first wiki; setE2E_WIKI_NAMEto force one. When no wiki exists, both skip.add-build-tag,upload-build-attachment,upload-pipeline-artifact— need a real current build (BUILD_BUILDID); they skip when run outside a pipeline.- All five GitHub issue scenarios — need
EXECUTOR_E2E_GITHUB_TOKENand a scratch repo (EXECUTOR_E2E_SCENARIO_ISSUE_REPOorEXECUTOR_E2E_ISSUE_REPO). They also skip when the token authenticates but cannot write issues on that repo, with the harness's auth diagnosis attached to the skip reason. set-github-issue-typeand, on the named-type path, the handoff — need a native issue type to exist. Issue types are an organisation-level construct (GET /orgs/{org}/issue-types) with no user-account equivalent, so a user-owned scratch repo can never expose one and these skip permanently there. The handoff scenario stays runnable by falling back to the documentedissue_type: ""clear operation;set-github-issue-type-clearskips only if GitHub rejects the clear outright.
Every object a scenario creates is prefixed
ado-aw-det-$(Build.BuildId)-<tool>. Cleanup runs unconditionally after each
scenario; the smoke-suite janitor (which prunes ado-aw-* artifacts) is the
backstop for anything a cleanup misses.
You need a write-capable ADO token (PAT) and a checkout-built binary:
cargo build --release --bin ado-aw
cd scripts/ado-script && npm ci && npm run build:executor-e2e && cd ../..
export SYSTEM_COLLECTIONURI="https://dev.azure.com/msazuresphere/"
export SYSTEM_TEAMPROJECT="AgentPlayground"
export SYSTEM_ACCESSTOKEN="<write-capable-PAT>"
export EXECUTOR_E2E_ADO_AW_BIN="$PWD/target/release/ado-aw"
export EXECUTOR_E2E_ADO_REPO="agent-definitions"
# Optional:
# export EXECUTOR_E2E_GITHUB_TOKEN="<fine-grained PAT: Issues rw on jamesadevine/ado-aw-issues>"
# export EXECUTOR_E2E_ISSUE_REPO="jamesadevine/ado-aw-issues"
# Optional: keep GitHub issue scenario scratch issues out of the failure-report repo
# export EXECUTOR_E2E_SCENARIO_ISSUE_REPO="<owner>/<scratch-repo>"
# Optional: force a native issue-type name (org-owned repos only)
# export E2E_GITHUB_ISSUE_TYPE="Bug"
# export E2E_QUEUE_PIPELINE_ID="<noop-target pipeline id>"
# Optional timeout tuning (milliseconds) for slow environments:
# export EXECUTOR_E2E_REST_TIMEOUT_MS=30000 # per ADO REST call (default 30000)
# export EXECUTOR_E2E_EXECUTE_TIMEOUT_MS=600000 # per `ado-aw execute` run (default 600000)
# export EXECUTOR_E2E_GIT_TIMEOUT_MS=300000 # per git subprocess call (default 300000)
node scripts/ado-script/test-bin/executor-e2e.jsBuild-scoped scenarios (add-build-tag, uploads) skip locally because there is
no current build. The harness exits non-zero if any scenario fails.
In https://dev.azure.com/msazuresphere/AgentPlayground:
Current registration: definition
2550in\executor-e2e, withE2E_QUEUE_PIPELINE_IDpointing at thequeue-targetdefinition registered fromqueue-target.yml.
-
Register the pipeline. New pipeline → GitHub through the
githubnextservice connection → existing YAML →tests/executor-e2e/azure-pipelines.yml. Place it in a\executor-e2efolder and skip the first run until variables are configured. In the live pull-request trigger settings, disable builds from forks and disable fork access to secrets/full tokens. Definition2550is audited bytests/smoke/trigger-policy.json. -
Grant the principal behind
agent-playground-writewrite access on theagent-definitionsrepo (Contribute, Create branch, Contribute to PRs) and on Build (add tags). The YAML maps its AAD token toSYSTEM_ACCESSTOKENthroughSC_WRITE_TOKEN. Seedocs/safe-output-permissions.mdif Stage 3 hits 401/403. -
Set the GitHub PAT secret on this pipeline only:
ado-aw secrets set EXECUTOR_E2E_GITHUB_TOKEN ` --org msazuresphere --project AgentPlayground ` --definition-ids <executor-e2e-pipeline-id> ` --value <fine-grained-pat-Issues-rw-on-jamesadevine/ado-aw-issues>
Do not place this token in a shared variable group.
The GitHub issue scenarios reuse this same token, so it needs Issues: write on the scratch repository — not just enough to file a failure report. When it can authenticate but not write, those scenarios skip with a diagnosis rather than failing the build.
-
Set
EXECUTOR_E2E_ISSUE_REPO=jamesadevine/ado-aw-issues. Confirm the target repository hasexecutor-e2e-failureandpipeline-failurelabels. (Optional) SetEXECUTOR_E2E_SCENARIO_ISSUE_REPOto a separate scratch repository so the GitHub issue scenarios do not accumulate closed issues alongside the failure reports.set-github-issue-typeand the named-type half of the handoff need a native issue type, which is an organisation-level construct. On a user-owned repo such asjamesadevine/ado-aw-issuesthey will skip on every run; pointEXECUTOR_E2E_SCENARIO_ISSUE_REPOat an org-owned repo with issue types defined to enable them. -
Set
E2E_QUEUE_PIPELINE_IDto thequeue-targetdefinition ID (registerqueue-target.ymlif it does not exist yet). It is a permanent, trigger-free, non-agentic pipeline that exists only to be queued, so this scenario no longer depends on the smoke suite's registration lifecycle. (Optional) SetE2E_WIKI_NAMEto enable the wiki scenarios. -
Trigger one manual run to seed the schedule.