11 KiB
Agent Challenge Evaluation Runbook
This runbook is for running and auditing external-agent trials against the
workflow product surface. It covers the shared harness in
examples/agent_challenges/, not a specific challenge implementation.
The goal is to measure whether an agent can use public workflow commands and instructions to build, deploy, and run a workflow. The harness is also a UX instrument: failed or contaminated trials often point to missing docs, confusing commands, or product gaps.
What The Harness Produces
Each trial writes four kinds of evidence:
- Raw result JSON in the challenge
results/directory. This is immutable raw evidence from the runner. - Machine report JSON beside the raw result, named
*.report.json. This is the bounded projection for analysis. - Human report Markdown beside the raw result, named
*.report.md. This is useful for reviewing results after workspace cleanup. - Human report Markdown inside the trial workspace, named
final-report.md. This is the file to read first during manual review.
Manual audits add one more file:
manual-audit.yamlinside the trial workspace. Re-running the audit command regenerates the workspace Markdown, result Markdown, and machine report projections without mutating the raw result JSON.
Run One Trial
Run from the repository root. Use --attach when an opencode server is already
running.
uv run python examples/agent_challenges/run_trials.py `
--challenge examples/agent_challenges/browser_click_challenge/challenge.yaml `
--instruction-profile skills `
--model opencode/deepseek-v4-flash-free `
--variant high `
--trials 1 `
--attach http://127.0.0.1:4096
For the report workflow challenge:
uv run python examples/agent_challenges/run_trials.py `
--challenge examples/agent_challenges/report_workflow_challenge/challenge.yaml `
--instruction-profile skills `
--model opencode/deepseek-v4-flash-free `
--variant high `
--trials 1 `
--attach http://127.0.0.1:4096
The runner prints summary JSON with the trial classification, result path, and
report paths. Read the corresponding final-report.md before trusting the
classification.
Use bounded concurrency when collecting larger samples:
uv run python examples/agent_challenges/run_trials.py `
--challenge examples/agent_challenges/report_workflow_challenge/challenge.yaml `
--instruction-profile skills `
--model opencode/deepseek-v4-flash-free `
--variant high `
--trials 5 `
--concurrency 2 `
--attach http://127.0.0.1:4096
--trials 5 --concurrency 2 creates five unique trial workspaces and lets at
most two OpenCode subprocesses run at once.
Run The Default Matrix
The Python matrix runner expands the bundled challenges, instruction profiles, and default model set, then schedules them with one global concurrency limit:
uv run python examples/agent_challenges/run_matrix.py `
--trials 5 `
--concurrency 2 `
--attach http://127.0.0.1:4096
The PowerShell helper is now only a convenience wrapper around the Python runner:
.\examples\agent_challenges\run_matrix.ps1 -Trials 5 -Concurrency 2
Avoid treating high-concurrency runs as the same dataset as sequential runs. Large concurrency can measure provider queueing, OpenCode timeouts, and local machine contention rather than workflow UX.
Instruction Profiles
Use profiles to separate product usability from instruction quality:
| Profile | Meaning | Typical Use |
|---|---|---|
none |
Base prompt plus challenge prompt only. | Tests discoverability with almost no agent instructions. |
skills |
Adds the workflow CLI skill bundle. | Tests the intended public agent instruction layer. |
all |
Allows broader docs/code exploration. | Tests whether the repository contains enough information to solve the task, but results are less clean. |
debug |
Adds the workflow CLI skill bundle and asks the agent to keep detailed UX notes before escalating beyond public surfaces. | Mines command, docs, schema, and validation friction; keep separate from normal benchmark scoring. |
The same model/challenge should be run across profiles when comparing the value of the instruction layer.
debug is opt-in and is not part of the default matrix. Use it when you want a
trial report to preserve failed commands, mistaken assumptions, confusing help
text, schema-shape problems, and any point where the agent became genuinely
blocked. A debug run may still complete the workflow, but its main value is the
ux_issues_found list in the self-report.
Suggested Matrix
Start small and grow only after the harness output is stable:
challenge in [browser_click, report_workflow]
profile in [none, skills, all]
model in [deepseek-v4-flash-free, mimo-v2.5-free, nemotron-3-ultra-free]
trials per cell = 3 to 5 while iterating, more for claims
Do not treat one successful run as a model-quality result. One run can be useful as a product UX finding, but not as an aggregate benchmark.
Manual Review Checklist
Open the trial workspace final-report.md and compare it to the raw transcript
when needed.
Check these items:
- Did the agent use the product path:
wf artifact create-from-planor draft commands,wf deploy saveorwf deploy create, andwf run start? - Did the run actually complete with the required output fields?
- Did the agent write a helper script that directly drives
WorkflowApior bypasses the CLI/server path? - Did the agent read implementation code under
src/or tests undertests/? - Did the agent read a ready-made solution, store internals, adjacent attempt, or generated workspace from another trial?
- Did the agent self-report those reads honestly in the YAML block?
- Did opaque shell commands hide important behavior that needs manual review?
- Did the report include a real
run_id, deployment id, and workflow file path?
Validity And Coverage
The report separates policy validity from policy coverage.
evaluation_validity answers whether the automatic evidence found a rule
violation:
clean: no observed disallowed reads or policy violations.contaminated: observed disallowed evidence, such as reading a ready-made solution or forbidden prior result.unauditable: reserved for missing or corrupt raw evidence.
policy_coverage answers how much of the evidence the automatic pass could
inspect:
complete: automatic policy checks could inspect the recorded tool evidence.partial: some behavior happened through opaque shell commands. This does not automatically invalidate the trial, but it requires manual review.
Manual audit is authoritative for final interpretation. A technically successful workflow can still be invalid as evaluation evidence if the agent copied from an existing solution or bypassed the public product path.
Save A Manual Audit
Use manual audit after reading the report and raw transcript.
uv run python examples/agent_challenges/save_manual_audit.py `
--from-result examples/agent_challenges/browser_click_challenge/results/opencode_deepseek-v4-flash-free-trial-001.json `
--manual-classification invalid `
--auditor codex `
--set-read product_code=true `
--set-read existing_solution=true `
--notes "Technical workflow run succeeded, but the trial is invalid because the agent inspected an existing solution."
Use --manual-classification pass when the run satisfies the challenge and no
disqualifying evidence is found. Use fail when the task was not completed.
Use invalid when the workflow ran but the evaluation is contaminated,
bypassed, or otherwise not usable as clean benchmark evidence.
Resume An Incomplete OpenCode Trial
When a result captures an OpenCode sessionID, the report includes an
OpenCode Resume section. Print the resume command with:
uv run python examples/agent_challenges/resume_trial.py `
--from-result examples/agent_challenges/browser_click_challenge/results/opencode_deepseek-v4-flash-free-trial-034.json
Run the resume and save a sidecar result with:
uv run python examples/agent_challenges/resume_trial.py `
--from-result examples/agent_challenges/browser_click_challenge/results/opencode_deepseek-v4-flash-free-trial-034.json `
--run
The resume command writes *.resume-001.json beside the original result and
does not mutate the original raw result. Use manual audit to decide whether the
resumed output completes the trial or only provides additional evidence.
Summarize Audited Results
After manual audits, generate a compact matrix table from the bounded report projections:
uv run python examples/agent_challenges/summarize_trials.py `
examples/agent_challenges/browser_click_challenge `
examples/agent_challenges/report_workflow_challenge
The table uses manual_audit.official_outcome when present, while keeping the
automatic task outcome, policy validity, duration, token count, attempt count,
and read flags visible. Use it as a working operator summary, not as a
statistical claim by itself.
Common Invalid Patterns
Mark the trial invalid or at least contaminated when any of these happen:
- The agent reads a complete existing solution, such as a fixture workflow plan for the same challenge.
- The agent writes a one-off Python runner that calls
WorkflowApidirectly instead of usingwfor the JSON-RPC server path. - The agent solves the browser/report task outside the workflow runtime.
- The agent uses prior trial artifacts or adjacent generated workspaces.
- The agent reads
.wf_*store internals under the current or another trial workspace innoneorskillsprofile. Use public commands such aswf draft inspect,wf artifact inspect,wf deploy validate, andwf run traceinstead. Store inspection is only acceptable inallprofile and should still be reported. - The agent reports
product_code: falseafter readingsrc/,tests/, or example implementation files needed to infer the answer.
Reading docs and skills is allowed unless the chosen profile says otherwise. Reading implementation code is not automatically a product failure, but it must be reported and usually makes the trial less useful as public-surface evidence.
What To Claim
Safe claims from small samples:
- A specific model run did or did not complete the challenge.
- A specific UX problem appeared, such as a confusing command name or missing schema documentation.
- The instruction layer helped or failed in a specific observed case.
Avoid stronger claims until the matrix has enough audited trials:
- Model A is better than Model B.
- The system generally reduces token usage.
- Agents can reliably author workflows without code reads.
- The benchmark is statistically meaningful.
Use the challenge evidence as product-design feedback first. Treat aggregate model benchmarking as a later result once trial counts and audit rules are stable.