25 KiB
Agent Challenge Harness V2 Implementation Plan
For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (
- [ ]) syntax for tracking.
Goal: Build a manifest-driven generic challenge runner with layered prompts, none|skills|all instruction profiles, one-hour trial ceilings, normalized OpenCode tool/token evidence, policy findings, and generated audit reports.
Architecture: Challenge manifests and templates are data; executable code lives in examples/agent_challenges. The runner creates one isolated trial workspace, composes base/profile/challenge prompts, executes OpenCode from that workspace, preserves raw JSONL, and delegates bounded normalization to focused metrics/policy/report modules. Automatic task classification and validity are provisional; manual audit remains authoritative.
Tech Stack: Python 3.14, dataclasses/Pydantic, PyYAML, subprocess, hashlib, pytest, existing OpenCode JSONL format.
File Structure
- Create
examples/agent_challenges/models.py: manifest/profile/result DTOs. - Create
examples/agent_challenges/manifests.py: YAML loading and path resolution. - Create
examples/agent_challenges/prompts.py: layered prompt rendering/provenance. - Create
examples/agent_challenges/metrics.py: JSONL event/tool/token normalization. - Create
examples/agent_challenges/policy.py: observed path categorization and validity. - Modify
examples/agent_challenges/workspace.py: profile-aware workspace creation and instruction bundle copying. - Modify
examples/agent_challenges/opencode_io.py: command uses persisted rendered prompt. - Modify
examples/agent_challenges/runner.py: manifest/profile orchestration, trial cwd, 3,600-second default, normalized evidence. - Modify
examples/agent_challenges/reports.py: generated report with evidence sections. - Modify
examples/agent_challenges/audit.py: consume v2 task/validity fields. - Create
examples/agent_challenges/run_trials.py: central CLI. - Create
examples/agent_challenges/save_trial_report.py: central report CLI. - Create
examples/agent_challenges/save_manual_audit.py: central audit CLI. - Create
examples/agent_challenges/base-prompt.md: invariant benchmark rules. - Create
examples/agent_challenges/profile-prompts/{none,skills,all}.md: policy fragments. - Create
tests/examples/test_agent_challenge_harness_v2.py: generic behavior tests.
Task 1: Define Manifest And Profile Models
Files:
-
Create:
examples/agent_challenges/models.py -
Create:
examples/agent_challenges/manifests.py -
Create:
tests/examples/test_agent_challenge_harness_v2.py -
Step 1: Write failing manifest/profile tests
Create tests/examples/test_agent_challenge_harness_v2.py with imports and:
from __future__ import annotations
import json
from pathlib import Path
import pytest
from examples.agent_challenges.manifests import load_challenge_manifest
from examples.agent_challenges.models import InstructionProfile
def _write_manifest(root: Path) -> Path:
(root / "workspace_template").mkdir(parents=True)
(root / "challenge-prompt.md").write_text("Build it.\n", encoding="utf-8")
path = root / "challenge.yaml"
path.write_text(
"""\
version: 1
id: fixture
prompt: challenge-prompt.md
workspace_template: workspace_template
source:
id: local.fixture
root: source
module: ops
registry: registry
store_root: .wf_fixture_store
server:
config: wf.config.json
default_port: 8779
report:
required_fields: [value, run_failed]
success_assertions:
value: expected
run_failed: false
""",
encoding="utf-8",
)
return path
def test_load_challenge_manifest_resolves_paths(tmp_path: Path) -> None:
manifest_path = _write_manifest(tmp_path)
loaded = load_challenge_manifest(manifest_path)
assert loaded.manifest.id == "fixture"
assert loaded.root == tmp_path.resolve()
assert loaded.prompt_path == (tmp_path / "challenge-prompt.md").resolve()
assert loaded.workspace_template == (tmp_path / "workspace_template").resolve()
assert loaded.manifest.report.success_assertions == {
"value": "expected",
"run_failed": False,
}
def test_instruction_profiles_are_exactly_the_supported_conditions() -> None:
assert [profile.value for profile in InstructionProfile] == [
"none",
"skills",
"all",
]
def test_invalid_manifest_rejects_parent_traversal(tmp_path: Path) -> None:
path = _write_manifest(tmp_path)
text = path.read_text(encoding="utf-8").replace(
"workspace_template: workspace_template",
"workspace_template: ../outside",
)
path.write_text(text, encoding="utf-8")
with pytest.raises(ValueError, match="workspace_template"):
load_challenge_manifest(path)
- Step 2: Run tests and verify imports fail
uv run pytest tests/examples/test_agent_challenge_harness_v2.py -q
Expected: import failures for models/manifests.
- Step 3: Implement Pydantic manifest DTOs
Create examples/agent_challenges/models.py:
from __future__ import annotations
from enum import StrEnum
from pathlib import Path
from typing import Any
from pydantic import BaseModel, ConfigDict, Field
class InstructionProfile(StrEnum):
NONE = "none"
SKILLS = "skills"
ALL = "all"
class SourceManifest(BaseModel):
model_config = ConfigDict(extra="forbid")
id: str = Field(min_length=1)
root: str = Field(min_length=1)
module: str = Field(min_length=1)
registry: str = Field(min_length=1)
class ServerManifest(BaseModel):
model_config = ConfigDict(extra="forbid")
config: str = Field(min_length=1)
default_port: int = Field(ge=1, le=65535)
class ReportManifest(BaseModel):
model_config = ConfigDict(extra="forbid")
required_fields: list[str] = Field(default_factory=list)
success_assertions: dict[str, Any] = Field(default_factory=dict)
class ChallengeManifest(BaseModel):
model_config = ConfigDict(extra="forbid")
version: int = Field(ge=1)
id: str = Field(pattern=r"^[a-z][a-z0-9_-]*$")
prompt: str
workspace_template: str
source: SourceManifest
store_root: str
server: ServerManifest
report: ReportManifest
class LoadedChallenge(BaseModel):
model_config = ConfigDict(arbitrary_types_allowed=True)
manifest_path: Path
root: Path
prompt_path: Path
workspace_template: Path
source_root: Path
server_config: Path
manifest: ChallengeManifest
- Step 4: Implement safe manifest loading
Create examples/agent_challenges/manifests.py:
from __future__ import annotations
from pathlib import Path
import yaml
from .models import ChallengeManifest, LoadedChallenge
def _inside(root: Path, relative: str, *, field: str) -> Path:
candidate = (root / relative).resolve()
if not candidate.is_relative_to(root):
raise ValueError(f"challenge {field} must stay inside challenge directory")
return candidate
def load_challenge_manifest(path: Path) -> LoadedChallenge:
manifest_path = path.resolve()
root = manifest_path.parent
loaded = yaml.safe_load(manifest_path.read_text(encoding="utf-8"))
manifest = ChallengeManifest.model_validate(loaded)
prompt_path = _inside(root, manifest.prompt, field="prompt")
workspace_template = _inside(
root, manifest.workspace_template, field="workspace_template"
)
source_root = (root / manifest.source.root).resolve()
server_config = (root / manifest.server.config).resolve()
if not prompt_path.is_file():
raise ValueError(f"challenge prompt does not exist: {prompt_path}")
if not workspace_template.is_dir():
raise ValueError(f"challenge workspace_template does not exist: {workspace_template}")
return LoadedChallenge(
manifest_path=manifest_path,
root=root,
prompt_path=prompt_path,
workspace_template=workspace_template,
source_root=source_root,
server_config=server_config,
manifest=manifest,
)
- Step 5: Run manifest tests and commit
uv run pytest tests/examples/test_agent_challenge_harness_v2.py -q
git add examples/agent_challenges/models.py examples/agent_challenges/manifests.py tests/examples/test_agent_challenge_harness_v2.py
git commit -m "feat: add agent challenge manifests"
Task 2: Add Layered Prompts And Profile-Aware Workspaces
Files:
-
Create:
examples/agent_challenges/base-prompt.md -
Create:
examples/agent_challenges/profile-prompts/none.md -
Create:
examples/agent_challenges/profile-prompts/skills.md -
Create:
examples/agent_challenges/profile-prompts/all.md -
Create:
examples/agent_challenges/prompts.py -
Modify:
examples/agent_challenges/workspace.py -
Test:
tests/examples/test_agent_challenge_harness_v2.py -
Step 1: Add failing prompt/profile tests
Append tests that call compose_trial_prompt and prepare_v2_trial_workspace:
from examples.agent_challenges.prompts import compose_trial_prompt
from examples.agent_challenges.workspace import prepare_v2_trial_workspace
def test_challenge_prompt_is_identical_across_profiles(tmp_path: Path) -> None:
challenge = load_challenge_manifest(_write_manifest(tmp_path / "challenge"))
rendered = {
profile: compose_trial_prompt(
challenge,
profile=profile,
wf_command_prefix="uv run wf --config wf.config.json --local",
server_context="Local mode.",
workspace_path=tmp_path / profile.value,
)
for profile in InstructionProfile
}
assert {value.challenge_sha256 for value in rendered.values()} == {
rendered[InstructionProfile.NONE].challenge_sha256
}
assert len({value.rendered_sha256 for value in rendered.values()}) == 3
assert "report the exact blocker" in rendered[InstructionProfile.NONE].text
assert ".agent/skills" in rendered[InstructionProfile.SKILLS].text
assert "inspect broader repository" in rendered[InstructionProfile.ALL].text
def test_skills_profile_copies_bundle_but_none_does_not(tmp_path: Path) -> None:
challenge = load_challenge_manifest(_write_manifest(tmp_path / "challenge"))
bundle = ROOT / "examples/agent_challenges/instruction_bundles/workflow_cli.yaml"
none_workspace = prepare_v2_trial_workspace(
challenge,
profile=InstructionProfile.NONE,
model="model",
index=1,
workspaces_dir=tmp_path / "workspaces",
instruction_bundle=bundle,
)
skills_workspace = prepare_v2_trial_workspace(
challenge,
profile=InstructionProfile.SKILLS,
model="model",
index=2,
workspaces_dir=tmp_path / "workspaces",
instruction_bundle=bundle,
)
assert not (none_workspace.root / ".agent/skills").exists()
assert (skills_workspace.root / ".agent/skills/wf-cli/SKILL.md").is_file()
assert skills_workspace.instruction_files
Add ROOT = Path(__file__).resolve().parents[2] if it is not already present.
- Step 2: Create invariant and profile prompts
Create base-prompt.md with placeholders:
# Workflow Agent Challenge
Use the repository's public `wf` product path to complete the challenge below.
Do not replace the workflow lifecycle with a helper script that imports internal
workflow APIs. Preserve exact commands, failures, run ids, and evidence in your
final answer.
Use this command prefix:
{{wf_command_prefix}}
{{server_context}}
Your writable trial workspace is `{{workspace_path}}`. Write attempt files only
inside it. End with the challenge's requested YAML self-report. The self-report
will be checked against observed tool calls and manually audited.
Create profile fragments:
none.md:
## Instruction Profile: none
Use challenge files, `wf --help`, `wf schema`, validation, inspect, and bounded
trace commands. Do not read repository skills, docs, examples, tests, source,
prior trials, or prior stores. If public surfaces are insufficient, report the
exact blocker and finish the task as failed rather than reverse-engineering the
implementation.
skills.md:
## Instruction Profile: skills
Use the supplied skills under `.agent/skills/` plus public `wf` commands. Do not
read repository examples, tests, source, prior trials, or prior stores. If the
skills and public surfaces are insufficient, report the exact blocker rather
than reverse-engineering implementation code.
all.md:
## Instruction Profile: all
Start with the supplied skills and public docs. If genuinely blocked, you may
inspect broader repository docs, examples, tests, and source. Report what you
read; observed tool calls will also be retained for audit.
- Step 3: Implement prompt composition and hashes
Create prompts.py with a RenderedPrompt frozen dataclass, SHA-256 helper,
placeholder replacement, and concatenation in this order:
base prompt
profile fragment
challenge prompt
The dataclass must expose text, base_sha256, profile_sha256,
challenge_sha256, and rendered_sha256. Reject unresolved {{...}}
placeholders with ValueError.
- Step 4: Extend workspace preparation
Add a V2TrialWorkspace dataclass with:
root: Path
config_path: Path
rendered_prompt_path: Path
instruction_files: tuple[Path, ...]
Implement prepare_v2_trial_workspace to:
-
create the unique workspace;
-
copy the challenge template;
-
write the local config;
-
for
skillsandall, load the explicit bundle manifest andcopy2each file under.agent/skills/<destination>; -
for
none, copy no instruction files; -
never use symbolic links or junctions.
-
Step 5: Run profile tests and commit
uv run pytest tests/examples/test_agent_challenge_harness_v2.py -k "prompt or profile" -q
git add examples/agent_challenges/base-prompt.md examples/agent_challenges/profile-prompts examples/agent_challenges/prompts.py examples/agent_challenges/workspace.py tests/examples/test_agent_challenge_harness_v2.py
git commit -m "feat: add challenge instruction profiles"
Task 3: Normalize OpenCode Tool And Token Events
Files:
-
Create:
examples/agent_challenges/metrics.py -
Test:
tests/examples/test_agent_challenge_harness_v2.py -
Step 1: Add a realistic JSONL fixture test
Append a test using four JSON lines: step_start, successful tool_use for
read, failed tool_use for bash, and step_finish with token/cost data.
Assert:
metrics.step_count == 1
metrics.tool_call_count == 2
metrics.failed_tool_call_count == 1
metrics.tool_counts == {"bash": 1, "read": 1}
metrics.tokens.total == 120
metrics.tokens.input == 20
metrics.tokens.output == 30
metrics.tokens.reasoning == 10
metrics.tokens.cache_read == 60
metrics.cost == 0.01
metrics.tool_calls[0].tool == "read"
metrics.tool_calls[0].output_chars == 4000
len(metrics.tool_calls[0].output_preview) <= 500
- Step 2: Run and verify the metrics import fails
uv run pytest tests/examples/test_agent_challenge_harness_v2.py -k metrics -q
Expected: import failure for metrics.
- Step 3: Implement event parsing DTOs
In metrics.py, define frozen dataclasses:
TokenMetrics(total, input, output, reasoning, cache_read, cache_write)
ToolCallEvidence(ordinal, call_id, tool, status, title, input, metadata,
output_chars, output_preview, output_sha256, failed)
TrialMetrics(step_count, tool_call_count, failed_tool_call_count, tool_counts,
tokens, cost, unknown_event_count, tool_calls)
Implement extract_trial_metrics(stdout: str, *, preview_chars: int = 500):
-
parse each nonblank line independently as JSON;
-
preserve malformed line count as unknown evidence instead of aborting;
-
count
step_finishevents and sum numeric token fields; -
read nested cache
read/writefields; -
sum numeric
costvalues; -
normalize
tool_use.partandpart.statefields; -
hash full string output with SHA-256 and store only the bounded preview in normalized metrics;
-
retain unknown event count.
-
Step 4: Add JSON serialization helper
Implement metrics_payload(metrics: TrialMetrics) -> dict[str, Any] using
dataclasses.asdict, with tool-count keys sorted for deterministic output.
- Step 5: Run metrics tests and commit
uv run pytest tests/examples/test_agent_challenge_harness_v2.py -k metrics -q
git add examples/agent_challenges/metrics.py tests/examples/test_agent_challenge_harness_v2.py
git commit -m "feat: extract opencode trial metrics"
Task 4: Derive Policy Evidence From Tool Calls
Files:
-
Create:
examples/agent_challenges/policy.py -
Test:
tests/examples/test_agent_challenge_harness_v2.py -
Step 1: Add path-category and validity tests
Create tool evidence fixtures for reads of:
- current workspace file;
.agent/skills/wf-cli/SKILL.md;- repository
src/wf_cli/app.py; - repository
tests/wf_cli/test_app.py; - another trial workspace;
- prior
.wf_*store; - opaque
bashcommand.
Assert:
-
nonewith source read iscontaminatedand records the path; -
skillspermits copied skill reads but source/test reads contaminate; -
allpermits source/test reads and setsescalated_to_product_code=Truewithout contamination; -
opaque shell commands produce
unauditableonly when no stronger contamination result exists. -
Step 2: Implement policy DTO and categorization
Create:
class EvaluationValidity(StrEnum):
CLEAN = "clean"
CONTAMINATED = "contaminated"
UNAUDITABLE = "unauditable"
@dataclass(frozen=True, slots=True)
class PolicyEvidence:
validity: EvaluationValidity
disallowed_reads: tuple[str, ...]
escalated_to_product_code: bool
opaque_shell_commands: tuple[str, ...]
reads_by_category: dict[str, tuple[str, ...]]
Implement evaluate_policy(profile, tool_calls, *, workspace_root, repository_root, workspaces_root) using normalized tool names/inputs.
Recognize structured read/search inputs for tools named read, glob,
grep, list, and equivalent names already observed in saved OpenCode output.
Classify resolved paths into workspace, supplied skills, docs, examples, tests,
source, adjacent attempts, prior stores, or outside/unknown.
For shell tools, store the command text. Recognize only clear literal path
arguments; do not claim complete shell parsing. Commands that cannot be
classified remain in opaque_shell_commands.
- Step 3: Run policy tests and commit
uv run pytest tests/examples/test_agent_challenge_harness_v2.py -k policy -q
git add examples/agent_challenges/policy.py tests/examples/test_agent_challenge_harness_v2.py
git commit -m "feat: derive challenge policy evidence"
Task 5: Integrate Manifest Runner, Workspace CWD, And One-Hour Ceiling
Files:
-
Modify:
examples/agent_challenges/opencode_io.py -
Modify:
examples/agent_challenges/runner.py -
Create:
examples/agent_challenges/run_trials.py -
Test:
tests/examples/test_agent_challenge_harness_v2.py -
Step 1: Add runner integration tests with injected subprocess
Add tests that inject a fake run function and assert:
-
default timeout is
3600; -
cwdis the generated trial workspace; -
exactly one profile is present in config;
-
rendered prompt is persisted;
-
result JSON contains
instruction_profile, prompt hashes, metrics, policy, repository commit/dirty marker, and raw stdout/stderr; -
timeout preserves partial stdout and extracted partial metrics.
-
Step 2: Replace command construction input
Change TrialConfig/build_opencode_command so command construction receives
the already-rendered prompt text or reads rendered_prompt_path. Prompt
composition must not happen inside opencode_io.py.
Keep:
opencode run [--attach URL] <rendered prompt> --format json --model ... --variant ...
- Step 3: Refactor runner entrypoint
The central run_trials.py CLI accepts:
--challenge PATH
--instruction-profile none|skills|all
--model MODEL
--variant VARIANT
--trials N
--timeout-seconds 3600
--attach URL
--results-dir PATH
--workspaces-dir PATH
--instruction-bundle PATH
--server-url URL
--start-server
--server-port PORT
Require one explicit or default profile, never a profile list. Use one generated
workspace per trial and pass that workspace as cwd to subprocess.run.
- Step 4: Write normalized evidence after every terminal path
For success, nonzero process exit, timeout, and parse error:
-
write raw result JSON first;
-
extract metrics from available stdout;
-
evaluate policy from available tool calls;
-
write
metrics.jsonin the workspace; -
include prompt/bundle hashes and repository provenance;
-
call report generation;
-
never discard raw evidence if later stages fail.
-
Step 5: Run runner tests and direct help smoke
uv run pytest tests/examples/test_agent_challenge_harness_v2.py -k "runner or timeout" -q
uv run python examples/agent_challenges/run_trials.py --help
Expected: tests pass and help lists --instruction-profile and default timeout.
- Step 6: Commit runner integration
git add examples/agent_challenges/opencode_io.py examples/agent_challenges/runner.py examples/agent_challenges/run_trials.py tests/examples/test_agent_challenge_harness_v2.py
git commit -m "feat: run profiled agent challenge trials"
Task 6: Generate Evidence-Rich Reports And Central Audit CLIs
Files:
-
Modify:
examples/agent_challenges/reports.py -
Modify:
examples/agent_challenges/audit.py -
Create:
examples/agent_challenges/save_trial_report.py -
Create:
examples/agent_challenges/save_manual_audit.py -
Test:
tests/examples/test_agent_challenge_harness_v2.py -
Step 1: Add report snapshot assertions
Build a small result payload and assert generated Markdown contains:
Instruction profile: skills
Task outcome: success
Evaluation validity: contaminated
Duration
Observed token metrics
Tool calls by tool
Disallowed reads
Agent self-report discrepancies
Final agent answer
Manual audit: pending
Assert full tool outputs are absent and previews are bounded.
- Step 2: Separate task outcome from validity in reports
Update reports.py to render these sections in order:
- Trial identity/provenance.
- Task outcome and evaluation validity.
- Duration/token/cost metrics.
- Tool/command summary.
- Observed reads and policy findings.
- Agent self-report and observed discrepancies.
- Final agent answer.
- Manual-audit status.
Keep report_from_result and save_report_from_result_payload as generic
helpers, but remove browser-click assumptions.
- Step 3: Update manual audit payload
manual-audit.yaml must include:
task_outcome: success
evaluation_validity: clean
official_outcome: pass
auditor_notes: "..."
The manual command can override automatic task/validity fields but must retain
the automatic values under automatic_evidence for comparison.
- Step 4: Add central wrapper CLIs
save_trial_report.py delegates to reports.main; save_manual_audit.py
delegates to audit.main. Both must support direct execution from repo root and
display useful --help text.
- Step 5: Run report/audit tests and commit
uv run pytest tests/examples/test_agent_challenge_harness_v2.py -k "report or audit" -q
git add examples/agent_challenges/reports.py examples/agent_challenges/audit.py examples/agent_challenges/save_trial_report.py examples/agent_challenges/save_manual_audit.py tests/examples/test_agent_challenge_harness_v2.py
git commit -m "feat: report agent challenge evidence"
Task 7: Final Generic Harness Verification
Files:
-
Modify:
docs/current_roadmap.md -
Move after completion:
docs/superpowers/plans/2026-06-22-agent-challenge-harness-v2.mdtodocs/historical/superpowers/plans/2026-06-22-agent-challenge-harness-v2.md -
Step 1: Run focused tests and static checks
uv run pytest tests/examples/test_agent_challenge_harness_v2.py tests/examples/test_agent_challenge_skill_bundle.py -q
uv run ruff check examples/agent_challenges tests/examples/test_agent_challenge_harness_v2.py tests/examples/test_agent_challenge_skill_bundle.py
uv run ruff format --check examples/agent_challenges tests/examples/test_agent_challenge_harness_v2.py tests/examples/test_agent_challenge_skill_bundle.py
uv run basedpyright --level error examples/agent_challenges tests/examples/test_agent_challenge_harness_v2.py tests/examples/test_agent_challenge_skill_bundle.py
git diff --check
Expected: all pass; only accepted Windows CRLF warnings may appear.
- Step 2: Add roadmap completion note
- Completed: the generic agent challenge harness now supports data-driven
manifests, layered prompts, explicit `none|skills|all` profiles, one-hour hard
ceilings, normalized OpenCode tool/token evidence, policy findings, and
manual-audited reports.
- Step 3: Archive and commit the plan
New-Item -ItemType Directory -Force docs/historical/superpowers/plans | Out-Null
Move-Item docs/superpowers/plans/2026-06-22-agent-challenge-harness-v2.md docs/historical/superpowers/plans/2026-06-22-agent-challenge-harness-v2.md
git add docs/current_roadmap.md docs/superpowers/plans/2026-06-22-agent-challenge-harness-v2.md docs/historical/superpowers/plans/2026-06-22-agent-challenge-harness-v2.md
git commit -m "docs: record agent challenge harness v2"