feat: add opencode browser click challenge harness

- Challenge prompt requiring workflow/deployment/run evidence
- CLI harness (run_opencode_trials.py) for running N agent trials
- Classification: success, workflow_not_used, run_failed, timeout, parse_error, unknown
- 8 unit tests covering build/parse/classify/path logic
- README with usage docs and optional Playwright MCP attachment
- Evidence index and roadmap updated
- Plan archived to historical/
This commit is contained in:
lda
2026-06-15 02:32:54 +07:00 Verified
parent 4c3ce651df
commit 10dabca241
10 changed files with 1199 additions and 0 deletions
@@ -0,0 +1,729 @@
# Opencode Browser Click Challenge Harness Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Build a local evidence harness that runs the browser-click workflow challenge through `opencode run`, captures trial outputs, and classifies failures without changing product runtime code.
**Architecture:** The harness lives under `examples/agent_challenges/browser_click_challenge/` and treats opencode as an external executable. Tests do not call opencode; they cover prompt loading, command construction, JSONL/result parsing, and deterministic classification from captured text.
**Tech Stack:** Python stdlib (`argparse`, `json`, `subprocess`, `dataclasses`, `pathlib`, `time`), pytest, existing browser-click workflow example, optional Playwright MCP attachment configured by command-line flags.
---
## File Structure
- Create `examples/agent_challenges/browser_click_challenge/prompt.md`
- The exact prompt sent to agents.
- It must require workflow usage, deployment run, before/after snapshots, cleanup, and final summary.
- Create `examples/agent_challenges/browser_click_challenge/README.md`
- How to prepare the local server/config.
- How to run one trial and multiple trials.
- How to attach Playwright MCP when desired.
- How success/failure is classified.
- Create `examples/agent_challenges/browser_click_challenge/results/.gitignore`
- Keeps the output folder in git.
- Ignores generated trial outputs in the folder itself.
- Create `examples/agent_challenges/browser_click_challenge/run_opencode_trials.py`
- CLI harness.
- Builds `opencode run` command.
- Runs N trials.
- Writes one JSON artifact per trial.
- Classifies result text.
- Create `tests/examples/test_opencode_browser_click_challenge.py`
- Unit tests only.
- No live opencode calls.
- Modify `docs/add/evidence-index.md`
- Add the challenge harness as planned/evaluation evidence.
- Modify `docs/current_roadmap.md`
- Add completed evidence-harness bullet after implementation.
---
## Task 1: Add Prompt And README
**Files:**
- Create: `examples/agent_challenges/browser_click_challenge/prompt.md`
- Create: `examples/agent_challenges/browser_click_challenge/README.md`
- Create: `examples/agent_challenges/browser_click_challenge/results/.gitignore`
- Test: no pytest yet; this task is docs/assets only
- [ ] **Step 1: Create challenge folder**
Run:
```powershell
New-Item -ItemType Directory -Force examples\agent_challenges\browser_click_challenge\results
New-Item -ItemType File -Force examples\agent_challenges\browser_click_challenge\results\.gitignore
```
Expected: the folder exists.
- [ ] **Step 2: Create prompt**
Create `examples/agent_challenges/browser_click_challenge/prompt.md`:
```markdown
# Browser Click Workflow Challenge
Build and successfully run a workflow that:
1. Opens a browser page or local web page with a visible button.
2. Waits for a human click or performs a clearly simulated click.
3. Captures a before snapshot and an after snapshot.
4. Returns both snapshots as workflow output.
Use this repository's workflow product path. That means you should use the
`wf` CLI and/or `wf-rpc-server`, create or reuse a workflow deployment, and run
the deployment through the workflow API. Do not solve the challenge with only a
standalone Playwright/Python script.
The repository already includes a deterministic source example at:
```text
examples/browser_click_workflow/
```
You may inspect and use it. A successful final answer must include:
- the commands you ran,
- the deployment id,
- the run id if one was produced,
- evidence that `before.clicked` is `false`,
- evidence that `after.clicked` is `true`,
- whether any server/browser process remains running.
If something fails, report the exact command and error instead of hiding it.
```
- [ ] **Step 3: Create README**
Create `examples/agent_challenges/browser_click_challenge/README.md`:
````markdown
# Opencode Browser Click Challenge Harness
This harness runs agent trials against the browser-click workflow challenge.
It is evidence tooling, not product runtime code.
The deterministic workflow example is:
```text
examples/browser_click_workflow/
```
## One Trial
From the repository root:
```powershell
uv run python examples/agent_challenges/browser_click_challenge/run_opencode_trials.py `
--model opencode/mimo-v2.5-free `
--variant high `
--trials 1
```
Results are written to:
```text
examples/agent_challenges/browser_click_challenge/results/
```
## Optional Playwright MCP Attachment
If you want the agent to have browser-control tools, pass:
```powershell
--attach http://127.0.0.1:4096
```
Start that MCP/tool endpoint separately. For example, one possible MCP server
command is:
```json
{
"command": "npx",
"args": ["-y", "@playwright/mcp@latest"]
}
```
The baseline challenge does not require Playwright MCP. The score is based on
whether the agent used the workflow product path and produced the expected
workflow output.
## Classification
Each trial is classified as one of:
- `success`: output shows workflow usage and before/after clicked states.
- `workflow_not_used`: output appears to solve the task without `wf`,
`wf-rpc-server`, deployment, or run evidence.
- `run_failed`: output includes workflow usage but reports a failure.
- `timeout`: the opencode process exceeded the configured timeout.
- `parse_error`: the harness could not read opencode JSON/JSONL output.
- `unknown`: no clear success or failure signal was found.
Committed tests cover harness logic only. They do not invoke opencode.
````
- [ ] **Step 4: Configure result ignores**
Write this content to
`examples/agent_challenges/browser_click_challenge/results/.gitignore`:
```gitignore
*
!.gitignore
```
- [ ] **Step 5: Commit**
Run:
```powershell
git add examples\agent_challenges\browser_click_challenge
git commit -m "docs: add browser click agent challenge prompt"
```
Expected: commit succeeds.
---
## Task 2: Add Harness Core With Tests
**Files:**
- Create: `examples/agent_challenges/browser_click_challenge/run_opencode_trials.py`
- Create: `tests/examples/test_opencode_browser_click_challenge.py`
- [ ] **Step 1: Write failing tests**
Create `tests/examples/test_opencode_browser_click_challenge.py`:
```python
from __future__ import annotations
import json
from pathlib import Path
from examples.agent_challenges.browser_click_challenge.run_opencode_trials import (
TrialConfig,
build_opencode_command,
classify_output,
parse_opencode_output,
trial_output_path,
)
def test_build_opencode_command_without_attach(tmp_path: Path) -> None:
prompt = tmp_path / "prompt.md"
prompt.write_text("hello", encoding="utf-8")
config = TrialConfig(
model="opencode/mimo-v2.5-free",
variant="high",
prompt_path=prompt,
attach_url=None,
timeout_seconds=120,
)
command = build_opencode_command(config)
assert command[:2] == ["opencode", "run"]
assert "--attach" not in command
assert "--format" in command
assert "json" in command
assert "--model" in command
assert "opencode/mimo-v2.5-free" in command
assert "hello" in command
def test_build_opencode_command_with_attach(tmp_path: Path) -> None:
prompt = tmp_path / "prompt.md"
prompt.write_text("hello", encoding="utf-8")
config = TrialConfig(
model="opencode/deepseek-v3.1-free",
variant="high",
prompt_path=prompt,
attach_url="http://127.0.0.1:4096",
timeout_seconds=120,
)
command = build_opencode_command(config)
assert "--attach" in command
assert "http://127.0.0.1:4096" in command
def test_parse_opencode_output_reads_json_object() -> None:
payload = {"text": "wf run start demo.default\nbefore.clicked false\nafter.clicked true"}
parsed = parse_opencode_output(json.dumps(payload))
assert parsed["text"] == payload["text"]
def test_parse_opencode_output_reads_last_jsonl_object() -> None:
payload = "\n".join(
[
json.dumps({"type": "log", "text": "starting"}),
json.dumps({"type": "message", "text": "final"}),
]
)
parsed = parse_opencode_output(payload)
assert parsed["text"] == "final"
def test_classify_output_success() -> None:
result = classify_output(
"""
uv run wf-rpc-server --config examples/browser_click_workflow/wf.config.json
uv run wf run start browser_click_case_study.default
deployment id: browser_click_case_study.default
run id: run_123
before.clicked is false
after.clicked is true
"""
)
assert result == "success"
def test_classify_output_workflow_not_used() -> None:
result = classify_output(
"""
I wrote a Playwright script.
before clicked false
after clicked true
"""
)
assert result == "workflow_not_used"
def test_classify_output_run_failed() -> None:
result = classify_output(
"""
wf run start browser_click_case_study.default
error: deployment validation failed
"""
)
assert result == "run_failed"
def test_trial_output_path_is_zero_padded(tmp_path: Path) -> None:
path = trial_output_path(tmp_path, model="opencode/mimo-v2.5-free", index=3)
assert path.name == "opencode_mimo-v2.5-free-trial-003.json"
```
- [ ] **Step 2: Run tests to verify failure**
Run:
```powershell
uv run pytest tests/examples/test_opencode_browser_click_challenge.py -q
```
Expected: FAIL because module/functions do not exist.
- [ ] **Step 3: Implement harness module**
Create `examples/agent_challenges/browser_click_challenge/run_opencode_trials.py`:
```python
from __future__ import annotations
import argparse
import json
import subprocess
import time
from dataclasses import asdict, dataclass
from pathlib import Path
from typing import Any, Literal
Classification = Literal[
"success",
"workflow_not_used",
"run_failed",
"timeout",
"parse_error",
"unknown",
]
ROOT = Path(__file__).resolve().parents[3]
CHALLENGE_DIR = Path(__file__).resolve().parent
DEFAULT_PROMPT = CHALLENGE_DIR / "prompt.md"
DEFAULT_RESULTS_DIR = CHALLENGE_DIR / "results"
@dataclass(frozen=True, slots=True)
class TrialConfig:
model: str
variant: str
prompt_path: Path
attach_url: str | None
timeout_seconds: int
def build_opencode_command(config: TrialConfig) -> list[str]:
prompt_text = config.prompt_path.read_text(encoding="utf-8")
command = [
"opencode",
"run",
]
if config.attach_url is not None:
command.extend(["--attach", config.attach_url])
command.extend(
[
prompt_text,
"--format",
"json",
"--model",
config.model,
"--variant",
config.variant,
]
)
return command
def parse_opencode_output(stdout: str) -> dict[str, Any]:
text = stdout.strip()
if not text:
raise ValueError("opencode produced no JSON output")
try:
parsed = json.loads(text)
except json.JSONDecodeError:
parsed = _parse_jsonl_tail(text)
if not isinstance(parsed, dict):
raise ValueError("opencode output was not a JSON object")
return parsed
def _parse_jsonl_tail(text: str) -> dict[str, Any]:
last_error: json.JSONDecodeError | None = None
for line in reversed(text.splitlines()):
stripped = line.strip()
if not stripped:
continue
try:
parsed = json.loads(stripped)
except json.JSONDecodeError as exc:
last_error = exc
continue
if isinstance(parsed, dict):
return parsed
if last_error is not None:
raise last_error
raise ValueError("opencode output did not contain JSON lines")
def classify_output(text: str) -> Classification:
lowered = text.lower()
workflow_markers = [
"wf ",
"wf-rpc-server",
"deployment",
"run id",
"run_",
]
used_workflow = any(marker in lowered for marker in workflow_markers)
failed = any(
marker in lowered
for marker in [
"error:",
"failed",
"traceback",
"exception",
"validation failed",
]
)
before_false = (
"before.clicked is false" in lowered
or '"before"' in lowered
and '"clicked": false' in lowered
)
after_true = (
"after.clicked is true" in lowered
or '"after"' in lowered
and '"clicked": true' in lowered
)
if used_workflow and before_false and after_true and not failed:
return "success"
if used_workflow and failed:
return "run_failed"
if not used_workflow and (before_false or after_true or "playwright" in lowered):
return "workflow_not_used"
return "unknown"
def trial_output_path(results_dir: Path, *, model: str, index: int) -> Path:
safe_model = model.replace("/", "_").replace(":", "_")
return results_dir / f"{safe_model}-trial-{index:03d}.json"
def run_trial(config: TrialConfig, *, index: int, results_dir: Path) -> dict[str, Any]:
command = build_opencode_command(config)
started = time.monotonic()
try:
completed = subprocess.run(
command,
cwd=ROOT,
text=True,
capture_output=True,
timeout=config.timeout_seconds,
check=False,
)
duration_seconds = time.monotonic() - started
except subprocess.TimeoutExpired as exc:
payload = {
"index": index,
"config": _jsonable_config(config),
"command": command,
"classification": "timeout",
"duration_seconds": config.timeout_seconds,
"returncode": None,
"stdout": exc.stdout or "",
"stderr": exc.stderr or "",
"parsed": None,
}
_write_trial_result(results_dir, config=config, index=index, payload=payload)
return payload
parsed: dict[str, Any] | None
try:
parsed = parse_opencode_output(completed.stdout)
text = _result_text(parsed)
classification = classify_output(text)
except Exception:
parsed = None
classification = "parse_error"
payload = {
"index": index,
"config": _jsonable_config(config),
"command": command,
"classification": classification,
"duration_seconds": duration_seconds,
"returncode": completed.returncode,
"stdout": completed.stdout,
"stderr": completed.stderr,
"parsed": parsed,
}
_write_trial_result(results_dir, config=config, index=index, payload=payload)
return payload
def _result_text(parsed: dict[str, Any]) -> str:
for key in ("text", "message", "content", "output"):
value = parsed.get(key)
if isinstance(value, str):
return value
return json.dumps(parsed, sort_keys=True)
def _jsonable_config(config: TrialConfig) -> dict[str, Any]:
payload = asdict(config)
payload["prompt_path"] = str(config.prompt_path)
return payload
def _write_trial_result(
results_dir: Path,
*,
config: TrialConfig,
index: int,
payload: dict[str, Any],
) -> None:
results_dir.mkdir(parents=True, exist_ok=True)
path = trial_output_path(results_dir, model=config.model, index=index)
path.write_text(json.dumps(payload, indent=2, sort_keys=True), encoding="utf-8")
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--model", default="opencode/mimo-v2.5-free")
parser.add_argument("--variant", default="high")
parser.add_argument("--trials", type=int, default=1)
parser.add_argument("--timeout-seconds", type=int, default=600)
parser.add_argument("--attach", dest="attach_url", default=None)
parser.add_argument("--prompt", type=Path, default=DEFAULT_PROMPT)
parser.add_argument("--results-dir", type=Path, default=DEFAULT_RESULTS_DIR)
args = parser.parse_args(argv)
if args.trials < 1:
parser.error("--trials must be >= 1")
config = TrialConfig(
model=args.model,
variant=args.variant,
prompt_path=args.prompt,
attach_url=args.attach_url,
timeout_seconds=args.timeout_seconds,
)
summaries: list[dict[str, Any]] = []
for index in range(1, args.trials + 1):
result = run_trial(config, index=index, results_dir=args.results_dir)
summaries.append(
{
"index": index,
"classification": result["classification"],
"returncode": result["returncode"],
"duration_seconds": round(float(result["duration_seconds"]), 3),
}
)
print(json.dumps(summaries[-1], sort_keys=True))
success_count = sum(1 for item in summaries if item["classification"] == "success")
print(json.dumps({"success_count": success_count, "trial_count": len(summaries)}))
return 0
if __name__ == "__main__":
raise SystemExit(main())
```
- [ ] **Step 4: Run tests**
Run:
```powershell
uv run pytest tests/examples/test_opencode_browser_click_challenge.py -q
```
Expected: PASS.
- [ ] **Step 5: Run lint/typecheck**
Run:
```powershell
uv run ruff check examples/agent_challenges/browser_click_challenge/run_opencode_trials.py tests/examples/test_opencode_browser_click_challenge.py
uv run basedpyright --level error examples/agent_challenges/browser_click_challenge/run_opencode_trials.py tests/examples/test_opencode_browser_click_challenge.py
```
Expected: both pass.
- [ ] **Step 6: Commit**
Run:
```powershell
git add examples\agent_challenges\browser_click_challenge\run_opencode_trials.py tests\examples\test_opencode_browser_click_challenge.py
git commit -m "test: add opencode browser click challenge harness"
```
Expected: commit succeeds.
---
## Task 3: Add Evidence Docs And Final Verification
**Files:**
- Modify: `docs/add/evidence-index.md`
- Modify: `docs/current_roadmap.md`
- Move: this plan to `docs/historical/superpowers/plans/2026-06-15-opencode-browser-click-challenge-harness.md`
- [ ] **Step 1: Update evidence index**
Add this bullet to `docs/add/evidence-index.md` near the browser-click workflow evidence:
```markdown
- `examples/agent_challenges/browser_click_challenge/`
- `tests/examples/test_opencode_browser_click_challenge.py`
```
- [ ] **Step 2: Update roadmap**
Add this bullet under the product smoke/evidence area in `docs/current_roadmap.md`:
```markdown
- Completed: an opencode browser-click challenge harness captures external
agent trials against the deterministic browser-click workflow example without
changing product runtime code.
```
- [ ] **Step 3: Archive plan**
Run:
```powershell
Move-Item -LiteralPath docs\superpowers\plans\2026-06-15-opencode-browser-click-challenge-harness.md -Destination docs\historical\superpowers\plans\2026-06-15-opencode-browser-click-challenge-harness.md
```
Expected: the plan file moves to `docs/historical/superpowers/plans/`.
- [ ] **Step 4: Run final verification**
Run:
```powershell
uv run pytest tests/examples/test_opencode_browser_click_challenge.py tests/docs -q
uv run ruff check examples/agent_challenges/browser_click_challenge tests/examples/test_opencode_browser_click_challenge.py tests/docs
uv run basedpyright --level error examples/agent_challenges/browser_click_challenge tests/examples/test_opencode_browser_click_challenge.py tests/docs
git status --short
```
Expected:
- pytest passes.
- ruff passes.
- basedpyright reports 0 errors.
- `git status --short` shows only intended docs/harness changes before commit.
- [ ] **Step 5: Commit**
Run:
```powershell
git add docs\add\evidence-index.md docs\current_roadmap.md docs\historical\superpowers\plans\2026-06-15-opencode-browser-click-challenge-harness.md
git commit -m "docs: record opencode browser click challenge harness"
```
Expected: commit succeeds.
---
## Manual Trial Command
After implementation, a human can run:
```powershell
uv run python examples/agent_challenges/browser_click_challenge/run_opencode_trials.py `
--model opencode/mimo-v2.5-free `
--variant high `
--trials 1
```
With a Playwright MCP/tool endpoint attached:
```powershell
uv run python examples/agent_challenges/browser_click_challenge/run_opencode_trials.py `
--model opencode/mimo-v2.5-free `
--variant high `
--trials 1 `
--attach http://127.0.0.1:4096
```
The harness writes detailed JSON files under:
```text
examples/agent_challenges/browser_click_challenge/results/
```
---
## Self-Review Checklist
- The harness does not call opencode during pytest.
- Product runtime code is untouched.
- The prompt requires workflow/deployment/run evidence.
- Classification is deterministic and string-based.
- Trial outputs are gitignored by the local `results/.gitignore`.
- Playwright MCP attachment is optional.
- Docs make clear this is evidence tooling, not product runtime.