Files
lda-wf/docs/historical/superpowers/plans/2026-06-15-opencode-browser-click-challenge-harness.md
T
lda Verified 10dabca241 feat: add opencode browser click challenge harness
- Challenge prompt requiring workflow/deployment/run evidence
- CLI harness (run_opencode_trials.py) for running N agent trials
- Classification: success, workflow_not_used, run_failed, timeout, parse_error, unknown
- 8 unit tests covering build/parse/classify/path logic
- README with usage docs and optional Playwright MCP attachment
- Evidence index and roadmap updated
- Plan archived to historical/
2026-06-15 02:32:54 +07:00

21 KiB

Opencode Browser Click Challenge Harness Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Build a local evidence harness that runs the browser-click workflow challenge through opencode run, captures trial outputs, and classifies failures without changing product runtime code.

Architecture: The harness lives under examples/agent_challenges/browser_click_challenge/ and treats opencode as an external executable. Tests do not call opencode; they cover prompt loading, command construction, JSONL/result parsing, and deterministic classification from captured text.

Tech Stack: Python stdlib (argparse, json, subprocess, dataclasses, pathlib, time), pytest, existing browser-click workflow example, optional Playwright MCP attachment configured by command-line flags.


File Structure

  • Create examples/agent_challenges/browser_click_challenge/prompt.md
    • The exact prompt sent to agents.
    • It must require workflow usage, deployment run, before/after snapshots, cleanup, and final summary.
  • Create examples/agent_challenges/browser_click_challenge/README.md
    • How to prepare the local server/config.
    • How to run one trial and multiple trials.
    • How to attach Playwright MCP when desired.
    • How success/failure is classified.
  • Create examples/agent_challenges/browser_click_challenge/results/.gitignore
    • Keeps the output folder in git.
    • Ignores generated trial outputs in the folder itself.
  • Create examples/agent_challenges/browser_click_challenge/run_opencode_trials.py
    • CLI harness.
    • Builds opencode run command.
    • Runs N trials.
    • Writes one JSON artifact per trial.
    • Classifies result text.
  • Create tests/examples/test_opencode_browser_click_challenge.py
    • Unit tests only.
    • No live opencode calls.
  • Modify docs/add/evidence-index.md
    • Add the challenge harness as planned/evaluation evidence.
  • Modify docs/current_roadmap.md
    • Add completed evidence-harness bullet after implementation.

Task 1: Add Prompt And README

Files:

  • Create: examples/agent_challenges/browser_click_challenge/prompt.md

  • Create: examples/agent_challenges/browser_click_challenge/README.md

  • Create: examples/agent_challenges/browser_click_challenge/results/.gitignore

  • Test: no pytest yet; this task is docs/assets only

  • Step 1: Create challenge folder

Run:

New-Item -ItemType Directory -Force examples\agent_challenges\browser_click_challenge\results
New-Item -ItemType File -Force examples\agent_challenges\browser_click_challenge\results\.gitignore

Expected: the folder exists.

  • Step 2: Create prompt

Create examples/agent_challenges/browser_click_challenge/prompt.md:

# Browser Click Workflow Challenge

Build and successfully run a workflow that:

1. Opens a browser page or local web page with a visible button.
2. Waits for a human click or performs a clearly simulated click.
3. Captures a before snapshot and an after snapshot.
4. Returns both snapshots as workflow output.

Use this repository's workflow product path. That means you should use the
`wf` CLI and/or `wf-rpc-server`, create or reuse a workflow deployment, and run
the deployment through the workflow API. Do not solve the challenge with only a
standalone Playwright/Python script.

The repository already includes a deterministic source example at:

```text
examples/browser_click_workflow/

You may inspect and use it. A successful final answer must include:

  • the commands you ran,
  • the deployment id,
  • the run id if one was produced,
  • evidence that before.clicked is false,
  • evidence that after.clicked is true,
  • whether any server/browser process remains running.

If something fails, report the exact command and error instead of hiding it.


- [ ] **Step 3: Create README**

Create `examples/agent_challenges/browser_click_challenge/README.md`:

````markdown
# Opencode Browser Click Challenge Harness

This harness runs agent trials against the browser-click workflow challenge.
It is evidence tooling, not product runtime code.

The deterministic workflow example is:

```text
examples/browser_click_workflow/

One Trial

From the repository root:

uv run python examples/agent_challenges/browser_click_challenge/run_opencode_trials.py `
  --model opencode/mimo-v2.5-free `
  --variant high `
  --trials 1

Results are written to:

examples/agent_challenges/browser_click_challenge/results/

Optional Playwright MCP Attachment

If you want the agent to have browser-control tools, pass:

--attach http://127.0.0.1:4096

Start that MCP/tool endpoint separately. For example, one possible MCP server command is:

{
  "command": "npx",
  "args": ["-y", "@playwright/mcp@latest"]
}

The baseline challenge does not require Playwright MCP. The score is based on whether the agent used the workflow product path and produced the expected workflow output.

Classification

Each trial is classified as one of:

  • success: output shows workflow usage and before/after clicked states.
  • workflow_not_used: output appears to solve the task without wf, wf-rpc-server, deployment, or run evidence.
  • run_failed: output includes workflow usage but reports a failure.
  • timeout: the opencode process exceeded the configured timeout.
  • parse_error: the harness could not read opencode JSON/JSONL output.
  • unknown: no clear success or failure signal was found.

Committed tests cover harness logic only. They do not invoke opencode.


- [ ] **Step 4: Configure result ignores**

Write this content to
`examples/agent_challenges/browser_click_challenge/results/.gitignore`:

```gitignore
*
!.gitignore
```

- [ ] **Step 5: Commit**

Run:

```powershell
git add examples\agent_challenges\browser_click_challenge
git commit -m "docs: add browser click agent challenge prompt"
```

Expected: commit succeeds.

---

## Task 2: Add Harness Core With Tests

**Files:**
- Create: `examples/agent_challenges/browser_click_challenge/run_opencode_trials.py`
- Create: `tests/examples/test_opencode_browser_click_challenge.py`

- [ ] **Step 1: Write failing tests**

Create `tests/examples/test_opencode_browser_click_challenge.py`:

```python
from __future__ import annotations

import json
from pathlib import Path

from examples.agent_challenges.browser_click_challenge.run_opencode_trials import (
    TrialConfig,
    build_opencode_command,
    classify_output,
    parse_opencode_output,
    trial_output_path,
)


def test_build_opencode_command_without_attach(tmp_path: Path) -> None:
    prompt = tmp_path / "prompt.md"
    prompt.write_text("hello", encoding="utf-8")
    config = TrialConfig(
        model="opencode/mimo-v2.5-free",
        variant="high",
        prompt_path=prompt,
        attach_url=None,
        timeout_seconds=120,
    )

    command = build_opencode_command(config)

    assert command[:2] == ["opencode", "run"]
    assert "--attach" not in command
    assert "--format" in command
    assert "json" in command
    assert "--model" in command
    assert "opencode/mimo-v2.5-free" in command
    assert "hello" in command


def test_build_opencode_command_with_attach(tmp_path: Path) -> None:
    prompt = tmp_path / "prompt.md"
    prompt.write_text("hello", encoding="utf-8")
    config = TrialConfig(
        model="opencode/deepseek-v3.1-free",
        variant="high",
        prompt_path=prompt,
        attach_url="http://127.0.0.1:4096",
        timeout_seconds=120,
    )

    command = build_opencode_command(config)

    assert "--attach" in command
    assert "http://127.0.0.1:4096" in command


def test_parse_opencode_output_reads_json_object() -> None:
    payload = {"text": "wf run start demo.default\nbefore.clicked false\nafter.clicked true"}

    parsed = parse_opencode_output(json.dumps(payload))

    assert parsed["text"] == payload["text"]


def test_parse_opencode_output_reads_last_jsonl_object() -> None:
    payload = "\n".join(
        [
            json.dumps({"type": "log", "text": "starting"}),
            json.dumps({"type": "message", "text": "final"}),
        ]
    )

    parsed = parse_opencode_output(payload)

    assert parsed["text"] == "final"


def test_classify_output_success() -> None:
    result = classify_output(
        """
        uv run wf-rpc-server --config examples/browser_click_workflow/wf.config.json
        uv run wf run start browser_click_case_study.default
        deployment id: browser_click_case_study.default
        run id: run_123
        before.clicked is false
        after.clicked is true
        """
    )

    assert result == "success"


def test_classify_output_workflow_not_used() -> None:
    result = classify_output(
        """
        I wrote a Playwright script.
        before clicked false
        after clicked true
        """
    )

    assert result == "workflow_not_used"


def test_classify_output_run_failed() -> None:
    result = classify_output(
        """
        wf run start browser_click_case_study.default
        error: deployment validation failed
        """
    )

    assert result == "run_failed"


def test_trial_output_path_is_zero_padded(tmp_path: Path) -> None:
    path = trial_output_path(tmp_path, model="opencode/mimo-v2.5-free", index=3)

    assert path.name == "opencode_mimo-v2.5-free-trial-003.json"
```

- [ ] **Step 2: Run tests to verify failure**

Run:

```powershell
uv run pytest tests/examples/test_opencode_browser_click_challenge.py -q
```

Expected: FAIL because module/functions do not exist.

- [ ] **Step 3: Implement harness module**

Create `examples/agent_challenges/browser_click_challenge/run_opencode_trials.py`:

```python
from __future__ import annotations

import argparse
import json
import subprocess
import time
from dataclasses import asdict, dataclass
from pathlib import Path
from typing import Any, Literal


Classification = Literal[
    "success",
    "workflow_not_used",
    "run_failed",
    "timeout",
    "parse_error",
    "unknown",
]

ROOT = Path(__file__).resolve().parents[3]
CHALLENGE_DIR = Path(__file__).resolve().parent
DEFAULT_PROMPT = CHALLENGE_DIR / "prompt.md"
DEFAULT_RESULTS_DIR = CHALLENGE_DIR / "results"


@dataclass(frozen=True, slots=True)
class TrialConfig:
    model: str
    variant: str
    prompt_path: Path
    attach_url: str | None
    timeout_seconds: int


def build_opencode_command(config: TrialConfig) -> list[str]:
    prompt_text = config.prompt_path.read_text(encoding="utf-8")
    command = [
        "opencode",
        "run",
    ]
    if config.attach_url is not None:
        command.extend(["--attach", config.attach_url])
    command.extend(
        [
            prompt_text,
            "--format",
            "json",
            "--model",
            config.model,
            "--variant",
            config.variant,
        ]
    )
    return command


def parse_opencode_output(stdout: str) -> dict[str, Any]:
    text = stdout.strip()
    if not text:
        raise ValueError("opencode produced no JSON output")

    try:
        parsed = json.loads(text)
    except json.JSONDecodeError:
        parsed = _parse_jsonl_tail(text)

    if not isinstance(parsed, dict):
        raise ValueError("opencode output was not a JSON object")
    return parsed


def _parse_jsonl_tail(text: str) -> dict[str, Any]:
    last_error: json.JSONDecodeError | None = None
    for line in reversed(text.splitlines()):
        stripped = line.strip()
        if not stripped:
            continue
        try:
            parsed = json.loads(stripped)
        except json.JSONDecodeError as exc:
            last_error = exc
            continue
        if isinstance(parsed, dict):
            return parsed
    if last_error is not None:
        raise last_error
    raise ValueError("opencode output did not contain JSON lines")


def classify_output(text: str) -> Classification:
    lowered = text.lower()
    workflow_markers = [
        "wf ",
        "wf-rpc-server",
        "deployment",
        "run id",
        "run_",
    ]
    used_workflow = any(marker in lowered for marker in workflow_markers)
    failed = any(
        marker in lowered
        for marker in [
            "error:",
            "failed",
            "traceback",
            "exception",
            "validation failed",
        ]
    )
    before_false = (
        "before.clicked is false" in lowered
        or '"before"' in lowered
        and '"clicked": false' in lowered
    )
    after_true = (
        "after.clicked is true" in lowered
        or '"after"' in lowered
        and '"clicked": true' in lowered
    )

    if used_workflow and before_false and after_true and not failed:
        return "success"
    if used_workflow and failed:
        return "run_failed"
    if not used_workflow and (before_false or after_true or "playwright" in lowered):
        return "workflow_not_used"
    return "unknown"


def trial_output_path(results_dir: Path, *, model: str, index: int) -> Path:
    safe_model = model.replace("/", "_").replace(":", "_")
    return results_dir / f"{safe_model}-trial-{index:03d}.json"


def run_trial(config: TrialConfig, *, index: int, results_dir: Path) -> dict[str, Any]:
    command = build_opencode_command(config)
    started = time.monotonic()
    try:
        completed = subprocess.run(
            command,
            cwd=ROOT,
            text=True,
            capture_output=True,
            timeout=config.timeout_seconds,
            check=False,
        )
        duration_seconds = time.monotonic() - started
    except subprocess.TimeoutExpired as exc:
        payload = {
            "index": index,
            "config": _jsonable_config(config),
            "command": command,
            "classification": "timeout",
            "duration_seconds": config.timeout_seconds,
            "returncode": None,
            "stdout": exc.stdout or "",
            "stderr": exc.stderr or "",
            "parsed": None,
        }
        _write_trial_result(results_dir, config=config, index=index, payload=payload)
        return payload

    parsed: dict[str, Any] | None
    try:
        parsed = parse_opencode_output(completed.stdout)
        text = _result_text(parsed)
        classification = classify_output(text)
    except Exception:
        parsed = None
        classification = "parse_error"

    payload = {
        "index": index,
        "config": _jsonable_config(config),
        "command": command,
        "classification": classification,
        "duration_seconds": duration_seconds,
        "returncode": completed.returncode,
        "stdout": completed.stdout,
        "stderr": completed.stderr,
        "parsed": parsed,
    }
    _write_trial_result(results_dir, config=config, index=index, payload=payload)
    return payload


def _result_text(parsed: dict[str, Any]) -> str:
    for key in ("text", "message", "content", "output"):
        value = parsed.get(key)
        if isinstance(value, str):
            return value
    return json.dumps(parsed, sort_keys=True)


def _jsonable_config(config: TrialConfig) -> dict[str, Any]:
    payload = asdict(config)
    payload["prompt_path"] = str(config.prompt_path)
    return payload


def _write_trial_result(
    results_dir: Path,
    *,
    config: TrialConfig,
    index: int,
    payload: dict[str, Any],
) -> None:
    results_dir.mkdir(parents=True, exist_ok=True)
    path = trial_output_path(results_dir, model=config.model, index=index)
    path.write_text(json.dumps(payload, indent=2, sort_keys=True), encoding="utf-8")


def main(argv: list[str] | None = None) -> int:
    parser = argparse.ArgumentParser(description=__doc__)
    parser.add_argument("--model", default="opencode/mimo-v2.5-free")
    parser.add_argument("--variant", default="high")
    parser.add_argument("--trials", type=int, default=1)
    parser.add_argument("--timeout-seconds", type=int, default=600)
    parser.add_argument("--attach", dest="attach_url", default=None)
    parser.add_argument("--prompt", type=Path, default=DEFAULT_PROMPT)
    parser.add_argument("--results-dir", type=Path, default=DEFAULT_RESULTS_DIR)
    args = parser.parse_args(argv)

    if args.trials < 1:
        parser.error("--trials must be >= 1")

    config = TrialConfig(
        model=args.model,
        variant=args.variant,
        prompt_path=args.prompt,
        attach_url=args.attach_url,
        timeout_seconds=args.timeout_seconds,
    )

    summaries: list[dict[str, Any]] = []
    for index in range(1, args.trials + 1):
        result = run_trial(config, index=index, results_dir=args.results_dir)
        summaries.append(
            {
                "index": index,
                "classification": result["classification"],
                "returncode": result["returncode"],
                "duration_seconds": round(float(result["duration_seconds"]), 3),
            }
        )
        print(json.dumps(summaries[-1], sort_keys=True))

    success_count = sum(1 for item in summaries if item["classification"] == "success")
    print(json.dumps({"success_count": success_count, "trial_count": len(summaries)}))
    return 0


if __name__ == "__main__":
    raise SystemExit(main())
```

- [ ] **Step 4: Run tests**

Run:

```powershell
uv run pytest tests/examples/test_opencode_browser_click_challenge.py -q
```

Expected: PASS.

- [ ] **Step 5: Run lint/typecheck**

Run:

```powershell
uv run ruff check examples/agent_challenges/browser_click_challenge/run_opencode_trials.py tests/examples/test_opencode_browser_click_challenge.py
uv run basedpyright --level error examples/agent_challenges/browser_click_challenge/run_opencode_trials.py tests/examples/test_opencode_browser_click_challenge.py
```

Expected: both pass.

- [ ] **Step 6: Commit**

Run:

```powershell
git add examples\agent_challenges\browser_click_challenge\run_opencode_trials.py tests\examples\test_opencode_browser_click_challenge.py
git commit -m "test: add opencode browser click challenge harness"
```

Expected: commit succeeds.

---

## Task 3: Add Evidence Docs And Final Verification

**Files:**
- Modify: `docs/add/evidence-index.md`
- Modify: `docs/current_roadmap.md`
- Move: this plan to `docs/historical/superpowers/plans/2026-06-15-opencode-browser-click-challenge-harness.md`

- [ ] **Step 1: Update evidence index**

Add this bullet to `docs/add/evidence-index.md` near the browser-click workflow evidence:

```markdown
- `examples/agent_challenges/browser_click_challenge/`
- `tests/examples/test_opencode_browser_click_challenge.py`
```

- [ ] **Step 2: Update roadmap**

Add this bullet under the product smoke/evidence area in `docs/current_roadmap.md`:

```markdown
- Completed: an opencode browser-click challenge harness captures external
  agent trials against the deterministic browser-click workflow example without
  changing product runtime code.
```

- [ ] **Step 3: Archive plan**

Run:

```powershell
Move-Item -LiteralPath docs\superpowers\plans\2026-06-15-opencode-browser-click-challenge-harness.md -Destination docs\historical\superpowers\plans\2026-06-15-opencode-browser-click-challenge-harness.md
```

Expected: the plan file moves to `docs/historical/superpowers/plans/`.

- [ ] **Step 4: Run final verification**

Run:

```powershell
uv run pytest tests/examples/test_opencode_browser_click_challenge.py tests/docs -q
uv run ruff check examples/agent_challenges/browser_click_challenge tests/examples/test_opencode_browser_click_challenge.py tests/docs
uv run basedpyright --level error examples/agent_challenges/browser_click_challenge tests/examples/test_opencode_browser_click_challenge.py tests/docs
git status --short
```

Expected:
- pytest passes.
- ruff passes.
- basedpyright reports 0 errors.
- `git status --short` shows only intended docs/harness changes before commit.

- [ ] **Step 5: Commit**

Run:

```powershell
git add docs\add\evidence-index.md docs\current_roadmap.md docs\historical\superpowers\plans\2026-06-15-opencode-browser-click-challenge-harness.md
git commit -m "docs: record opencode browser click challenge harness"
```

Expected: commit succeeds.

---

## Manual Trial Command

After implementation, a human can run:

```powershell
uv run python examples/agent_challenges/browser_click_challenge/run_opencode_trials.py `
  --model opencode/mimo-v2.5-free `
  --variant high `
  --trials 1
```

With a Playwright MCP/tool endpoint attached:

```powershell
uv run python examples/agent_challenges/browser_click_challenge/run_opencode_trials.py `
  --model opencode/mimo-v2.5-free `
  --variant high `
  --trials 1 `
  --attach http://127.0.0.1:4096
```

The harness writes detailed JSON files under:

```text
examples/agent_challenges/browser_click_challenge/results/
```

---

## Self-Review Checklist

- The harness does not call opencode during pytest.
- Product runtime code is untouched.
- The prompt requires workflow/deployment/run evidence.
- Classification is deterministic and string-based.
- Trial outputs are gitignored by the local `results/.gitignore`.
- Playwright MCP attachment is optional.
- Docs make clear this is evidence tooling, not product runtime.