test: add concurrent agent challenge matrix
This commit is contained in:
@@ -1142,7 +1142,7 @@ The evidence supporting the thesis claims includes:
|
|||||||
| Interrupted runs resume at explicit boundaries | `tests/wf_api/test_run_api.py` and resume-concurrency tests | Stopped run state is persisted and resumed through the run API | Pass in focused test suite |
|
| Interrupted runs resume at explicit boundaries | `tests/wf_api/test_run_api.py` and resume-concurrency tests | Stopped run state is persisted and resumed through the run API | Pass in focused test suite |
|
||||||
| Python source lifecycle works | `tests/examples/test_report_workflow_example.py` | Python capability -> artifact -> deployment -> run completes | Pass in focused test suite |
|
| Python source lifecycle works | `tests/examples/test_report_workflow_example.py` | Python capability -> artifact -> deployment -> run completes | Pass in focused test suite |
|
||||||
| Serial multi-node workflow works | `tests/examples/test_browser_click_workflow_example.py` | `open_click_page` -> `wait_for_click` -> `collect_snapshots` completes with before/after evidence | Pass in focused test suite |
|
| Serial multi-node workflow works | `tests/examples/test_browser_click_workflow_example.py` | `open_click_page` -> `wait_for_click` -> `collect_snapshots` completes with before/after evidence | Pass in focused test suite |
|
||||||
| Agent challenge harness implementation exists; no aggregate agent-performance claim | `examples/agent_challenges/browser_click_challenge/` and Appendix C | The harness can prompt, classify, and support manual audit of browser-click workflow trials | Harness tests pass; aggregate model results pending |
|
| Agent challenge harness implementation exists; no aggregate agent-performance claim | `examples/agent_challenges/browser_click_challenge/`, `examples/agent_challenges/report_workflow_challenge/`, and Appendix C | The shared harness can prompt, classify, collect tool/token evidence, and support manual audit across browser-click and report-workflow trials | Harness tests pass; aggregate model results pending |
|
||||||
| CLI and JSON-RPC share the API surface | `tests/wf_transport_rpc_http/` and `tests/wf_cli/` | Transport and CLI delegate to the same workflow operations | Pass in focused test suite |
|
| CLI and JSON-RPC share the API surface | `tests/wf_transport_rpc_http/` and `tests/wf_cli/` | Transport and CLI delegate to the same workflow operations | Pass in focused test suite |
|
||||||
|
|
||||||
The table summarizes repository evidence; it is not a substitute for rerunning
|
The table summarizes repository evidence; it is not a substitute for rerunning
|
||||||
|
|||||||
@@ -207,10 +207,10 @@ stable.
|
|||||||
- Completed supplemental browser-click workflow example with serial multi-node lifecycle evidence.
|
- Completed supplemental browser-click workflow example with serial multi-node lifecycle evidence.
|
||||||
- Completed: an opencode browser-click challenge harness captures external
|
- Completed: an opencode browser-click challenge harness captures external
|
||||||
agent trials against the deterministic browser-click workflow example without
|
agent trials against the deterministic browser-click workflow example without
|
||||||
changing product runtime code.
|
changing product runtime code. The old staged-server modes were replaced by
|
||||||
- Next thesis/evaluation polish: treat `skills/`, runbooks, and challenge
|
per-trial local configs in the generic V2 harness.
|
||||||
prompt templates as the agent instruction layer. Strengthen them before
|
- Completed: `skills/`, runbooks, and challenge prompt templates are treated as
|
||||||
claiming aggregate agent-operability results, and track when trials rely on
|
the agent instruction layer. Challenge reports now track when trials rely on
|
||||||
product code, prior stores, adjacent attempts, or existing example solutions.
|
product code, prior stores, adjacent attempts, or existing example solutions.
|
||||||
- Completed: workflow/CLI agent instructions now form an explicit copyable
|
- Completed: workflow/CLI agent instructions now form an explicit copyable
|
||||||
bundle for controlled challenge profiles, use `wf schema` for public shape
|
bundle for controlled challenge profiles, use `wf schema` for public shape
|
||||||
@@ -225,6 +225,10 @@ stable.
|
|||||||
every V2 trial and regenerate both after audit without mutating raw evidence.
|
every V2 trial and regenerate both after audit without mutating raw evidence.
|
||||||
Implementation:
|
Implementation:
|
||||||
[`report projections`](historical/superpowers/plans/2026-06-23-agent-challenge-report-projections.md).
|
[`report projections`](historical/superpowers/plans/2026-06-23-agent-challenge-report-projections.md).
|
||||||
|
- Completed: challenge trial collection now supports bounded concurrency through
|
||||||
|
`run_trials.py --concurrency` and a Python matrix runner,
|
||||||
|
`examples/agent_challenges/run_matrix.py`. The PowerShell matrix helper now
|
||||||
|
delegates to the Python runner.
|
||||||
- Completed: shared agent challenge evaluation runbook documents trial
|
- Completed: shared agent challenge evaluation runbook documents trial
|
||||||
execution, instruction profiles, manual audit, and the distinction between
|
execution, instruction profiles, manual audit, and the distinction between
|
||||||
evaluation validity and policy coverage:
|
evaluation validity and policy coverage:
|
||||||
|
|||||||
@@ -11,20 +11,22 @@ commands, or product gaps.
|
|||||||
|
|
||||||
## What The Harness Produces
|
## What The Harness Produces
|
||||||
|
|
||||||
Each trial writes three kinds of evidence:
|
Each trial writes four kinds of evidence:
|
||||||
|
|
||||||
- Raw result JSON in the challenge `results/` directory. This is immutable raw
|
- Raw result JSON in the challenge `results/` directory. This is immutable raw
|
||||||
evidence from the runner.
|
evidence from the runner.
|
||||||
- Machine report JSON beside the raw result, named `*.report.json`. This is the
|
- Machine report JSON beside the raw result, named `*.report.json`. This is the
|
||||||
bounded projection for analysis.
|
bounded projection for analysis.
|
||||||
|
- Human report Markdown beside the raw result, named `*.report.md`. This is
|
||||||
|
useful for reviewing results after workspace cleanup.
|
||||||
- Human report Markdown inside the trial workspace, named `final-report.md`.
|
- Human report Markdown inside the trial workspace, named `final-report.md`.
|
||||||
This is the file to read first during manual review.
|
This is the file to read first during manual review.
|
||||||
|
|
||||||
Manual audits add one more file:
|
Manual audits add one more file:
|
||||||
|
|
||||||
- `manual-audit.yaml` inside the trial workspace. Re-running the audit command
|
- `manual-audit.yaml` inside the trial workspace. Re-running the audit command
|
||||||
regenerates the human and machine report projections without mutating the raw
|
regenerates the workspace Markdown, result Markdown, and machine report
|
||||||
result JSON.
|
projections without mutating the raw result JSON.
|
||||||
|
|
||||||
## Run One Trial
|
## Run One Trial
|
||||||
|
|
||||||
@@ -57,6 +59,45 @@ The runner prints summary JSON with the trial classification, result path, and
|
|||||||
report paths. Read the corresponding `final-report.md` before trusting the
|
report paths. Read the corresponding `final-report.md` before trusting the
|
||||||
classification.
|
classification.
|
||||||
|
|
||||||
|
Use bounded concurrency when collecting larger samples:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
uv run python examples/agent_challenges/run_trials.py `
|
||||||
|
--challenge examples/agent_challenges/report_workflow_challenge/challenge.yaml `
|
||||||
|
--instruction-profile skills `
|
||||||
|
--model opencode/deepseek-v4-flash-free `
|
||||||
|
--variant high `
|
||||||
|
--trials 5 `
|
||||||
|
--concurrency 2 `
|
||||||
|
--attach http://127.0.0.1:4096
|
||||||
|
```
|
||||||
|
|
||||||
|
`--trials 5 --concurrency 2` creates five unique trial workspaces and lets at
|
||||||
|
most two OpenCode subprocesses run at once.
|
||||||
|
|
||||||
|
## Run The Default Matrix
|
||||||
|
|
||||||
|
The Python matrix runner expands the bundled challenges, instruction profiles,
|
||||||
|
and default model set, then schedules them with one global concurrency limit:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
uv run python examples/agent_challenges/run_matrix.py `
|
||||||
|
--trials 5 `
|
||||||
|
--concurrency 2 `
|
||||||
|
--attach http://127.0.0.1:4096
|
||||||
|
```
|
||||||
|
|
||||||
|
The PowerShell helper is now only a convenience wrapper around the Python
|
||||||
|
runner:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
.\examples\agent_challenges\run_matrix.ps1 -Trials 5 -Concurrency 2
|
||||||
|
```
|
||||||
|
|
||||||
|
Avoid treating high-concurrency runs as the same dataset as sequential runs.
|
||||||
|
Large concurrency can measure provider queueing, OpenCode timeouts, and local
|
||||||
|
machine contention rather than workflow UX.
|
||||||
|
|
||||||
## Instruction Profiles
|
## Instruction Profiles
|
||||||
|
|
||||||
Use profiles to separate product usability from instruction quality:
|
Use profiles to separate product usability from instruction quality:
|
||||||
@@ -140,10 +181,10 @@ uv run python examples/agent_challenges/save_manual_audit.py `
|
|||||||
--notes "Technical workflow run succeeded, but the trial is invalid because the agent inspected an existing solution."
|
--notes "Technical workflow run succeeded, but the trial is invalid because the agent inspected an existing solution."
|
||||||
```
|
```
|
||||||
|
|
||||||
Use `--manual-classification valid` when the run satisfies the challenge and no
|
Use `--manual-classification pass` when the run satisfies the challenge and no
|
||||||
disqualifying evidence is found. Use `invalid` when the workflow ran but the
|
disqualifying evidence is found. Use `fail` when the task was not completed.
|
||||||
evaluation is contaminated or bypassed. Use another classification only when the
|
Use `invalid` when the workflow ran but the evaluation is contaminated,
|
||||||
challenge manifest or report schema explicitly defines it.
|
bypassed, or otherwise not usable as clean benchmark evidence.
|
||||||
|
|
||||||
## Common Invalid Patterns
|
## Common Invalid Patterns
|
||||||
|
|
||||||
@@ -181,4 +222,3 @@ Avoid stronger claims until the matrix has enough audited trials:
|
|||||||
Use the challenge evidence as product-design feedback first. Treat aggregate
|
Use the challenge evidence as product-design feedback first. Treat aggregate
|
||||||
model benchmarking as a later result once trial counts and audit rules are
|
model benchmarking as a later result once trial counts and audit rules are
|
||||||
stable.
|
stable.
|
||||||
|
|
||||||
|
|||||||
@@ -100,7 +100,7 @@ Challenge directories become data-only:
|
|||||||
```text
|
```text
|
||||||
examples/agent_challenges/
|
examples/agent_challenges/
|
||||||
run_trials.py
|
run_trials.py
|
||||||
extract_metrics.py
|
metrics.py
|
||||||
save_trial_report.py
|
save_trial_report.py
|
||||||
save_manual_audit.py
|
save_manual_audit.py
|
||||||
base-prompt.md
|
base-prompt.md
|
||||||
|
|||||||
@@ -253,12 +253,12 @@ def manual_audit_from_v2_result(
|
|||||||
output_path: Path
|
output_path: Path
|
||||||
if workspace is not None:
|
if workspace is not None:
|
||||||
output_path = workspace / output_name
|
output_path = workspace / output_name
|
||||||
|
else:
|
||||||
|
output_path = Path(output_name)
|
||||||
output_path.write_text(
|
output_path.write_text(
|
||||||
yaml.safe_dump(audit, sort_keys=False, allow_unicode=True),
|
yaml.safe_dump(audit, sort_keys=False, allow_unicode=True),
|
||||||
encoding="utf-8",
|
encoding="utf-8",
|
||||||
)
|
)
|
||||||
else:
|
|
||||||
output_path = Path(output_name)
|
|
||||||
|
|
||||||
return output_path, audit
|
return output_path, audit
|
||||||
|
|
||||||
@@ -271,6 +271,7 @@ class V2AuditPaths:
|
|||||||
audit: Path
|
audit: Path
|
||||||
markdown: Path
|
markdown: Path
|
||||||
machine: Path
|
machine: Path
|
||||||
|
results_markdown: Path | None = None
|
||||||
|
|
||||||
|
|
||||||
def save_v2_manual_audit(
|
def save_v2_manual_audit(
|
||||||
@@ -350,10 +351,12 @@ def save_v2_manual_audit(
|
|||||||
|
|
||||||
markdown_path = workspace_path / "final-report.md"
|
markdown_path = workspace_path / "final-report.md"
|
||||||
machine_path = result_path.with_suffix(".report.json")
|
machine_path = result_path.with_suffix(".report.json")
|
||||||
|
results_markdown_path = result_path.with_suffix(".report.md")
|
||||||
write_trial_report_projections(
|
write_trial_report_projections(
|
||||||
trial_report,
|
trial_report,
|
||||||
markdown_path=markdown_path,
|
markdown_path=markdown_path,
|
||||||
machine_path=machine_path,
|
machine_path=machine_path,
|
||||||
|
extra_markdown_paths=[results_markdown_path],
|
||||||
)
|
)
|
||||||
|
|
||||||
audit_path = workspace_path / "manual-audit.yaml"
|
audit_path = workspace_path / "manual-audit.yaml"
|
||||||
@@ -366,4 +369,5 @@ def save_v2_manual_audit(
|
|||||||
audit=audit_path,
|
audit=audit_path,
|
||||||
markdown=markdown_path,
|
markdown=markdown_path,
|
||||||
machine=machine_path,
|
machine=machine_path,
|
||||||
|
results_markdown=results_markdown_path,
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -47,15 +47,15 @@ The hard timeout ceiling is 3,600 seconds per trial.
|
|||||||
|
|
||||||
## Default Behavior
|
## Default Behavior
|
||||||
|
|
||||||
By default the harness does not start a `wf-rpc-server`. It prompts agents to
|
The V2 harness does not start a `wf-rpc-server`. It prompts agents to use a
|
||||||
use a per-trial configured local CLI path. Use `--start-server` when the trial
|
per-trial configured local CLI path.
|
||||||
should exercise the JSON-RPC server path.
|
|
||||||
|
|
||||||
## Workspace Layout
|
## Workspace Layout
|
||||||
|
|
||||||
- `workspace_template/` contains files copied into each isolated trial workspace.
|
- `workspace_template/` contains files copied into each isolated trial workspace.
|
||||||
- `workspaces/` holds per-trial workspaces (gitignored).
|
- `workspaces/` holds per-trial workspaces (gitignored).
|
||||||
- `results/` holds per-trial result JSON files (gitignored).
|
- `results/` holds per-trial raw result JSON and report projections
|
||||||
|
(gitignored).
|
||||||
- `challenge.yaml` declares the manifest, source, server, and report schema.
|
- `challenge.yaml` declares the manifest, source, server, and report schema.
|
||||||
- `challenge-prompt.md` contains the task-specific prompt.
|
- `challenge-prompt.md` contains the task-specific prompt.
|
||||||
|
|
||||||
|
|||||||
@@ -73,6 +73,15 @@ def opencode_text_results(stdout: str) -> list[dict[str, Any]]:
|
|||||||
events.append({"text": event_text, "event": event})
|
events.append({"text": event_text, "event": event})
|
||||||
return events
|
return events
|
||||||
|
|
||||||
|
if isinstance(parsed, list):
|
||||||
|
events = []
|
||||||
|
for event in parsed:
|
||||||
|
if not isinstance(event, dict):
|
||||||
|
continue
|
||||||
|
event_text = _event_text(event)
|
||||||
|
if event_text is not None:
|
||||||
|
events.append({"text": event_text, "event": event})
|
||||||
|
return events
|
||||||
if not isinstance(parsed, dict):
|
if not isinstance(parsed, dict):
|
||||||
return []
|
return []
|
||||||
event_text = _event_text(parsed)
|
event_text = _event_text(parsed)
|
||||||
|
|||||||
@@ -74,7 +74,12 @@ def _classify_path(
|
|||||||
return "search_intent"
|
return "search_intent"
|
||||||
|
|
||||||
try:
|
try:
|
||||||
p = Path(path_str).resolve()
|
raw_path = Path(path_str)
|
||||||
|
p = (
|
||||||
|
raw_path.resolve()
|
||||||
|
if raw_path.is_absolute()
|
||||||
|
else (Path(workspace_root) / raw_path).resolve()
|
||||||
|
)
|
||||||
except OSError, ValueError:
|
except OSError, ValueError:
|
||||||
return "unknown"
|
return "unknown"
|
||||||
|
|
||||||
@@ -173,7 +178,6 @@ def evaluate_policy(
|
|||||||
"workspace",
|
"workspace",
|
||||||
"supplied_skills",
|
"supplied_skills",
|
||||||
"search_intent",
|
"search_intent",
|
||||||
"example_implementation",
|
|
||||||
"unknown",
|
"unknown",
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -191,7 +195,7 @@ def evaluate_policy(
|
|||||||
if category == "existing_solution":
|
if category == "existing_solution":
|
||||||
disallowed_reads.append(path_str)
|
disallowed_reads.append(path_str)
|
||||||
elif profile == InstructionProfile.NONE:
|
elif profile == InstructionProfile.NONE:
|
||||||
if category not in ("workspace", "unknown"):
|
if category not in ("workspace", "search_intent", "unknown"):
|
||||||
disallowed_reads.append(path_str)
|
disallowed_reads.append(path_str)
|
||||||
elif profile == InstructionProfile.SKILLS:
|
elif profile == InstructionProfile.SKILLS:
|
||||||
if category not in allowed_skills_categories:
|
if category not in allowed_skills_categories:
|
||||||
|
|||||||
@@ -93,6 +93,10 @@ class TrialReport(StrictReportModel):
|
|||||||
|
|
||||||
_MAX_FINAL_TEXT_CHARS = 8_000
|
_MAX_FINAL_TEXT_CHARS = 8_000
|
||||||
_MAX_COMMAND_DETAIL_CHARS = 1_000
|
_MAX_COMMAND_DETAIL_CHARS = 1_000
|
||||||
|
_MAX_SELF_REPORT_STRING_CHARS = 2_000
|
||||||
|
_MAX_SELF_REPORT_LIST_ITEMS = 50
|
||||||
|
_MAX_SELF_REPORT_DICT_ITEMS = 50
|
||||||
|
_MAX_SELF_REPORT_DEPTH = 4
|
||||||
|
|
||||||
|
|
||||||
def _build_identity(
|
def _build_identity(
|
||||||
@@ -243,7 +247,7 @@ def _build_trial_report(
|
|||||||
agent_self_report: dict[str, Any] | None = None
|
agent_self_report: dict[str, Any] | None = None
|
||||||
challenge_report = result.get("challenge_report")
|
challenge_report = result.get("challenge_report")
|
||||||
if isinstance(challenge_report, dict):
|
if isinstance(challenge_report, dict):
|
||||||
agent_self_report = challenge_report
|
agent_self_report = _bounded_report_mapping(challenge_report)
|
||||||
|
|
||||||
final_agent_answer: str | None = None
|
final_agent_answer: str | None = None
|
||||||
parsed = result.get("parsed")
|
parsed = result.get("parsed")
|
||||||
@@ -531,6 +535,46 @@ def _dict_any(value: object) -> dict[str, Any]:
|
|||||||
return value if isinstance(value, dict) else {}
|
return value if isinstance(value, dict) else {}
|
||||||
|
|
||||||
|
|
||||||
|
def _bounded_report_mapping(value: dict[object, object]) -> dict[str, Any]:
|
||||||
|
"""Project agent YAML to bounded JSON-like data for reports.
|
||||||
|
|
||||||
|
The raw challenge report is untrusted model output. Keep useful keys, but
|
||||||
|
cap strings, container sizes, and nesting so generated report projections
|
||||||
|
cannot balloon or carry arbitrary deep payloads.
|
||||||
|
"""
|
||||||
|
bounded: dict[str, Any] = {}
|
||||||
|
for index, (key, item) in enumerate(value.items()):
|
||||||
|
if index >= _MAX_SELF_REPORT_DICT_ITEMS:
|
||||||
|
break
|
||||||
|
if not isinstance(key, str):
|
||||||
|
continue
|
||||||
|
bounded[key] = _bounded_report_value(item, depth=0)
|
||||||
|
return bounded
|
||||||
|
|
||||||
|
|
||||||
|
def _bounded_report_value(value: object, *, depth: int) -> Any:
|
||||||
|
if depth >= _MAX_SELF_REPORT_DEPTH:
|
||||||
|
return _str(value)[:_MAX_SELF_REPORT_STRING_CHARS]
|
||||||
|
if isinstance(value, str):
|
||||||
|
return value[:_MAX_SELF_REPORT_STRING_CHARS]
|
||||||
|
if isinstance(value, bool | int | float) or value is None:
|
||||||
|
return value
|
||||||
|
if isinstance(value, (list, tuple)):
|
||||||
|
return [
|
||||||
|
_bounded_report_value(item, depth=depth + 1)
|
||||||
|
for item in value[:_MAX_SELF_REPORT_LIST_ITEMS]
|
||||||
|
]
|
||||||
|
if isinstance(value, dict):
|
||||||
|
bounded: dict[str, Any] = {}
|
||||||
|
for index, (key, item) in enumerate(value.items()):
|
||||||
|
if index >= _MAX_SELF_REPORT_DICT_ITEMS:
|
||||||
|
break
|
||||||
|
if isinstance(key, str):
|
||||||
|
bounded[key] = _bounded_report_value(item, depth=depth + 1)
|
||||||
|
return bounded
|
||||||
|
return _str(value)[:_MAX_SELF_REPORT_STRING_CHARS]
|
||||||
|
|
||||||
|
|
||||||
def _parse_errors(result: dict[str, object]) -> dict[str, dict[str, str]]:
|
def _parse_errors(result: dict[str, object]) -> dict[str, dict[str, str]]:
|
||||||
errors: dict[str, dict[str, str]] = {}
|
errors: dict[str, dict[str, str]] = {}
|
||||||
parse_error = result.get("parse_error")
|
parse_error = result.get("parse_error")
|
||||||
|
|||||||
@@ -44,7 +44,8 @@ The hard timeout ceiling is 3,600 seconds per trial.
|
|||||||
|
|
||||||
- `workspace_template/` holds local store ignore rules (gitignored contents).
|
- `workspace_template/` holds local store ignore rules (gitignored contents).
|
||||||
- `workspaces/` holds per-trial workspaces (gitignored).
|
- `workspaces/` holds per-trial workspaces (gitignored).
|
||||||
- `results/` holds per-trial result JSON files (gitignored).
|
- `results/` holds per-trial raw result JSON and report projections
|
||||||
|
(gitignored).
|
||||||
- `challenge.yaml` declares the manifest, source, server, and report schema.
|
- `challenge.yaml` declares the manifest, source, server, and report schema.
|
||||||
- `challenge-prompt.md` contains the task-specific prompt.
|
- `challenge-prompt.md` contains the task-specific prompt.
|
||||||
|
|
||||||
|
|||||||
@@ -7,6 +7,8 @@ from collections.abc import Sequence
|
|||||||
from dataclasses import dataclass
|
from dataclasses import dataclass
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
|
import yaml
|
||||||
|
|
||||||
from examples.agent_challenges.classification import extract_challenge_report
|
from examples.agent_challenges.classification import extract_challenge_report
|
||||||
from examples.agent_challenges.opencode_io import parse_opencode_output, result_text
|
from examples.agent_challenges.opencode_io import parse_opencode_output, result_text
|
||||||
from examples.agent_challenges.report_models import TrialReport
|
from examples.agent_challenges.report_models import TrialReport
|
||||||
@@ -245,8 +247,13 @@ def render_trial_report_markdown(report: TrialReport) -> str:
|
|||||||
lines.append("")
|
lines.append("")
|
||||||
if report.agent_self_report is not None:
|
if report.agent_self_report is not None:
|
||||||
lines.append("```yaml")
|
lines.append("```yaml")
|
||||||
for key, value in report.agent_self_report.items():
|
lines.append(
|
||||||
lines.append(f"{key}: {value}")
|
yaml.safe_dump(
|
||||||
|
report.agent_self_report,
|
||||||
|
sort_keys=False,
|
||||||
|
allow_unicode=True,
|
||||||
|
).rstrip()
|
||||||
|
)
|
||||||
lines.append("```")
|
lines.append("```")
|
||||||
else:
|
else:
|
||||||
lines.append("No agent self-report captured.")
|
lines.append("No agent self-report captured.")
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
param(
|
param(
|
||||||
[string]$AttachUrl = "http://127.0.0.1:8192",
|
[string]$AttachUrl = "http://127.0.0.1:8192",
|
||||||
[int]$Trials = 5,
|
[int]$Trials = 5,
|
||||||
|
[int]$Concurrency = 2,
|
||||||
[int]$TimeoutSeconds = 3600,
|
[int]$TimeoutSeconds = 3600,
|
||||||
[object[]]$models = @( # object[] because ModelProfile is not known yet
|
[object[]]$models = @( # object[] because ModelProfile is not known yet
|
||||||
[ModelProfile]::new("opencode/deepseek-v4-flash-free", "max"),
|
[ModelProfile]::new("opencode/deepseek-v4-flash-free", "max"),
|
||||||
@@ -33,35 +34,28 @@ $ErrorActionPreference = "Stop"
|
|||||||
|
|
||||||
[ModelProfile[]]$models = $models # cast should fail if there are any non-ModelProfile objects in the array
|
[ModelProfile[]]$models = $models # cast should fail if there are any non-ModelProfile objects in the array
|
||||||
|
|
||||||
# Run from the repository root no matter where the script is invoked.
|
Push-Location (Get-Item $PSScriptRoot).Parent.Parent.FullName
|
||||||
Set-Location (Get-Item $PSScriptRoot).Parent.Parent.FullName
|
|
||||||
|
|
||||||
$profiles = @(
|
try {
|
||||||
"none",
|
$argsList = @(
|
||||||
"skills",
|
"run",
|
||||||
"all"
|
"python",
|
||||||
|
"examples/agent_challenges/run_matrix.py",
|
||||||
|
"--trials",
|
||||||
|
"$Trials",
|
||||||
|
"--concurrency",
|
||||||
|
"$Concurrency",
|
||||||
|
"--attach",
|
||||||
|
"$AttachUrl",
|
||||||
|
"--timeout-seconds",
|
||||||
|
"$TimeoutSeconds"
|
||||||
)
|
)
|
||||||
|
|
||||||
$challenges = @(
|
|
||||||
"examples/agent_challenges/browser_click_challenge/challenge.yaml",
|
|
||||||
"examples/agent_challenges/report_workflow_challenge/challenge.yaml"
|
|
||||||
)
|
|
||||||
|
|
||||||
foreach ($challenge in $challenges) {
|
|
||||||
foreach ($challengeProfile in $profiles) {
|
|
||||||
foreach ($model in $models) {
|
foreach ($model in $models) {
|
||||||
Write-Host ""
|
$argsList += "--model"
|
||||||
Write-Host "==> challenge=$challenge profile=$challengeProfile model=$model trials=$Trials"
|
$argsList += "$($model.Model)=$($model.Variant)"
|
||||||
uv run python examples/agent_challenges/run_trials.py `
|
|
||||||
--challenge $challenge `
|
|
||||||
--instruction-profile $challengeProfile `
|
|
||||||
--model $model.Model `
|
|
||||||
--variant $model.Variant `
|
|
||||||
--trials $Trials `
|
|
||||||
--attach $AttachUrl `
|
|
||||||
--timeout-seconds $TimeoutSeconds
|
|
||||||
}
|
}
|
||||||
|
uv @argsList
|
||||||
}
|
}
|
||||||
}
|
finally {
|
||||||
|
|
||||||
Pop-Location
|
Pop-Location
|
||||||
|
}
|
||||||
|
|||||||
@@ -0,0 +1,219 @@
|
|||||||
|
"""Run a challenge/profile/model matrix with bounded global concurrency."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
try:
|
||||||
|
from .manifests import load_challenge_manifest
|
||||||
|
from .models import InstructionProfile, LoadedChallenge
|
||||||
|
from .runner import run_v2_trial
|
||||||
|
from .workspace import PROJECT_ROOT, starting_trial_index
|
||||||
|
except ImportError:
|
||||||
|
_project_root = Path(__file__).resolve().parents[2]
|
||||||
|
if str(_project_root) not in sys.path:
|
||||||
|
sys.path.insert(0, str(_project_root))
|
||||||
|
from examples.agent_challenges.manifests import load_challenge_manifest
|
||||||
|
from examples.agent_challenges.models import InstructionProfile, LoadedChallenge
|
||||||
|
from examples.agent_challenges.runner import run_v2_trial
|
||||||
|
from examples.agent_challenges.workspace import PROJECT_ROOT, starting_trial_index
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_CHALLENGES = (
|
||||||
|
PROJECT_ROOT / "examples/agent_challenges/browser_click_challenge/challenge.yaml",
|
||||||
|
PROJECT_ROOT / "examples/agent_challenges/report_workflow_challenge/challenge.yaml",
|
||||||
|
)
|
||||||
|
DEFAULT_PROFILES = (
|
||||||
|
InstructionProfile.NONE,
|
||||||
|
InstructionProfile.SKILLS,
|
||||||
|
InstructionProfile.ALL,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class ModelProfile:
|
||||||
|
model: str
|
||||||
|
variant: str
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_MODELS = (
|
||||||
|
ModelProfile("opencode/deepseek-v4-flash-free", "max"),
|
||||||
|
ModelProfile("opencode/mimo-v2.5-free", "high"),
|
||||||
|
ModelProfile("opencode/nemotron-3-ultra-free", "high"),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True, slots=True)
|
||||||
|
class MatrixTask:
|
||||||
|
challenge: LoadedChallenge
|
||||||
|
profile: InstructionProfile
|
||||||
|
model: str
|
||||||
|
variant: str
|
||||||
|
index: int
|
||||||
|
workspaces_dir: Path
|
||||||
|
results_dir: Path
|
||||||
|
|
||||||
|
|
||||||
|
def parse_model_profile(raw: str) -> ModelProfile:
|
||||||
|
"""Parse MODEL or MODEL=VARIANT values from CLI flags."""
|
||||||
|
model, separator, variant = raw.partition("=")
|
||||||
|
if not model:
|
||||||
|
raise ValueError("model cannot be empty")
|
||||||
|
return ModelProfile(model=model, variant=variant if separator else "high")
|
||||||
|
|
||||||
|
|
||||||
|
def build_matrix_tasks(
|
||||||
|
*,
|
||||||
|
challenges: list[LoadedChallenge],
|
||||||
|
profiles: list[InstructionProfile],
|
||||||
|
models: list[ModelProfile],
|
||||||
|
trials: int,
|
||||||
|
) -> list[MatrixTask]:
|
||||||
|
"""Allocate unique trial indices for each challenge/model across profiles."""
|
||||||
|
tasks: list[MatrixTask] = []
|
||||||
|
for challenge in challenges:
|
||||||
|
results_dir = challenge.root / "results"
|
||||||
|
workspaces_dir = challenge.root / "workspaces"
|
||||||
|
for model in models:
|
||||||
|
next_index = starting_trial_index(
|
||||||
|
model=model.model,
|
||||||
|
results_dir=results_dir,
|
||||||
|
workspaces_dir=workspaces_dir,
|
||||||
|
)
|
||||||
|
for profile in profiles:
|
||||||
|
for _ in range(trials):
|
||||||
|
tasks.append(
|
||||||
|
MatrixTask(
|
||||||
|
challenge=challenge,
|
||||||
|
profile=profile,
|
||||||
|
model=model.model,
|
||||||
|
variant=model.variant,
|
||||||
|
index=next_index,
|
||||||
|
workspaces_dir=workspaces_dir,
|
||||||
|
results_dir=results_dir,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
next_index += 1
|
||||||
|
return tasks
|
||||||
|
|
||||||
|
|
||||||
|
def _summary_from_result(task: MatrixTask, result: dict[str, Any]) -> dict[str, object]:
|
||||||
|
return {
|
||||||
|
"challenge": task.challenge.manifest.id,
|
||||||
|
"instruction_profile": task.profile.value,
|
||||||
|
"model": task.model,
|
||||||
|
"variant": task.variant,
|
||||||
|
"index": task.index,
|
||||||
|
"task_outcome": result["task_outcome"],
|
||||||
|
"evaluation_validity": result["evaluation_validity"],
|
||||||
|
"duration_seconds": result["duration_seconds"],
|
||||||
|
"result_path": result.get("result_path"),
|
||||||
|
"report_paths": result.get("report_paths"),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: list[str] | None = None) -> int:
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument(
|
||||||
|
"--challenge",
|
||||||
|
action="append",
|
||||||
|
type=Path,
|
||||||
|
default=None,
|
||||||
|
help="Path to challenge.yaml. May be repeated. Defaults to all bundled challenges.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--instruction-profile",
|
||||||
|
action="append",
|
||||||
|
choices=[p.value for p in InstructionProfile],
|
||||||
|
default=None,
|
||||||
|
help="Instruction profile. May be repeated. Defaults to none, skills, all.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--model",
|
||||||
|
action="append",
|
||||||
|
default=None,
|
||||||
|
help=(
|
||||||
|
"Model profile as MODEL or MODEL=VARIANT. May be repeated. "
|
||||||
|
"Defaults to the standard free OpenCode model set."
|
||||||
|
),
|
||||||
|
)
|
||||||
|
parser.add_argument("--trials", type=int, default=5)
|
||||||
|
parser.add_argument("--concurrency", type=int, default=2)
|
||||||
|
parser.add_argument("--timeout-seconds", type=int, default=3600)
|
||||||
|
parser.add_argument("--attach", dest="attach_url", default=None)
|
||||||
|
parser.add_argument(
|
||||||
|
"--instruction-bundle",
|
||||||
|
type=Path,
|
||||||
|
default=PROJECT_ROOT
|
||||||
|
/ "examples"
|
||||||
|
/ "agent_challenges"
|
||||||
|
/ "instruction_bundles"
|
||||||
|
/ "workflow_cli.yaml",
|
||||||
|
)
|
||||||
|
args = parser.parse_args(argv)
|
||||||
|
|
||||||
|
if args.trials < 1:
|
||||||
|
parser.error("--trials must be >= 1")
|
||||||
|
if args.concurrency < 1:
|
||||||
|
parser.error("--concurrency must be >= 1")
|
||||||
|
|
||||||
|
try:
|
||||||
|
challenges = [
|
||||||
|
load_challenge_manifest(path)
|
||||||
|
for path in (args.challenge or DEFAULT_CHALLENGES)
|
||||||
|
]
|
||||||
|
profiles = [
|
||||||
|
InstructionProfile(value)
|
||||||
|
for value in (
|
||||||
|
args.instruction_profile or [p.value for p in DEFAULT_PROFILES]
|
||||||
|
)
|
||||||
|
]
|
||||||
|
models = [parse_model_profile(raw) for raw in (args.model or [])] or list(
|
||||||
|
DEFAULT_MODELS
|
||||||
|
)
|
||||||
|
except ValueError as exc:
|
||||||
|
parser.error(str(exc))
|
||||||
|
|
||||||
|
tasks = build_matrix_tasks(
|
||||||
|
challenges=challenges,
|
||||||
|
profiles=profiles,
|
||||||
|
models=models,
|
||||||
|
trials=args.trials,
|
||||||
|
)
|
||||||
|
|
||||||
|
def _run_task(task: MatrixTask) -> dict[str, object]:
|
||||||
|
result = run_v2_trial(
|
||||||
|
task.challenge,
|
||||||
|
profile=task.profile,
|
||||||
|
model=task.model,
|
||||||
|
variant=task.variant,
|
||||||
|
index=task.index,
|
||||||
|
workspaces_dir=task.workspaces_dir,
|
||||||
|
results_dir=task.results_dir,
|
||||||
|
instruction_bundle=args.instruction_bundle,
|
||||||
|
timeout_seconds=args.timeout_seconds,
|
||||||
|
attach_url=args.attach_url,
|
||||||
|
)
|
||||||
|
return _summary_from_result(task, result)
|
||||||
|
|
||||||
|
summaries: list[dict[str, object]] = []
|
||||||
|
with ThreadPoolExecutor(max_workers=min(args.concurrency, len(tasks))) as pool:
|
||||||
|
futures = [pool.submit(_run_task, task) for task in tasks]
|
||||||
|
for future in as_completed(futures):
|
||||||
|
summary = future.result()
|
||||||
|
summaries.append(summary)
|
||||||
|
print(json.dumps(summary, sort_keys=True))
|
||||||
|
|
||||||
|
success_count = sum(1 for item in summaries if item["task_outcome"] == "success")
|
||||||
|
print(json.dumps({"success_count": success_count, "trial_count": len(summaries)}))
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
@@ -5,6 +5,7 @@ from __future__ import annotations
|
|||||||
import argparse
|
import argparse
|
||||||
import json
|
import json
|
||||||
import sys
|
import sys
|
||||||
|
from concurrent.futures import ThreadPoolExecutor, as_completed
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
try:
|
try:
|
||||||
@@ -43,6 +44,12 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
parser.add_argument("--model", default="opencode/mimo-v2.5-free")
|
parser.add_argument("--model", default="opencode/mimo-v2.5-free")
|
||||||
parser.add_argument("--variant", default="high")
|
parser.add_argument("--variant", default="high")
|
||||||
parser.add_argument("--trials", type=int, default=1)
|
parser.add_argument("--trials", type=int, default=1)
|
||||||
|
parser.add_argument(
|
||||||
|
"--concurrency",
|
||||||
|
type=int,
|
||||||
|
default=1,
|
||||||
|
help="Maximum number of trials for this model/profile/challenge to run at once.",
|
||||||
|
)
|
||||||
parser.add_argument("--timeout-seconds", type=int, default=3600)
|
parser.add_argument("--timeout-seconds", type=int, default=3600)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--attach",
|
"--attach",
|
||||||
@@ -61,13 +68,12 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
/ "instruction_bundles"
|
/ "instruction_bundles"
|
||||||
/ "workflow_cli.yaml",
|
/ "workflow_cli.yaml",
|
||||||
)
|
)
|
||||||
parser.add_argument("--server-url", default=None)
|
|
||||||
parser.add_argument("--start-server", action="store_true", default=False)
|
|
||||||
parser.add_argument("--server-port", type=int, default=8772)
|
|
||||||
args = parser.parse_args(argv)
|
args = parser.parse_args(argv)
|
||||||
|
|
||||||
if args.trials < 1:
|
if args.trials < 1:
|
||||||
parser.error("--trials must be >= 1")
|
parser.error("--trials must be >= 1")
|
||||||
|
if args.concurrency < 1:
|
||||||
|
parser.error("--concurrency must be >= 1")
|
||||||
|
|
||||||
challenge = load_challenge_manifest(args.challenge)
|
challenge = load_challenge_manifest(args.challenge)
|
||||||
profile = InstructionProfile(args.instruction_profile)
|
profile = InstructionProfile(args.instruction_profile)
|
||||||
@@ -81,8 +87,7 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
workspaces_dir=workspaces_dir,
|
workspaces_dir=workspaces_dir,
|
||||||
)
|
)
|
||||||
|
|
||||||
summaries: list[dict[str, object]] = []
|
def _run_one(index: int) -> dict[str, object]:
|
||||||
for index in range(first_index, first_index + args.trials):
|
|
||||||
result = run_v2_trial(
|
result = run_v2_trial(
|
||||||
challenge,
|
challenge,
|
||||||
profile=profile,
|
profile=profile,
|
||||||
@@ -95,8 +100,7 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
timeout_seconds=args.timeout_seconds,
|
timeout_seconds=args.timeout_seconds,
|
||||||
attach_url=args.attach_url,
|
attach_url=args.attach_url,
|
||||||
)
|
)
|
||||||
summaries.append(
|
return {
|
||||||
{
|
|
||||||
"index": index,
|
"index": index,
|
||||||
"task_outcome": result["task_outcome"],
|
"task_outcome": result["task_outcome"],
|
||||||
"evaluation_validity": result["evaluation_validity"],
|
"evaluation_validity": result["evaluation_validity"],
|
||||||
@@ -104,8 +108,15 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
"result_path": result.get("result_path"),
|
"result_path": result.get("result_path"),
|
||||||
"report_paths": result.get("report_paths"),
|
"report_paths": result.get("report_paths"),
|
||||||
}
|
}
|
||||||
)
|
|
||||||
print(json.dumps(summaries[-1], sort_keys=True))
|
summaries: list[dict[str, object]] = []
|
||||||
|
indices = list(range(first_index, first_index + args.trials))
|
||||||
|
with ThreadPoolExecutor(max_workers=min(args.concurrency, args.trials)) as pool:
|
||||||
|
futures = [pool.submit(_run_one, index) for index in indices]
|
||||||
|
for future in as_completed(futures):
|
||||||
|
summary = future.result()
|
||||||
|
summaries.append(summary)
|
||||||
|
print(json.dumps(summary, sort_keys=True))
|
||||||
|
|
||||||
success_count = sum(1 for s in summaries if s["task_outcome"] == "success")
|
success_count = sum(1 for s in summaries if s["task_outcome"] == "success")
|
||||||
print(json.dumps({"success_count": success_count, "trial_count": len(summaries)}))
|
print(json.dumps({"success_count": success_count, "trial_count": len(summaries)}))
|
||||||
|
|||||||
@@ -664,5 +664,8 @@ def run_v2_trial(
|
|||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
report_generation_error = str(exc)
|
report_generation_error = str(exc)
|
||||||
result["report_generation_error"] = report_generation_error
|
result["report_generation_error"] = report_generation_error
|
||||||
|
result_path.write_text(
|
||||||
|
json.dumps(result, indent=2, sort_keys=True), encoding="utf-8"
|
||||||
|
)
|
||||||
|
|
||||||
return result
|
return result
|
||||||
|
|||||||
@@ -48,11 +48,13 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
if _is_v2_result(result_path):
|
if _is_v2_result(result_path):
|
||||||
if args.from_report is not None:
|
if args.from_report is not None:
|
||||||
parser.error("--from-report is not supported for V2 results")
|
parser.error("--from-report is not supported for V2 results")
|
||||||
read_overrides = dict(_parse_bool_assignment(item) for item in args.set_read)
|
try:
|
||||||
|
read_overrides = dict(
|
||||||
|
_parse_bool_assignment(item) for item in args.set_read
|
||||||
|
)
|
||||||
evidence_overrides = dict(
|
evidence_overrides = dict(
|
||||||
_parse_value_assignment(item) for item in args.set_evidence
|
_parse_value_assignment(item) for item in args.set_evidence
|
||||||
)
|
)
|
||||||
try:
|
|
||||||
paths = save_v2_manual_audit(
|
paths = save_v2_manual_audit(
|
||||||
result_path,
|
result_path,
|
||||||
official_outcome=args.manual_classification,
|
official_outcome=args.manual_classification,
|
||||||
@@ -69,6 +71,11 @@ def main(argv: list[str] | None = None) -> int:
|
|||||||
"audit": paths.audit.as_posix(),
|
"audit": paths.audit.as_posix(),
|
||||||
"markdown": paths.markdown.as_posix(),
|
"markdown": paths.markdown.as_posix(),
|
||||||
"machine": paths.machine.as_posix(),
|
"machine": paths.machine.as_posix(),
|
||||||
|
"results_markdown": (
|
||||||
|
paths.results_markdown.as_posix()
|
||||||
|
if paths.results_markdown is not None
|
||||||
|
else None
|
||||||
|
),
|
||||||
}
|
}
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -75,7 +75,9 @@ def server_command(*, port: int, config_arg: str) -> list[str]:
|
|||||||
|
|
||||||
|
|
||||||
def _safe_model_name(model: str) -> str:
|
def _safe_model_name(model: str) -> str:
|
||||||
return model.replace("/", "_").replace(":", "_")
|
return (
|
||||||
|
model.replace("/", "_").replace("\\", "_").replace(":", "_").replace("..", "_")
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def prepare_trial_workspace(
|
def prepare_trial_workspace(
|
||||||
@@ -208,6 +210,7 @@ def _load_instruction_bundle(
|
|||||||
if not isinstance(loaded, dict) or not isinstance(loaded.get("files"), list):
|
if not isinstance(loaded, dict) or not isinstance(loaded.get("files"), list):
|
||||||
raise ValueError(f"invalid instruction bundle: {bundle_path}")
|
raise ValueError(f"invalid instruction bundle: {bundle_path}")
|
||||||
entries: list[tuple[str, str]] = []
|
entries: list[tuple[str, str]] = []
|
||||||
|
project_root = PROJECT_ROOT.resolve()
|
||||||
for entry in loaded["files"]:
|
for entry in loaded["files"]:
|
||||||
if not isinstance(entry, dict):
|
if not isinstance(entry, dict):
|
||||||
raise ValueError(f"invalid bundle entry: {entry}")
|
raise ValueError(f"invalid bundle entry: {entry}")
|
||||||
@@ -215,6 +218,19 @@ def _load_instruction_bundle(
|
|||||||
destination = entry.get("destination")
|
destination = entry.get("destination")
|
||||||
if not isinstance(source, str) or not isinstance(destination, str):
|
if not isinstance(source, str) or not isinstance(destination, str):
|
||||||
raise ValueError(f"bundle entry missing source/destination: {entry}")
|
raise ValueError(f"bundle entry missing source/destination: {entry}")
|
||||||
|
source_path = Path(source)
|
||||||
|
destination_path = Path(destination)
|
||||||
|
if source_path.is_absolute() or destination_path.is_absolute():
|
||||||
|
raise ValueError(f"bundle paths must be relative: {entry}")
|
||||||
|
resolved_source = (project_root / source_path).resolve()
|
||||||
|
resolved_destination = (
|
||||||
|
project_root / ".agent" / "skills" / destination_path
|
||||||
|
).resolve()
|
||||||
|
trusted_destination_root = (project_root / ".agent" / "skills").resolve()
|
||||||
|
if not resolved_source.is_relative_to(project_root):
|
||||||
|
raise ValueError(f"bundle source escapes project root: {source}")
|
||||||
|
if not resolved_destination.is_relative_to(trusted_destination_root):
|
||||||
|
raise ValueError(f"bundle destination escapes skill root: {destination}")
|
||||||
entries.append((source, destination))
|
entries.append((source, destination))
|
||||||
return entries
|
return entries
|
||||||
|
|
||||||
|
|||||||
@@ -6,7 +6,7 @@ validated, runnable deployment.
|
|||||||
## Primary Path
|
## Primary Path
|
||||||
|
|
||||||
1. List sources.
|
1. List sources.
|
||||||
- CLI: `wf cap list --format ids`
|
- CLI: `wf source list`
|
||||||
2. List workflow capabilities.
|
2. List workflow capabilities.
|
||||||
- CLI: `wf cap list`
|
- CLI: `wf cap list`
|
||||||
3. Inspect one capability.
|
3. Inspect one capability.
|
||||||
@@ -18,10 +18,12 @@ validated, runnable deployment.
|
|||||||
for common edits.
|
for common edits.
|
||||||
- Use JSON Patch only for general structural edits.
|
- Use JSON Patch only for general structural edits.
|
||||||
6. Save an artifact.
|
6. Save an artifact.
|
||||||
- Draft artifact: `wf draft save`
|
- Draft artifact:
|
||||||
- Complete raw plan: `wf artifact create-from-plan`
|
`wf draft save <workspace_id> --artifact <artifact_id> --version 1 --title "Workflow Title"`
|
||||||
|
- Complete raw plan:
|
||||||
|
`wf artifact create-from-plan workflow.plan.json --artifact <artifact_id> --version 1 --title "Workflow Title" --outcome ok`
|
||||||
7. Save and validate a deployment.
|
7. Save and validate a deployment.
|
||||||
- `wf deploy save` (or `wf deploy create` alias)
|
- `wf deploy save <deployment_id> --artifact <artifact_id> --version 1 --binding <logical_source>=<concrete_source>` (or `wf deploy create` alias)
|
||||||
8. Run the deployment.
|
8. Run the deployment.
|
||||||
9. Inspect the run summary first; read bounded traces only when needed.
|
9. Inspect the run summary first; read bounded traces only when needed.
|
||||||
|
|
||||||
|
|||||||
@@ -156,7 +156,9 @@ def _compact_node(schema: object, related: set[str]) -> object:
|
|||||||
branches = schema.get(source_key)
|
branches = schema.get(source_key)
|
||||||
if not isinstance(branches, list):
|
if not isinstance(branches, list):
|
||||||
continue
|
continue
|
||||||
result["one_of"] = [_compact_node(branch, related) for branch in branches]
|
result["one_of" if source_key == "oneOf" else "any_of"] = [
|
||||||
|
_compact_node(branch, related) for branch in branches
|
||||||
|
]
|
||||||
break
|
break
|
||||||
discriminator = schema.get("discriminator")
|
discriminator = schema.get("discriminator")
|
||||||
if isinstance(discriminator, dict) and isinstance(
|
if isinstance(discriminator, dict) and isinstance(
|
||||||
|
|||||||
@@ -18,9 +18,11 @@ def test_big_doc_links_case_study_and_embeds_evidence_index() -> None:
|
|||||||
links = markdown_links(doc)
|
links = markdown_links(doc)
|
||||||
|
|
||||||
assert any(link.startswith("../../examples/report_workflow") for link in links)
|
assert any(link.startswith("../../examples/report_workflow") for link in links)
|
||||||
assert "Evidence Index" in doc
|
assert re.search(r"^# Evidence Index$", doc, flags=re.MULTILINE)
|
||||||
assert "Core Workflow Lifecycle" in doc
|
assert re.search(r"^## Core Workflow Lifecycle$", doc, flags=re.MULTILINE)
|
||||||
assert "Agent Challenge Evaluation Protocol" in doc
|
assert re.search(
|
||||||
|
r"^## Agent Challenge Evaluation Protocol$", doc, flags=re.MULTILINE
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def test_project_map_links_big_doc() -> None:
|
def test_project_map_links_big_doc() -> None:
|
||||||
|
|||||||
@@ -1,5 +1,6 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import re
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
ROOT = Path(__file__).resolve().parents[2]
|
ROOT = Path(__file__).resolve().parents[2]
|
||||||
@@ -9,6 +10,10 @@ def _read_doc(relative_path: str) -> str:
|
|||||||
return (ROOT / relative_path).read_text(encoding="utf-8")
|
return (ROOT / relative_path).read_text(encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def _markdown_links(text: str) -> set[str]:
|
||||||
|
return set(re.findall(r"(?<!!)\[[^\]]+\]\(([^)]+)\)", text))
|
||||||
|
|
||||||
|
|
||||||
def test_cli_docs_link_source_provider_guide() -> None:
|
def test_cli_docs_link_source_provider_guide() -> None:
|
||||||
text = _read_doc("docs/wf_cli.md")
|
text = _read_doc("docs/wf_cli.md")
|
||||||
|
|
||||||
@@ -32,15 +37,19 @@ def test_source_provider_guide_links_python_runbook() -> None:
|
|||||||
def test_project_map_links_agent_challenge_runbook() -> None:
|
def test_project_map_links_agent_challenge_runbook() -> None:
|
||||||
text = _read_doc("docs/project_map.md")
|
text = _read_doc("docs/project_map.md")
|
||||||
|
|
||||||
assert "runbooks/agent-challenge-evaluation.md" in text
|
assert "runbooks/agent-challenge-evaluation.md" in _markdown_links(text)
|
||||||
|
|
||||||
|
|
||||||
def test_agent_challenge_readmes_link_shared_runbook() -> None:
|
def test_agent_challenge_readmes_link_shared_runbook() -> None:
|
||||||
browser = _read_doc("examples/agent_challenges/browser_click_challenge/README.md")
|
browser = _read_doc("examples/agent_challenges/browser_click_challenge/README.md")
|
||||||
report = _read_doc("examples/agent_challenges/report_workflow_challenge/README.md")
|
report = _read_doc("examples/agent_challenges/report_workflow_challenge/README.md")
|
||||||
|
|
||||||
assert "docs/runbooks/agent-challenge-evaluation.md" in browser
|
assert "../../../docs/runbooks/agent-challenge-evaluation.md" in _markdown_links(
|
||||||
assert "docs/runbooks/agent-challenge-evaluation.md" in report
|
browser
|
||||||
|
)
|
||||||
|
assert "../../../docs/runbooks/agent-challenge-evaluation.md" in _markdown_links(
|
||||||
|
report
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def test_agent_challenge_runbook_defines_validity_and_coverage() -> None:
|
def test_agent_challenge_runbook_defines_validity_and_coverage() -> None:
|
||||||
|
|||||||
@@ -381,6 +381,21 @@ def test_opencode_text_results_preserve_report_before_later_summary() -> None:
|
|||||||
]
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_opencode_text_results_accept_json_array_output() -> None:
|
||||||
|
from examples.agent_challenges.opencode_io import opencode_text_results
|
||||||
|
|
||||||
|
stdout = json.dumps(
|
||||||
|
[
|
||||||
|
{"type": "text", "part": {"text": "first"}},
|
||||||
|
{"type": "text", "part": {"text": "second"}},
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
results = opencode_text_results(stdout)
|
||||||
|
|
||||||
|
assert [result["text"] for result in results] == ["first", "second"]
|
||||||
|
|
||||||
|
|
||||||
def test_policy_evidence_classifies_reads(tmp_path: Path) -> None:
|
def test_policy_evidence_classifies_reads(tmp_path: Path) -> None:
|
||||||
from examples.agent_challenges.metrics import ToolCallEvidence
|
from examples.agent_challenges.metrics import ToolCallEvidence
|
||||||
from examples.agent_challenges.policy import evaluate_policy
|
from examples.agent_challenges.policy import evaluate_policy
|
||||||
@@ -679,6 +694,42 @@ def test_policy_records_broad_globs_without_contaminating(tmp_path: Path) -> Non
|
|||||||
assert policy.disallowed_reads == ()
|
assert policy.disallowed_reads == ()
|
||||||
|
|
||||||
|
|
||||||
|
def test_policy_none_records_broad_globs_without_contaminating(
|
||||||
|
tmp_path: Path,
|
||||||
|
) -> None:
|
||||||
|
from examples.agent_challenges.metrics import ToolCallEvidence
|
||||||
|
from examples.agent_challenges.policy import evaluate_policy
|
||||||
|
|
||||||
|
repository = tmp_path / "repo"
|
||||||
|
workspace = repository / "examples" / "challenge" / "workspaces" / "trial"
|
||||||
|
workspace.mkdir(parents=True)
|
||||||
|
call = ToolCallEvidence(
|
||||||
|
ordinal=1,
|
||||||
|
call_id="glob-1",
|
||||||
|
tool="glob",
|
||||||
|
status="completed",
|
||||||
|
title="Search workspace files",
|
||||||
|
input={"pattern": "*.json"},
|
||||||
|
metadata={},
|
||||||
|
output_chars=10,
|
||||||
|
output_preview="",
|
||||||
|
output_sha256="abc",
|
||||||
|
failed=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
policy = evaluate_policy(
|
||||||
|
"none",
|
||||||
|
[call],
|
||||||
|
workspace_root=workspace,
|
||||||
|
repository_root=repository,
|
||||||
|
workspaces_root=workspace.parent,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert policy.validity.value == "clean"
|
||||||
|
assert policy.reads_by_category["search_intent"] == ("*.json",)
|
||||||
|
assert policy.disallowed_reads == ()
|
||||||
|
|
||||||
|
|
||||||
def test_policy_classifies_example_implementation_reads_separately(
|
def test_policy_classifies_example_implementation_reads_separately(
|
||||||
tmp_path: Path,
|
tmp_path: Path,
|
||||||
) -> None:
|
) -> None:
|
||||||
@@ -711,10 +762,43 @@ def test_policy_classifies_example_implementation_reads_separately(
|
|||||||
workspaces_root=workspace.parent,
|
workspaces_root=workspace.parent,
|
||||||
)
|
)
|
||||||
|
|
||||||
assert policy.validity.value == "clean"
|
assert policy.validity.value == "contaminated"
|
||||||
assert policy.escalated_to_product_code is True
|
assert policy.escalated_to_product_code is True
|
||||||
assert policy.reads_by_category["example_implementation"] == (str(ops_path),)
|
assert policy.reads_by_category["example_implementation"] == (str(ops_path),)
|
||||||
assert policy.disallowed_reads == ()
|
assert policy.disallowed_reads == (str(ops_path),)
|
||||||
|
|
||||||
|
|
||||||
|
def test_policy_anchors_relative_paths_to_workspace(tmp_path: Path) -> None:
|
||||||
|
from examples.agent_challenges.metrics import ToolCallEvidence
|
||||||
|
from examples.agent_challenges.policy import evaluate_policy
|
||||||
|
|
||||||
|
repository = tmp_path / "repo"
|
||||||
|
workspace = repository / "examples" / "challenge" / "workspaces" / "trial"
|
||||||
|
workspace.mkdir(parents=True)
|
||||||
|
call = ToolCallEvidence(
|
||||||
|
ordinal=1,
|
||||||
|
call_id="read-1",
|
||||||
|
tool="read",
|
||||||
|
status="completed",
|
||||||
|
title="Read workspace plan",
|
||||||
|
input={"path": "workflow.plan.json"},
|
||||||
|
metadata={},
|
||||||
|
output_chars=10,
|
||||||
|
output_preview="",
|
||||||
|
output_sha256="abc",
|
||||||
|
failed=False,
|
||||||
|
)
|
||||||
|
|
||||||
|
policy = evaluate_policy(
|
||||||
|
"none",
|
||||||
|
[call],
|
||||||
|
workspace_root=workspace,
|
||||||
|
repository_root=repository,
|
||||||
|
workspaces_root=workspace.parent,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert policy.validity.value == "clean"
|
||||||
|
assert policy.reads_by_category["workspace"] == ("workflow.plan.json",)
|
||||||
|
|
||||||
|
|
||||||
def test_policy_classifies_ready_made_example_plans_as_existing_solution(
|
def test_policy_classifies_ready_made_example_plans_as_existing_solution(
|
||||||
@@ -814,6 +898,41 @@ def test_v2_runner_default_timeout_and_workspace_cwd(tmp_path: Path) -> None:
|
|||||||
assert "repository_commit" in result
|
assert "repository_commit" in result
|
||||||
|
|
||||||
|
|
||||||
|
def test_safe_model_name_replaces_windows_path_separators() -> None:
|
||||||
|
from examples.agent_challenges.workspace import _safe_model_name
|
||||||
|
|
||||||
|
safe = _safe_model_name(r"..\bad/model:name")
|
||||||
|
|
||||||
|
assert "\\" not in safe
|
||||||
|
assert "/" not in safe
|
||||||
|
assert ":" not in safe
|
||||||
|
assert ".." not in safe
|
||||||
|
|
||||||
|
|
||||||
|
def test_instruction_bundle_rejects_path_traversal(tmp_path: Path) -> None:
|
||||||
|
import yaml
|
||||||
|
|
||||||
|
from examples.agent_challenges.workspace import _load_instruction_bundle
|
||||||
|
|
||||||
|
bundle = tmp_path / "bundle.yaml"
|
||||||
|
bundle.write_text(
|
||||||
|
yaml.safe_dump(
|
||||||
|
{
|
||||||
|
"files": [
|
||||||
|
{
|
||||||
|
"source": "../outside.md",
|
||||||
|
"destination": "wf-cli/SKILL.md",
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
|
||||||
|
with pytest.raises(ValueError, match="escapes project root"):
|
||||||
|
_load_instruction_bundle(bundle)
|
||||||
|
|
||||||
|
|
||||||
def test_v2_runner_timeout_preserves_partial_evidence(tmp_path: Path) -> None:
|
def test_v2_runner_timeout_preserves_partial_evidence(tmp_path: Path) -> None:
|
||||||
import subprocess as sp
|
import subprocess as sp
|
||||||
|
|
||||||
@@ -951,8 +1070,11 @@ def test_v2_manual_audit_includes_automatic_evidence(tmp_path: Path) -> None:
|
|||||||
result_payload,
|
result_payload,
|
||||||
official_outcome="pass",
|
official_outcome="pass",
|
||||||
auditor_notes="Reviewed and passed.",
|
auditor_notes="Reviewed and passed.",
|
||||||
|
output_name=str(tmp_path / "manual-audit.yaml"),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
assert workspace == tmp_path / "manual-audit.yaml"
|
||||||
|
assert workspace.is_file()
|
||||||
assert audit["manual_audit"]["task_outcome"] == "success"
|
assert audit["manual_audit"]["task_outcome"] == "success"
|
||||||
assert audit["manual_audit"]["evaluation_validity"] == "contaminated"
|
assert audit["manual_audit"]["evaluation_validity"] == "contaminated"
|
||||||
assert audit["manual_audit"]["official_outcome"] == "pass"
|
assert audit["manual_audit"]["official_outcome"] == "pass"
|
||||||
|
|||||||
@@ -0,0 +1,98 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from examples.agent_challenges.manifests import load_challenge_manifest
|
||||||
|
from examples.agent_challenges.models import InstructionProfile
|
||||||
|
from examples.agent_challenges.run_matrix import (
|
||||||
|
ModelProfile,
|
||||||
|
build_matrix_tasks,
|
||||||
|
parse_model_profile,
|
||||||
|
)
|
||||||
|
|
||||||
|
from .test_agent_challenge_harness_v2 import _write_manifest
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_model_profile_defaults_variant() -> None:
|
||||||
|
parsed = parse_model_profile("opencode/mimo-v2.5-free")
|
||||||
|
|
||||||
|
assert parsed == ModelProfile("opencode/mimo-v2.5-free", "high")
|
||||||
|
|
||||||
|
|
||||||
|
def test_parse_model_profile_accepts_explicit_variant() -> None:
|
||||||
|
parsed = parse_model_profile("opencode/deepseek-v4-flash-free=max")
|
||||||
|
|
||||||
|
assert parsed == ModelProfile("opencode/deepseek-v4-flash-free", "max")
|
||||||
|
|
||||||
|
|
||||||
|
def test_matrix_tasks_allocate_indices_across_profiles(tmp_path: Path) -> None:
|
||||||
|
challenge = load_challenge_manifest(_write_manifest(tmp_path / "challenge"))
|
||||||
|
(challenge.root / "results").mkdir()
|
||||||
|
existing = (
|
||||||
|
challenge.root / "results" / "opencode_deepseek-v4-flash-free-trial-002.json"
|
||||||
|
)
|
||||||
|
existing.write_text("{}", encoding="utf-8")
|
||||||
|
|
||||||
|
tasks = build_matrix_tasks(
|
||||||
|
challenges=[challenge],
|
||||||
|
profiles=[InstructionProfile.NONE, InstructionProfile.SKILLS],
|
||||||
|
models=[ModelProfile("opencode/deepseek-v4-flash-free", "max")],
|
||||||
|
trials=2,
|
||||||
|
)
|
||||||
|
|
||||||
|
assert [task.index for task in tasks] == [3, 4, 5, 6]
|
||||||
|
assert [task.profile for task in tasks] == [
|
||||||
|
InstructionProfile.NONE,
|
||||||
|
InstructionProfile.NONE,
|
||||||
|
InstructionProfile.SKILLS,
|
||||||
|
InstructionProfile.SKILLS,
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_run_trials_concurrency_invokes_all_indices(
|
||||||
|
monkeypatch,
|
||||||
|
tmp_path: Path,
|
||||||
|
capsys,
|
||||||
|
) -> None:
|
||||||
|
from examples.agent_challenges import run_trials
|
||||||
|
|
||||||
|
manifest = _write_manifest(tmp_path / "challenge")
|
||||||
|
seen: list[int] = []
|
||||||
|
|
||||||
|
def fake_run_v2_trial(*args: object, **kwargs: object) -> dict[str, object]:
|
||||||
|
index = kwargs["index"]
|
||||||
|
assert isinstance(index, int)
|
||||||
|
seen.append(index)
|
||||||
|
return {
|
||||||
|
"task_outcome": "success",
|
||||||
|
"evaluation_validity": "clean",
|
||||||
|
"duration_seconds": 1.0,
|
||||||
|
"result_path": f"trial-{index}.json",
|
||||||
|
"report_paths": {},
|
||||||
|
}
|
||||||
|
|
||||||
|
monkeypatch.setattr(run_trials, "run_v2_trial", fake_run_v2_trial)
|
||||||
|
|
||||||
|
exit_code = run_trials.main(
|
||||||
|
[
|
||||||
|
"--challenge",
|
||||||
|
str(manifest),
|
||||||
|
"--instruction-profile",
|
||||||
|
"none",
|
||||||
|
"--model",
|
||||||
|
"opencode/test",
|
||||||
|
"--trials",
|
||||||
|
"3",
|
||||||
|
"--concurrency",
|
||||||
|
"2",
|
||||||
|
"--instruction-bundle",
|
||||||
|
str(
|
||||||
|
Path(__file__).resolve().parents[2]
|
||||||
|
/ "examples/agent_challenges/instruction_bundles/workflow_cli.yaml"
|
||||||
|
),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
assert exit_code == 0
|
||||||
|
assert sorted(seen) == [1, 2, 3]
|
||||||
|
assert '"trial_count": 3' in capsys.readouterr().out
|
||||||
@@ -142,7 +142,9 @@ def test_markdown_projection_has_stable_headings(tmp_path: Path) -> None:
|
|||||||
from examples.agent_challenges.report_models import build_trial_report
|
from examples.agent_challenges.report_models import build_trial_report
|
||||||
from examples.agent_challenges.reports import render_trial_report_markdown
|
from examples.agent_challenges.reports import render_trial_report_markdown
|
||||||
|
|
||||||
report = build_trial_report(_raw_result(tmp_path), audit=None)
|
result = _raw_result(tmp_path)
|
||||||
|
result["challenge_report"] = {"run_failed": False, "notes": "ok"}
|
||||||
|
report = build_trial_report(result, audit=None)
|
||||||
md = render_trial_report_markdown(report)
|
md = render_trial_report_markdown(report)
|
||||||
|
|
||||||
expected_headings = [
|
expected_headings = [
|
||||||
@@ -163,6 +165,8 @@ def test_markdown_projection_has_stable_headings(tmp_path: Path) -> None:
|
|||||||
assert "The deployment succeeded" in md
|
assert "The deployment succeeded" in md
|
||||||
assert "large raw stream" not in md
|
assert "large raw stream" not in md
|
||||||
assert "full tool output" not in md
|
assert "full tool output" not in md
|
||||||
|
assert "```yaml\n" in md
|
||||||
|
assert "run_failed: false" in md
|
||||||
|
|
||||||
|
|
||||||
def test_markdown_command_items_indent_by_marker_width(tmp_path: Path) -> None:
|
def test_markdown_command_items_indent_by_marker_width(tmp_path: Path) -> None:
|
||||||
@@ -292,6 +296,8 @@ def test_manual_audit_regenerates_projections(tmp_path: Path) -> None:
|
|||||||
assert paths.audit.is_file()
|
assert paths.audit.is_file()
|
||||||
assert paths.markdown.is_file()
|
assert paths.markdown.is_file()
|
||||||
assert paths.machine.is_file()
|
assert paths.machine.is_file()
|
||||||
|
assert paths.results_markdown is not None
|
||||||
|
assert paths.results_markdown.is_file()
|
||||||
|
|
||||||
audit_yaml = paths.audit.read_text(encoding="utf-8")
|
audit_yaml = paths.audit.read_text(encoding="utf-8")
|
||||||
assert "official_outcome: pass" in audit_yaml
|
assert "official_outcome: pass" in audit_yaml
|
||||||
@@ -300,6 +306,7 @@ def test_manual_audit_regenerates_projections(tmp_path: Path) -> None:
|
|||||||
md = paths.markdown.read_text(encoding="utf-8")
|
md = paths.markdown.read_text(encoding="utf-8")
|
||||||
assert "Official outcome: pass" in md
|
assert "Official outcome: pass" in md
|
||||||
assert "Agent inspected a ready-made workflow plan." in md
|
assert "Agent inspected a ready-made workflow plan." in md
|
||||||
|
assert paths.results_markdown.read_text(encoding="utf-8") == md
|
||||||
|
|
||||||
machine = json.loads(paths.machine.read_text(encoding="utf-8"))
|
machine = json.loads(paths.machine.read_text(encoding="utf-8"))
|
||||||
assert machine["manual_audit"]["official_outcome"] == "pass"
|
assert machine["manual_audit"]["official_outcome"] == "pass"
|
||||||
@@ -314,9 +321,12 @@ def test_manual_audit_invalid_outcome_raises_and_preserves_projections(
|
|||||||
result_path = _write_v2_result_with_projections(tmp_path)
|
result_path = _write_v2_result_with_projections(tmp_path)
|
||||||
md_path = tmp_path / "final-report.md"
|
md_path = tmp_path / "final-report.md"
|
||||||
machine_path = tmp_path / "results" / "trial.report.json"
|
machine_path = tmp_path / "results" / "trial.report.json"
|
||||||
|
results_md_path = tmp_path / "results" / "trial.report.md"
|
||||||
|
results_md_path.write_text("existing report markdown\n", encoding="utf-8")
|
||||||
|
|
||||||
md_before = md_path.read_bytes()
|
md_before = md_path.read_bytes()
|
||||||
machine_before = machine_path.read_bytes()
|
machine_before = machine_path.read_bytes()
|
||||||
|
results_md_before = results_md_path.read_bytes()
|
||||||
|
|
||||||
with pytest.raises(ValueError, match="official_outcome"):
|
with pytest.raises(ValueError, match="official_outcome"):
|
||||||
save_v2_manual_audit(
|
save_v2_manual_audit(
|
||||||
@@ -327,6 +337,24 @@ def test_manual_audit_invalid_outcome_raises_and_preserves_projections(
|
|||||||
|
|
||||||
assert md_path.read_bytes() == md_before
|
assert md_path.read_bytes() == md_before
|
||||||
assert machine_path.read_bytes() == machine_before
|
assert machine_path.read_bytes() == machine_before
|
||||||
|
assert results_md_path.read_bytes() == results_md_before
|
||||||
|
|
||||||
|
|
||||||
|
def test_trial_report_bounds_agent_self_report_payload(tmp_path: Path) -> None:
|
||||||
|
from examples.agent_challenges.report_models import build_trial_report
|
||||||
|
|
||||||
|
result = _raw_result(tmp_path)
|
||||||
|
result["challenge_report"] = {
|
||||||
|
"run_failed": False,
|
||||||
|
"notes": "x" * 10_000,
|
||||||
|
"nested": {"items": list(range(100))},
|
||||||
|
}
|
||||||
|
|
||||||
|
report = build_trial_report(result, audit=None)
|
||||||
|
|
||||||
|
assert report.agent_self_report is not None
|
||||||
|
assert len(report.agent_self_report["notes"]) == 2_000
|
||||||
|
assert len(report.agent_self_report["nested"]["items"]) == 50
|
||||||
|
|
||||||
|
|
||||||
def test_discrepancy_detects_run_failed_contradiction(tmp_path: Path) -> None:
|
def test_discrepancy_detects_run_failed_contradiction(tmp_path: Path) -> None:
|
||||||
|
|||||||
@@ -1,5 +1,6 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
|
|
||||||
from examples.agent_challenges.manifests import load_challenge_manifest
|
from examples.agent_challenges.manifests import load_challenge_manifest
|
||||||
@@ -61,7 +62,12 @@ def test_report_challenge_workspace_template_contains_safe_input_files(
|
|||||||
.read_text(encoding="utf-8")
|
.read_text(encoding="utf-8")
|
||||||
.startswith("# Weekly Project Update")
|
.startswith("# Weekly Project Update")
|
||||||
)
|
)
|
||||||
assert '"text"' in (workspace.root / "run-input.json").read_text(encoding="utf-8")
|
run_input = json.loads(
|
||||||
|
(workspace.root / "run-input.json").read_text(encoding="utf-8")
|
||||||
|
)
|
||||||
|
assert isinstance(run_input, dict)
|
||||||
|
assert isinstance(run_input.get("text"), str)
|
||||||
|
assert "path" not in run_input
|
||||||
assert "not workflow solutions" in (workspace.root / "TASK_FILES.md").read_text(
|
assert "not workflow solutions" in (workspace.root / "TASK_FILES.md").read_text(
|
||||||
encoding="utf-8"
|
encoding="utf-8"
|
||||||
)
|
)
|
||||||
|
|||||||
@@ -80,7 +80,7 @@ def test_schema_compact_component_is_queryable() -> None:
|
|||||||
payload = _json_result("NodeUse")
|
payload = _json_result("NodeUse")
|
||||||
|
|
||||||
assert payload["name"] == "NodeUse"
|
assert payload["name"] == "NodeUse"
|
||||||
assert payload["properties"]["input"]["items"]["one_of"] == [
|
assert payload["properties"]["input"]["items"]["any_of"] == [
|
||||||
"InputPathBinding",
|
"InputPathBinding",
|
||||||
"InputValueBinding",
|
"InputValueBinding",
|
||||||
]
|
]
|
||||||
@@ -129,13 +129,23 @@ def test_schema_catalog_resolves_aliases_and_components() -> None:
|
|||||||
def test_compact_outline_replaces_local_refs_with_names() -> None:
|
def test_compact_outline_replaces_local_refs_with_names() -> None:
|
||||||
payload = compact_schema_outline("NodeUse")
|
payload = compact_schema_outline("NodeUse")
|
||||||
|
|
||||||
assert payload["properties"]["input"]["items"]["one_of"] == [
|
assert payload["properties"]["input"]["items"]["any_of"] == [
|
||||||
"InputPathBinding",
|
"InputPathBinding",
|
||||||
"InputValueBinding",
|
"InputValueBinding",
|
||||||
]
|
]
|
||||||
assert "$ref" not in json.dumps(payload)
|
assert "$ref" not in json.dumps(payload)
|
||||||
|
|
||||||
|
|
||||||
|
def test_compact_outline_preserves_any_of_keyword() -> None:
|
||||||
|
payload = compact_schema_outline("NodeUse")
|
||||||
|
|
||||||
|
assert payload["properties"]["input"]["items"]["any_of"] == [
|
||||||
|
"InputPathBinding",
|
||||||
|
"InputValueBinding",
|
||||||
|
]
|
||||||
|
assert "one_of" not in payload["properties"]["input"]["items"]
|
||||||
|
|
||||||
|
|
||||||
def test_verbose_component_uses_generated_definitions() -> None:
|
def test_verbose_component_uses_generated_definitions() -> None:
|
||||||
payload = verbose_schema_document("NodeUse")
|
payload = verbose_schema_document("NodeUse")
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user