fix: harden workflow demo and evaluation tooling
This commit is contained in:
+4
-2
@@ -10,8 +10,10 @@ implementation plans are kept for context, not as active instructions.
|
||||
and verification commands.
|
||||
- [`add/2026-06-workflow-platform-presentation.md`](add/2026-06-workflow-platform-presentation.md):
|
||||
concise presentation narrative for the current product shape and demo flow.
|
||||
- [`thesis/thesis-outline.md`](thesis/thesis-outline.md): thesis/report scaffold,
|
||||
argument structure, evidence checklist, limitations, and future work.
|
||||
- [`thesis/system-design-implementation.md`](thesis/system-design-implementation.md):
|
||||
maintained thesis/report document.
|
||||
- [`thesis/evidence-index.md`](thesis/evidence-index.md): claim-to-code and
|
||||
claim-to-test evidence map for the thesis.
|
||||
- [`current_roadmap.md`](current_roadmap.md): active next-work list after the
|
||||
core type-shape cleanup.
|
||||
- [`wf_core_architecture.md`](wf_core_architecture.md): kernel package
|
||||
|
||||
@@ -134,7 +134,8 @@ clear operator feedback before adding more architecture.
|
||||
`ux_issues_found: []` so debug-profile reports do not fail by omission.
|
||||
- Completed: `wf draft bind` now reuses existing workflow input/state schema
|
||||
fields when binding to step-local inputs, avoiding redundant-schema failures
|
||||
found by debug challenge runs.
|
||||
found by debug challenge runs. Implementation:
|
||||
[`idempotent draft bind inputs`](historical/superpowers/plans/2026-06-29-idempotent-draft-bind-inputs.md).
|
||||
- Keep status read-only; do not mutate registry, auth, config, or stores.
|
||||
|
||||
## Priority 2: Durable Run/Resume Hardening
|
||||
|
||||
@@ -2,8 +2,7 @@
|
||||
|
||||
Date: 2026-07-01
|
||||
|
||||
Status: Approved direction. Needs an executable implementation plan before code
|
||||
changes.
|
||||
Status: Implemented. This document is the live interrupt contract.
|
||||
|
||||
Related:
|
||||
|
||||
@@ -27,23 +26,18 @@ This is required for the Workflow Console and useful for CLI, JSON-RPC, MCP, and
|
||||
agent clients. It prevents every client from needing workflow-specific code just
|
||||
to render and answer an approval step.
|
||||
|
||||
## Current Gap
|
||||
## Implemented Contract
|
||||
|
||||
The current core interrupt model has useful mechanics but not a complete public
|
||||
contract.
|
||||
The core interrupt model now exposes a complete public pause/resume contract.
|
||||
|
||||
- `InterruptNode` stores `kind`, request bindings, resume bindings, and resume
|
||||
outcomes.
|
||||
outcomes, plus `request_schema` and `resume_schema`.
|
||||
- Runtime builds an `InterruptRequest` with id, frame id, node id, kind,
|
||||
payload, route, and resumability.
|
||||
- `workflow.runs.inspect` serializes the current runtime interrupt request.
|
||||
- Resume validates the selected outcome against `InterruptNode.outcomes` and
|
||||
applies resume bindings.
|
||||
|
||||
What is missing is a machine-readable schema for the request payload and resume
|
||||
payload. A client can see data, but it cannot know whether the response should
|
||||
be `{ "approved": true }`, `{ "selected_issue_ids": [...] }`, or something else
|
||||
without reading workflow code or challenge-specific docs.
|
||||
payload, route, resumability, outcomes, request schema, resume schema, and a
|
||||
`typed` marker.
|
||||
- `workflow.runs.inspect` serializes that persisted interrupt request.
|
||||
- Runtime validates request payloads before pausing and validates resume
|
||||
payloads against the persisted pause-time schema before mutating state.
|
||||
|
||||
This is close to the LangGraph-style `interrupt(value)` and `Command(resume=...)`
|
||||
pattern: flexible and simple, but the response contract is mostly app
|
||||
@@ -51,7 +45,7 @@ convention. `wf` should keep the flexibility while making the contract explicit.
|
||||
|
||||
## Design Summary
|
||||
|
||||
Add JSON Schema contracts to interrupt nodes and carry them through persisted run
|
||||
JSON Schema contracts are carried from interrupt nodes through persisted run
|
||||
inspection and resume validation.
|
||||
|
||||
```mermaid
|
||||
@@ -73,7 +67,7 @@ sequenceDiagram
|
||||
|
||||
## Core Model
|
||||
|
||||
Extend `InterruptNode` with two optional schema fields:
|
||||
`InterruptNode` carries two schema fields:
|
||||
|
||||
```python
|
||||
class InterruptNode(BaseModel):
|
||||
|
||||
@@ -347,12 +347,12 @@ The UI selects all issues by default and allows individual deselection.
|
||||
|
||||
## Self-Describing Interrupt Dependency
|
||||
|
||||
The current interrupt payload exposes `kind` and data but not a complete
|
||||
machine-readable response contract. The Workflow Console requires explicit
|
||||
request and resume schemas so it can render and validate arbitrary interrupts
|
||||
without reading workflow code.
|
||||
Interrupted run inspection now exposes `kind`, payload, outcomes,
|
||||
`request_schema`, `resume_schema`, and `typed`. Runtime validates request and
|
||||
resume payloads against those contracts. The Workflow Console can therefore
|
||||
render and validate arbitrary interrupts without reading workflow code.
|
||||
|
||||
The prerequisite contract is defined in
|
||||
The implemented contract is defined in
|
||||
[Self-describing interrupt contracts](2026-07-01-self-describing-interrupt-contracts.md).
|
||||
|
||||
## Replay And Failure Handling
|
||||
|
||||
@@ -35,6 +35,29 @@ if (-not (Test-Path $pandoc_diagram)) {
|
||||
Write-Error "pandoc diagram wrapper not found at $pandoc_diagram. Make sure it exists, then rerun this command."
|
||||
exit 1
|
||||
}
|
||||
|
||||
function Test-RenderNeedsAgentResults([string[]] $arguments) {
|
||||
for ($index = 0; $index -lt $arguments.Count; $index++) {
|
||||
$candidate = $null
|
||||
if ($arguments[$index] -in @("-i", "--input")) {
|
||||
if ($index + 1 -lt $arguments.Count) {
|
||||
$candidate = $arguments[$index + 1]
|
||||
$index++
|
||||
}
|
||||
}
|
||||
elseif ($arguments[$index].EndsWith(".md")) {
|
||||
$candidate = $arguments[$index]
|
||||
}
|
||||
if (
|
||||
$candidate -and
|
||||
(Test-Path -LiteralPath $candidate) -and
|
||||
(Select-String -LiteralPath $candidate -SimpleMatch "include-agent-challenge-results" -Quiet)
|
||||
) {
|
||||
return $true
|
||||
}
|
||||
}
|
||||
return $false
|
||||
}
|
||||
. $pandoc_diagram
|
||||
|
||||
$pandoc_crossref = Get-Command pandoc-crossref -ErrorAction SilentlyContinue
|
||||
@@ -62,7 +85,7 @@ if (-not (Test-Path $include_markdown_filter)) {
|
||||
}
|
||||
|
||||
$agent_results = Join-Path $PSScriptRoot "agent-challenge-results.md"
|
||||
if (-not (Test-Path $agent_results)) {
|
||||
if ((Test-RenderNeedsAgentResults $RemainingArgs) -and -not (Test-Path $agent_results)) {
|
||||
Write-Error "agent-challenge-results.md is missing. Run generate_agent_challenge_evaluation.py first."
|
||||
exit 1
|
||||
}
|
||||
|
||||
@@ -54,10 +54,17 @@ def _parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
args = _parse_args(argv)
|
||||
for path in generate(
|
||||
manifest_path=args.manifest.resolve(), output_dir=args.output_dir.resolve()
|
||||
):
|
||||
print(path.relative_to(ROOT))
|
||||
output_dir = args.output_dir.resolve()
|
||||
for path in generate(manifest_path=args.manifest.resolve(), output_dir=output_dir):
|
||||
resolved = path.resolve()
|
||||
try:
|
||||
display_path = resolved.relative_to(ROOT)
|
||||
except ValueError:
|
||||
try:
|
||||
display_path = resolved.relative_to(output_dir)
|
||||
except ValueError:
|
||||
display_path = resolved
|
||||
print(display_path)
|
||||
return 0
|
||||
|
||||
|
||||
|
||||
@@ -16,10 +16,13 @@ if (-not (Test-Path $file)) {
|
||||
$file = (Resolve-Path $file).Path
|
||||
$resourcePath = [System.IO.Path]::GetDirectoryName($file)
|
||||
|
||||
$evaluationGenerator = Join-Path $PSScriptRoot "generate_agent_challenge_evaluation.py"
|
||||
& uv run python $evaluationGenerator
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
throw "agent challenge evaluation generation failed with exit code $LASTEXITCODE"
|
||||
$needsAgentResults = Select-String -LiteralPath $file -SimpleMatch "include-agent-challenge-results" -Quiet
|
||||
if ($needsAgentResults) {
|
||||
$evaluationGenerator = Join-Path $PSScriptRoot "generate_agent_challenge_evaluation.py"
|
||||
& uv run python $evaluationGenerator
|
||||
if ($LASTEXITCODE -ne 0) {
|
||||
throw "agent challenge evaluation generation failed with exit code $LASTEXITCODE"
|
||||
}
|
||||
}
|
||||
|
||||
# name without extension
|
||||
|
||||
@@ -1994,7 +1994,7 @@ product.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `docs/thesis/thesis-outline.md`
|
||||
- `docs/historical/thesis/thesis-outline.md`
|
||||
- `docs/current_roadmap.md`
|
||||
- Absence of scheduler, visual-editor, and secret-manager production packages
|
||||
in the current source tree.
|
||||
|
||||
@@ -17,6 +17,8 @@ FIGURE_STEMS = (
|
||||
"agent-challenge-token-volume",
|
||||
)
|
||||
_MANUAL_OUTCOMES = frozenset({"pass", "invalid", "fail"})
|
||||
_TASK_OUTCOMES = frozenset({"success", "failed", "timeout", "runner_error"})
|
||||
_EVALUATION_PROFILES = frozenset({"none", "skills", "all"})
|
||||
_CHALLENGE_ORDER = {"browser": 0, "report": 1}
|
||||
_MODEL_ORDER = {"deepseek": 0, "mimo": 1}
|
||||
_PROFILE_ORDER = {"none": 0, "skills": 1, "all": 2}
|
||||
@@ -105,6 +107,19 @@ def _load_trial(
|
||||
raise ValueError(
|
||||
f"unsupported manual outcome {manual_outcome!r}: {report_path}"
|
||||
)
|
||||
profile = _string(run.get("profile"), field="runs[].profile")
|
||||
if profile not in _EVALUATION_PROFILES:
|
||||
raise ValueError(f"unsupported evaluation profile {profile!r}: {report_path}")
|
||||
task_outcome = _string(run.get("task_outcome"), field="runs[].task_outcome")
|
||||
if task_outcome not in _TASK_OUTCOMES:
|
||||
raise ValueError(f"unsupported task outcome {task_outcome!r}: {report_path}")
|
||||
raw_audit_notes = run.get("audit_notes")
|
||||
if raw_audit_notes is None:
|
||||
audit_notes = ""
|
||||
elif isinstance(raw_audit_notes, str):
|
||||
audit_notes = raw_audit_notes
|
||||
else:
|
||||
raise ValueError(f"runs[].audit_notes must be a string: {report_path}")
|
||||
|
||||
duration = run.get("duration_seconds")
|
||||
if not isinstance(duration, int | float):
|
||||
@@ -117,7 +132,7 @@ def _load_trial(
|
||||
_string(run.get("challenge"), field="runs[].challenge")
|
||||
),
|
||||
model=short_model_name(_string(run.get("model"), field="runs[].model")),
|
||||
profile=_string(run.get("profile"), field="runs[].profile"),
|
||||
profile=profile,
|
||||
trial_index=_integer(run.get("trial_index"), field="runs[].trial_index"),
|
||||
repository_commit=_string(
|
||||
run.get("repository_commit"), field="runs[].repository_commit"
|
||||
@@ -126,10 +141,10 @@ def _load_trial(
|
||||
run.get("base_prompt_hash"), field="runs[].base_prompt_hash"
|
||||
),
|
||||
manual_outcome=manual_outcome,
|
||||
task_outcome=_string(run.get("task_outcome"), field="runs[].task_outcome"),
|
||||
task_outcome=task_outcome,
|
||||
duration_seconds=float(duration),
|
||||
tokens_total=_integer(run.get("tokens_total"), field="runs[].tokens_total"),
|
||||
audit_notes=str(run.get("audit_notes") or ""),
|
||||
audit_notes=audit_notes,
|
||||
)
|
||||
|
||||
|
||||
|
||||
@@ -24,6 +24,7 @@ _MODEL_STYLES = {
|
||||
_PROFILE_ORDER = {"none": 0, "skills": 1, "all": 2}
|
||||
_CHALLENGE_ORDER = {"browser": 0, "report": 1}
|
||||
_MODEL_ORDER = {"deepseek": 0, "mimo": 1}
|
||||
_TASK_OUTCOME_ORDER = ("success", "failed", "timeout", "runner_error")
|
||||
|
||||
|
||||
def _configure_matplotlib() -> Any:
|
||||
@@ -170,8 +171,17 @@ def _outcomes_by_cell(cohort: EvaluationCohort, plt: Any) -> Figure:
|
||||
def _automatic_vs_manual(cohort: EvaluationCohort, plt: Any) -> Figure:
|
||||
from matplotlib.colors import LinearSegmentedColormap
|
||||
|
||||
task_labels = ("success", "failed")
|
||||
task_labels = tuple(
|
||||
outcome
|
||||
for outcome in _TASK_OUTCOME_ORDER
|
||||
if any(trial.task_outcome == outcome for trial in cohort.trials)
|
||||
)
|
||||
manual_labels = ("pass", "invalid", "fail")
|
||||
unknown_task_outcomes = sorted(
|
||||
{trial.task_outcome for trial in cohort.trials} - set(_TASK_OUTCOME_ORDER)
|
||||
)
|
||||
if unknown_task_outcomes:
|
||||
raise ValueError(f"unsupported task outcomes: {unknown_task_outcomes}")
|
||||
matrix = [
|
||||
[
|
||||
sum(
|
||||
@@ -211,7 +221,7 @@ def _automatic_vs_manual(cohort: EvaluationCohort, plt: Any) -> Figure:
|
||||
axis.set_ylabel("Automatic task outcome")
|
||||
axis.set_title("Automatic completion does not imply clean evaluation evidence")
|
||||
axis.set_xticks([value - 0.5 for value in range(1, len(manual_labels))], minor=True)
|
||||
axis.set_yticks([0.5], minor=True)
|
||||
axis.set_yticks([value - 0.5 for value in range(1, len(task_labels))], minor=True)
|
||||
axis.grid(which="minor", color="white", linewidth=2)
|
||||
axis.tick_params(which="minor", bottom=False, left=False)
|
||||
figure.colorbar(image, ax=axis, label="Trial count", shrink=0.82)
|
||||
@@ -275,7 +285,11 @@ def _scatter_metric(
|
||||
profiles = ("none", "skills", "all")
|
||||
wave_offsets = {1: -0.055, 2: 0.0, 3: 0.055}
|
||||
for trial in sorted(trials, key=_trial_sort_key):
|
||||
style = _MODEL_STYLES[trial.model]
|
||||
style = _MODEL_STYLES.get(trial.model)
|
||||
if style is None:
|
||||
raise ValueError(f"unsupported evaluation model {trial.model!r}")
|
||||
if trial.profile not in profiles:
|
||||
raise ValueError(f"unsupported evaluation profile {trial.profile!r}")
|
||||
x = (
|
||||
profiles.index(trial.profile)
|
||||
+ float(style["offset"])
|
||||
|
||||
@@ -97,20 +97,8 @@ def resume_command_from_result(
|
||||
variant: str | None = None,
|
||||
prompt_mode: PromptMode = "auto",
|
||||
) -> list[str]:
|
||||
"""Build a resume command from new metadata or recover it from old raw results."""
|
||||
"""Rebuild a resume command from validated metadata or old raw results."""
|
||||
opencode = result.get("opencode")
|
||||
if isinstance(opencode, dict):
|
||||
command = opencode.get("resume_command")
|
||||
has_override = (
|
||||
any(value is not None for value in (session_id, attach_url, model, variant))
|
||||
or prompt_mode != "auto"
|
||||
)
|
||||
if (
|
||||
not has_override
|
||||
and isinstance(command, list)
|
||||
and all(isinstance(part, str) for part in command)
|
||||
):
|
||||
return command
|
||||
|
||||
stdout = result.get("stdout")
|
||||
stdout_text = stdout if isinstance(stdout, str) else ""
|
||||
|
||||
@@ -249,13 +249,10 @@ def render_trial_report_markdown(report: TrialReport) -> str:
|
||||
)
|
||||
lines.append("")
|
||||
|
||||
if report.opencode is not None:
|
||||
if report.opencode is not None and report.opencode.session_id:
|
||||
lines.append("## OpenCode Resume")
|
||||
lines.append("")
|
||||
if report.opencode.session_id:
|
||||
lines.append(f"- Session: `{report.opencode.session_id}`")
|
||||
else:
|
||||
lines.append("- Session: not captured")
|
||||
lines.append(f"- Session: `{report.opencode.session_id}`")
|
||||
if report.opencode.attach_url:
|
||||
lines.append(f"- Attach URL: `{report.opencode.attach_url}`")
|
||||
if report.opencode.resume_command:
|
||||
|
||||
@@ -154,7 +154,7 @@ def main(argv: list[str] | None = None) -> int:
|
||||
)
|
||||
else:
|
||||
print(_display_command(command))
|
||||
except ValueError as exc:
|
||||
except (OSError, ValueError) as exc:
|
||||
parser.error(str(exc))
|
||||
return 0
|
||||
|
||||
|
||||
@@ -6,6 +6,7 @@ param(
|
||||
[object[]]$models = @( # object[] because ModelProfile is not known yet
|
||||
[ModelProfile]::new("opencode/deepseek-v4-flash-free", "max"),
|
||||
[ModelProfile]::new("opencode/mimo-v2.5-free", "high")
|
||||
# Disabled after repeated provider timeouts and resource-exhausted failures.
|
||||
# [ModelProfile]::new("opencode/nemotron-3-ultra-free", "high")
|
||||
)
|
||||
)
|
||||
|
||||
@@ -45,6 +45,7 @@ class ModelProfile:
|
||||
DEFAULT_MODELS = (
|
||||
ModelProfile("opencode/deepseek-v4-flash-free", "max"),
|
||||
ModelProfile("opencode/mimo-v2.5-free", "high"),
|
||||
# Disabled after repeated provider timeouts and resource-exhausted failures.
|
||||
# ModelProfile("opencode/nemotron-3-ultra-free", "high"),
|
||||
)
|
||||
|
||||
|
||||
@@ -42,8 +42,11 @@ uv run wf --config examples/lda_report_workflow/wf.config.json --local run inspe
|
||||
|
||||
Resume with selected issues:
|
||||
|
||||
Copy one or more actual issue ids from the interrupt payload and replace the
|
||||
placeholder below.
|
||||
|
||||
```powershell
|
||||
uv run wf --config examples/lda_report_workflow/wf.config.json --local run resume <run_id> --payload '{"approved":true,"selected_issue_ids":["risk-1"],"comment":"Create selected issues."}'
|
||||
uv run wf --config examples/lda_report_workflow/wf.config.json --local run resume <run_id> --payload '{"approved":true,"selected_issue_ids":["<issue_id_from_interrupt>"],"comment":"Create selected issues."}'
|
||||
```
|
||||
|
||||
## Cleanup
|
||||
|
||||
@@ -71,6 +71,11 @@ def build_workflow() -> Workflow:
|
||||
input=[{"path": "input.selected_documents", "target": "names"}],
|
||||
output=[{"source": "documents", "target": "state.documents"}],
|
||||
)
|
||||
reset_board = builder.use_ref(
|
||||
"local.issue_board.reset_issue_board",
|
||||
id="reset_board",
|
||||
input=[{"path": "input.board_path", "target": "board_path"}],
|
||||
)
|
||||
analyze = builder.use_ref(
|
||||
"local.lda_report.analyze_documents",
|
||||
id="analyze",
|
||||
@@ -164,7 +169,8 @@ def build_workflow() -> Workflow:
|
||||
end_completed = builder.end("completed", id="end_completed")
|
||||
end_cancelled = builder.end("cancelled", id="end_cancelled")
|
||||
|
||||
builder.set_entry_point(read_docs)
|
||||
builder.set_entry_point(reset_board)
|
||||
builder.connect(reset_board, "ok", read_docs)
|
||||
builder.connect(read_docs, "ok", analyze)
|
||||
builder.connect(analyze, "ok", build_report)
|
||||
builder.connect(build_report, "ok", draft_issues)
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
{
|
||||
"edges": [
|
||||
{
|
||||
"from": "reset_board",
|
||||
"outcome": "ok",
|
||||
"to": "read_docs"
|
||||
},
|
||||
{
|
||||
"from": "read_docs",
|
||||
"outcome": "ok",
|
||||
@@ -86,6 +91,21 @@
|
||||
"timeout_seconds": null,
|
||||
"type": "node"
|
||||
},
|
||||
{
|
||||
"desc": null,
|
||||
"id": "reset_board",
|
||||
"input": [
|
||||
{
|
||||
"path": "input.board_path",
|
||||
"target": "board_path"
|
||||
}
|
||||
],
|
||||
"node": "local.issue_board.reset_issue_board",
|
||||
"output": [],
|
||||
"retry": null,
|
||||
"timeout_seconds": null,
|
||||
"type": "node"
|
||||
},
|
||||
{
|
||||
"desc": null,
|
||||
"id": "analyze",
|
||||
@@ -368,7 +388,7 @@
|
||||
"required": [],
|
||||
"type": "object"
|
||||
},
|
||||
"start": "read_docs",
|
||||
"start": "reset_board",
|
||||
"state_schema": {
|
||||
"properties": {
|
||||
"analysis": {
|
||||
|
||||
@@ -115,7 +115,7 @@ def resume_interrupt(
|
||||
)
|
||||
|
||||
validate_payload_against_schema(
|
||||
step.resume_schema,
|
||||
run.interrupt.resume_schema,
|
||||
resume_payload,
|
||||
f"interrupt resume for {step.id}",
|
||||
)
|
||||
|
||||
@@ -130,7 +130,7 @@ async def test_interrupt_request_payload_validates_against_schema() -> None:
|
||||
input_schema=_schema(),
|
||||
state_schema=StateSchema.from_field_map({}),
|
||||
output_schema=_schema(),
|
||||
outcomes=["ok"],
|
||||
outcomes=["submitted"],
|
||||
start="ask",
|
||||
nodes=[
|
||||
InterruptNode(
|
||||
@@ -147,8 +147,11 @@ async def test_interrupt_request_payload_validates_against_schema() -> None:
|
||||
"additionalProperties": False,
|
||||
},
|
||||
),
|
||||
EndNode(id="end", type="end", outcome="submitted"),
|
||||
],
|
||||
edges=[
|
||||
Edge.model_validate({"from": "ask", "outcome": "submitted", "to": "end"})
|
||||
],
|
||||
edges=[Edge.model_validate({"from": "ask", "outcome": "submitted", "to": END})],
|
||||
)
|
||||
|
||||
run = await execute_workflow_result_async(workflow, {}, {})
|
||||
@@ -159,6 +162,55 @@ async def test_interrupt_request_payload_validates_against_schema() -> None:
|
||||
assert "interrupt request for ask" in run.error
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_interrupt_resume_uses_persisted_pause_schema() -> None:
|
||||
workflow = Workflow(
|
||||
name="persisted_resume_contract",
|
||||
input_schema=_schema(),
|
||||
state_schema=StateSchema.from_field_map({}),
|
||||
output_schema=_schema(),
|
||||
outcomes=["submitted"],
|
||||
start="ask",
|
||||
nodes=[
|
||||
InterruptNode(
|
||||
id="ask",
|
||||
type="interrupt",
|
||||
kind="approval",
|
||||
resume_schema={
|
||||
"type": "object",
|
||||
"properties": {"approved": {"type": "boolean"}},
|
||||
"required": ["approved"],
|
||||
"additionalProperties": False,
|
||||
},
|
||||
),
|
||||
EndNode(id="end", type="end", outcome="submitted"),
|
||||
],
|
||||
edges=[
|
||||
Edge.model_validate({"from": "ask", "outcome": "submitted", "to": "end"})
|
||||
],
|
||||
)
|
||||
interrupted = await execute_workflow_async(workflow, {}, {})
|
||||
changed_workflow = workflow.model_copy(deep=True)
|
||||
changed_step = changed_workflow.nodes[0]
|
||||
assert isinstance(changed_step, InterruptNode)
|
||||
changed_step.resume_schema = {
|
||||
"type": "object",
|
||||
"properties": {"approved": {"type": "string"}},
|
||||
"required": ["approved"],
|
||||
"additionalProperties": False,
|
||||
}
|
||||
|
||||
resumed = await resume_workflow_result_async(
|
||||
changed_workflow,
|
||||
interrupted,
|
||||
{},
|
||||
resume_payload={"approved": True},
|
||||
)
|
||||
|
||||
assert resumed.status == RunStatus.COMPLETED
|
||||
assert resumed.outcome == "submitted"
|
||||
|
||||
|
||||
async def test_interrupt_resume_payload_validates_before_state_mutation() -> None:
|
||||
workflow = Workflow(
|
||||
name="resume_validation",
|
||||
|
||||
@@ -38,6 +38,18 @@ def test_project_map_links_big_doc() -> None:
|
||||
assert any(link.endswith("evidence-index.md") for link in links)
|
||||
|
||||
|
||||
def test_docs_index_points_to_live_thesis_and_archives_scaffolds() -> None:
|
||||
docs = ROOT / "docs"
|
||||
index = (docs / "README.md").read_text(encoding="utf-8")
|
||||
|
||||
assert "thesis/system-design-implementation.md" in index
|
||||
assert "thesis/evidence-index.md" in index
|
||||
assert not (docs / "thesis" / "thesis-outline.md").exists()
|
||||
assert not (docs / "thesis" / "diagrams.md").exists()
|
||||
assert (docs / "historical" / "thesis" / "thesis-outline.md").is_file()
|
||||
assert (docs / "historical" / "thesis" / "diagrams.md").is_file()
|
||||
|
||||
|
||||
def test_big_doc_keeps_mcp_as_source_family() -> None:
|
||||
doc = (ROOT / "docs" / "thesis" / "system-design-implementation.md").read_text(
|
||||
encoding="utf-8"
|
||||
@@ -92,6 +104,9 @@ def test_thesis_bundle_has_reproducible_agent_evaluation_assets() -> None:
|
||||
< generate_script.index("--filter=pandoc-crossref")
|
||||
)
|
||||
assert "generate_agent_challenge_evaluation.py" in combined_build_script
|
||||
assert "Test-RenderNeedsAgentResults" in generate_script
|
||||
assert "include-agent-challenge-results" in combined_build_script
|
||||
assert "if ($needsAgentResults)" in combined_build_script
|
||||
assert "--resource-path" in combined_build_script
|
||||
for stem in figure_stems:
|
||||
assert f"figures/{stem}.svg" in results
|
||||
|
||||
@@ -2,8 +2,11 @@ from __future__ import annotations
|
||||
|
||||
import json
|
||||
from collections import Counter
|
||||
from dataclasses import replace
|
||||
from pathlib import Path
|
||||
|
||||
import pytest
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
COHORT_PATH = ROOT / "docs" / "thesis" / "agent-challenge-cohort.json"
|
||||
|
||||
@@ -60,6 +63,80 @@ def test_primary_cohort_snapshot_loads_without_local_report_files(
|
||||
assert len(cohort.trials) == 36
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("field", "value", "message"),
|
||||
[
|
||||
("profile", "debug", "unsupported evaluation profile"),
|
||||
("task_outcome", "unknown", "unsupported task outcome"),
|
||||
("audit_notes", ["not", "text"], "audit_notes must be a string"),
|
||||
],
|
||||
)
|
||||
def test_evaluation_cohort_rejects_unknown_or_mistyped_run_values(
|
||||
tmp_path: Path, field: str, value: object, message: str
|
||||
) -> None:
|
||||
from examples.agent_challenges.evaluation import load_evaluation_cohort
|
||||
|
||||
manifest = json.loads(COHORT_PATH.read_text(encoding="utf-8"))
|
||||
manifest["runs"][0][field] = value
|
||||
snapshot = tmp_path / "agent-challenge-cohort.json"
|
||||
snapshot.write_text(json.dumps(manifest), encoding="utf-8")
|
||||
|
||||
with pytest.raises(ValueError, match=message):
|
||||
load_evaluation_cohort(snapshot, repository_root=tmp_path)
|
||||
|
||||
|
||||
def test_evaluation_figures_reject_unknown_direct_trial_values() -> None:
|
||||
import matplotlib.pyplot as plt
|
||||
|
||||
from examples.agent_challenges.evaluation import load_evaluation_cohort
|
||||
from examples.agent_challenges.evaluation_figures import (
|
||||
_automatic_vs_manual,
|
||||
_scatter_metric,
|
||||
)
|
||||
|
||||
cohort = load_evaluation_cohort(COHORT_PATH, repository_root=ROOT)
|
||||
figure, axis = plt.subplots()
|
||||
try:
|
||||
with pytest.raises(ValueError, match="unsupported evaluation model"):
|
||||
_scatter_metric(
|
||||
axis,
|
||||
[replace(cohort.trials[0], model="unknown-model")],
|
||||
metric="duration",
|
||||
)
|
||||
with pytest.raises(ValueError, match="unsupported evaluation profile"):
|
||||
_scatter_metric(
|
||||
axis,
|
||||
[replace(cohort.trials[0], profile="unknown-profile")],
|
||||
metric="duration",
|
||||
)
|
||||
invalid_cohort = replace(
|
||||
cohort,
|
||||
trials=(replace(cohort.trials[0], task_outcome="unknown"),),
|
||||
)
|
||||
with pytest.raises(ValueError, match="unsupported task outcomes"):
|
||||
_automatic_vs_manual(invalid_cohort, plt)
|
||||
finally:
|
||||
plt.close(figure)
|
||||
|
||||
|
||||
def test_evaluation_generator_prints_paths_outside_repository(
|
||||
tmp_path: Path,
|
||||
monkeypatch: pytest.MonkeyPatch,
|
||||
capsys: pytest.CaptureFixture[str],
|
||||
) -> None:
|
||||
from docs.thesis import generate_agent_challenge_evaluation as generator
|
||||
|
||||
generated = tmp_path / "agent-challenge-results.md"
|
||||
monkeypatch.setattr(
|
||||
generator,
|
||||
"generate",
|
||||
lambda **_kwargs: (generated,),
|
||||
)
|
||||
|
||||
assert generator.main(["--output-dir", str(tmp_path)]) == 0
|
||||
assert capsys.readouterr().out.strip() == "agent-challenge-results.md"
|
||||
|
||||
|
||||
def test_evaluation_renderer_writes_stable_svg_and_pdf_names(tmp_path: Path) -> None:
|
||||
from examples.agent_challenges.evaluation import (
|
||||
FIGURE_STEMS,
|
||||
|
||||
@@ -129,6 +129,13 @@ def test_challenge_prompt_is_identical_across_profiles(tmp_path: Path) -> None:
|
||||
assert "genuinely blocked" in rendered[InstructionProfile.DEBUG].text
|
||||
assert "debug profile only" in rendered[InstructionProfile.DEBUG].text
|
||||
assert "ux_issues_found" in rendered[InstructionProfile.DEBUG].text
|
||||
for profile in (
|
||||
InstructionProfile.NONE,
|
||||
InstructionProfile.SKILLS,
|
||||
InstructionProfile.ALL,
|
||||
):
|
||||
assert "debug profile only" not in rendered[profile].text
|
||||
assert "ux_issues_found" not in rendered[profile].text
|
||||
none_prompt = rendered[InstructionProfile.NONE].text.replace("\n", " ")
|
||||
assert "inline in your" in none_prompt
|
||||
assert "without an inline self-report is invalid" in none_prompt
|
||||
|
||||
@@ -616,3 +616,20 @@ def test_trial_report_renders_opencode_resume_metadata(tmp_path: Path) -> None:
|
||||
assert "## OpenCode Resume" in rendered
|
||||
assert "ses_report" in rendered
|
||||
assert "opencode run --session ses_report" in rendered
|
||||
|
||||
|
||||
def test_trial_report_skips_resume_section_without_session_id(tmp_path: Path) -> None:
|
||||
from examples.agent_challenges.report_models import build_trial_report
|
||||
from examples.agent_challenges.reports import render_trial_report_markdown
|
||||
|
||||
result = _raw_result(tmp_path)
|
||||
result["opencode"] = {
|
||||
"model": "opencode/deepseek-v4-flash-free",
|
||||
"variant": "max",
|
||||
"session_id": None,
|
||||
}
|
||||
|
||||
report = build_trial_report(result, audit=None)
|
||||
rendered = render_trial_report_markdown(report)
|
||||
|
||||
assert "## OpenCode Resume" not in rendered
|
||||
|
||||
@@ -183,6 +183,25 @@ def test_resume_command_prompt_mode_overrides_stored_command() -> None:
|
||||
assert any("Continue this same trial" in part for part in command)
|
||||
|
||||
|
||||
def test_resume_command_never_executes_persisted_argv() -> None:
|
||||
command = resume_command_from_result(
|
||||
{
|
||||
"task_outcome": "timeout",
|
||||
"stdout": "",
|
||||
"opencode": {
|
||||
"model": "opencode/mimo-v2.5-free",
|
||||
"variant": "high",
|
||||
"session_id": "ses_metadata",
|
||||
"resume_command": ["powershell", "-Command", "Write-Host tampered"],
|
||||
},
|
||||
}
|
||||
)
|
||||
|
||||
assert command[0:4] == ["opencode", "run", "--session", "ses_metadata"]
|
||||
assert "powershell" not in command
|
||||
assert "Write-Host tampered" not in command
|
||||
|
||||
|
||||
def test_resume_result_path_uses_next_resume_index(tmp_path: Path) -> None:
|
||||
original = tmp_path / "trial.json"
|
||||
original.write_text("{}", encoding="utf-8")
|
||||
@@ -232,6 +251,20 @@ def test_resume_trial_prints_resume_command(
|
||||
assert "opencode run --session ses_cli" in output
|
||||
|
||||
|
||||
def test_resume_trial_reports_missing_result_as_cli_error(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||
) -> None:
|
||||
from examples.agent_challenges.resume_trial import main
|
||||
|
||||
missing = tmp_path / "missing.json"
|
||||
|
||||
with pytest.raises(SystemExit) as exc_info:
|
||||
main(["--from-result", str(missing), "--print-command"])
|
||||
|
||||
assert exc_info.value.code == 2
|
||||
assert "missing.json" in capsys.readouterr().err
|
||||
|
||||
|
||||
def test_resume_trial_prints_command_for_old_raw_result(
|
||||
tmp_path: Path, capsys: pytest.CaptureFixture[str]
|
||||
) -> None:
|
||||
|
||||
@@ -181,6 +181,7 @@ def test_lda_report_workflow_builder_generates_committed_raw_plan() -> None:
|
||||
|
||||
assert workflow.name == "lda_report_case_study"
|
||||
assert validated.name == "lda_report_case_study"
|
||||
assert validated.start == "reset_board"
|
||||
assert any(node.id == "review_issues" for node in validated.nodes)
|
||||
assert payload == committed
|
||||
|
||||
@@ -219,7 +220,20 @@ async def test_lda_report_workflow_artifact_interrupt_resume_path(
|
||||
}
|
||||
)
|
||||
run_input = json.loads((EXAMPLE_DIR / "run-input.json").read_text(encoding="utf-8"))
|
||||
run_input["board_path"] = str(tmp_path / "issue-board.json")
|
||||
board_path = tmp_path / "issue-board.json"
|
||||
board_path.write_text(
|
||||
json.dumps(
|
||||
[
|
||||
{
|
||||
"id": "ISSUE-099",
|
||||
"title": "Stale demo issue",
|
||||
"url": "local://issue-board/ISSUE-099",
|
||||
}
|
||||
]
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
run_input["board_path"] = str(board_path)
|
||||
started = await server.api.run_deployment(
|
||||
deployment_id="lda_report_case_study.default",
|
||||
workflow_input=run_input,
|
||||
@@ -228,14 +242,14 @@ async def test_lda_report_workflow_artifact_interrupt_resume_path(
|
||||
assert started["status"] == "interrupted"
|
||||
assert started["interrupt"]["kind"] == "issue_review"
|
||||
assert started["interrupt"]["typed"] is True
|
||||
assert started["interrupt"]["request_schema"]["required"] == [
|
||||
assert set(started["interrupt"]["request_schema"]["required"]) == {
|
||||
"report_markdown",
|
||||
"proposed_issues",
|
||||
]
|
||||
assert started["interrupt"]["resume_schema"]["required"] == [
|
||||
}
|
||||
assert set(started["interrupt"]["resume_schema"]["required"]) == {
|
||||
"approved",
|
||||
"selected_issue_ids",
|
||||
]
|
||||
}
|
||||
proposed_ids = [
|
||||
issue["id"] for issue in started["interrupt"]["payload"]["proposed_issues"]
|
||||
]
|
||||
@@ -255,6 +269,7 @@ async def test_lda_report_workflow_artifact_interrupt_resume_path(
|
||||
assert resumed["outcome"] == "completed"
|
||||
assert resumed["output"]["approved"] is True
|
||||
assert resumed["output"]["created_issues"]
|
||||
assert resumed["output"]["created_issues"][0]["id"] == "ISSUE-001"
|
||||
assert resumed["output"]["markdown"].startswith(
|
||||
"# lda.chat Thesis And Project Readiness Report"
|
||||
)
|
||||
|
||||
@@ -154,6 +154,7 @@ async def test_workflow_tools_have_human_metadata(tmp_path: Path) -> None:
|
||||
assert "Debug traces" in read_trace_schema.get("description", "")
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_create_artifact_from_plan_exposes_plan_as_plain_object(
|
||||
tmp_path: Path,
|
||||
) -> None:
|
||||
@@ -173,6 +174,7 @@ async def test_create_artifact_from_plan_exposes_plan_as_plain_object(
|
||||
assert plan_schema.get("additionalProperties") is True
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_draft_tools_expose_plain_object_and_patch_array_schemas(
|
||||
tmp_path: Path,
|
||||
) -> None:
|
||||
@@ -203,6 +205,7 @@ async def test_draft_tools_expose_plain_object_and_patch_array_schemas(
|
||||
assert "$defs" not in patch_patch_schema
|
||||
|
||||
|
||||
@pytest.mark.asyncio
|
||||
async def test_admin_tools_have_human_metadata(tmp_path: Path) -> None:
|
||||
config = BrokerConfig(
|
||||
store_root=tmp_path / "unified_admin_metadata_store",
|
||||
|
||||
Reference in New Issue
Block a user