docs: strengthen thesis evidence and citations

This commit is contained in:
lda
2026-06-16 14:25:39 +07:00 Verified
parent e24f28929a
commit 94dac0b08b
7 changed files with 312 additions and 210 deletions
+1
View File
@@ -1,3 +1,4 @@
*.html
*.pdf
*.tex
*.typ
+4 -109
View File
@@ -1,112 +1,7 @@
# Thesis Evidence Index
This file maps thesis claims to implementation evidence. It is not prose for the
final report; it is a guardrail against unsupported claims.
The claim-to-evidence map now lives inline in
[`system-design-implementation.md`](system-design-implementation.md), Appendix B:
Evidence Index.
## Core Workflow Lifecycle
Claim: The platform separates mutable drafts, immutable artifacts, deployments,
runs, and traces.
Evidence:
- `src/wf_artifacts/models.py` — artifact/deployment models.
- `src/wf_artifacts/runs/` — run records and run store.
- `src/wf_api/service.py` — facade for workflow lifecycle operations.
- `tests/wf_api/test_artifact_api.py`
- `tests/wf_api/test_run_api.py`
## Source Provider Boundary
Claim: Workflow execution consumes source-provided capabilities without making
the core runtime MCP-specific.
Evidence:
- `src/wf_platform/sources.py` — neutral source DTOs and source policy.
- `src/wf_server/config.py` — server composition for configured sources.
- `src/wf_sources_mcp/` — MCP source family.
- `src/wf_sources_python/` — Python source family.
- `docs/source_architecture.md`
## Agent-Operable Surface
Claim: External agents can operate the workflow lifecycle through stable CLI/API
surfaces.
Evidence:
- `src/wf_cli/`
- `src/wf_transport_rpc_http/`
- `tests/wf_cli/`
- `tests/wf_transport_rpc_http/`
- `docs/wf_cli.md`
- `examples/agent_challenges/browser_click_challenge/` — challenge harness for
CLI-operability trials; aggregate model results require manual audit.
## Validation And Diagnostics
Claim: Validation and diagnostics make failed workflow states repairable.
Evidence:
- `src/wf_artifacts/validation.py`
- `src/wf_api/next_actions.py`
- `src/wf_api/source_admin.py`
- `tests/artifacts/test_validation.py`
- `tests/wf_api/test_source_admin_api.py`
## Stateful MCP Source Correctness
Claim: MCP-backed sources can preserve stateful sessions across workflow calls.
Evidence:
- `src/wf_sources_mcp/runtime/`
- `src/wf_sources_mcp/client/`
- `tests/wf_sources_mcp/test_runtime.py`
- `tests/wf_transport_rpc_http/test_mcp_backed_server_rpc.py`
## Python Source Case Study
Claim: The source-provider model is not MCP-only.
Evidence:
- `examples/report_workflow/`
- `src/wf_sources_python/`
- `tests/examples/test_report_workflow_example.py`
- `tests/wf_sources_python/test_loader.py`
- `examples/browser_click_workflow/`
- `tests/examples/test_browser_click_workflow_example.py`
- `examples/agent_challenges/browser_click_challenge/`
- `tests/examples/test_opencode_browser_click_challenge.py`
## Agent Challenge Evaluation Protocol
Claim: The project has a repeatable protocol for evaluating whether external
agents can use the product-facing CLI lifecycle, but aggregate model results are
not yet claimed.
Evidence:
- `examples/agent_challenges/browser_click_challenge/workspace_template/prompt.md`
— challenge prompt and self-report schema.
- `examples/agent_challenges/browser_click_challenge/run_opencode_trials.py`
trial runner.
- `examples/agent_challenges/browser_click_challenge/classification.py`
convenience classifier for returned reports.
- `examples/agent_challenges/browser_click_challenge/reports.py` — report
extraction and saving.
- `tests/examples/test_opencode_browser_click_challenge.py`
## Limitations
Claim: This is a prototype platform substrate, not a finished automation product.
Evidence:
- `docs/add/thesis-outline.md`
- `docs/current_roadmap.md`
- absence of scheduler/visual-editor/secret-manager production packages in
current source tree.
This file remains as a stable pointer for older roadmap and project-map links.
+4 -1
View File
@@ -36,7 +36,10 @@ $metatempfile = New-TemporaryFile
try {
Set-Content -Path $metatempfile -Value $metadata
& $pandoc_diagram --metadata-file=$metatempfile --embed-resources --standalone --citeproc @RemainingArgs
& $pandoc_diagram `
--metadata-file=$metatempfile `
--embed-resources --standalone --citeproc `
@RemainingArgs
}
finally {
Remove-Item $metatempfile
+43 -8
View File
@@ -46,14 +46,6 @@
urldate = {2026-06-16}
}
@online{swebench-verified,
title = {SWE-bench Verified},
author = {{SWE-bench}},
year = {2025},
url = {https://www.swebench.com/verified.html},
urldate = {2026-06-16}
}
@online{nist-agent-cheating-2025,
title = {Examples of Cheating in CAISI's Agent Evaluations},
author = {{National Institute of Standards and Technology}},
@@ -70,3 +62,46 @@
urldate = {2026-06-16}
}
@article{react-2022,
title = {ReAct: Synergizing Reasoning and Acting in Language Models},
author = {Yao, Shunyu and Zhao, Jeffrey and Yu, Dian and Du, Nan and Shafran, Izhak and Narasimhan, Karthik and Cao, Yuan},
year = {2022},
journal = {arXiv preprint arXiv:2210.03629},
url = {https://arxiv.org/abs/2210.03629},
urldate = {2026-06-16}
}
@article{toolformer-2023,
title = {Toolformer: Language Models Can Teach Themselves to Use Tools},
author = {Schick, Timo and Dwivedi-Yu, Jane and Dessi, Roberto and Raileanu, Roberta and Lomeli, Maria and Zettlemoyer, Luke and Cancedda, Nicola and Scialom, Thomas},
year = {2023},
journal = {arXiv preprint arXiv:2302.04761},
url = {https://arxiv.org/abs/2302.04761},
urldate = {2026-06-16}
}
@online{openai-structured-outputs-2024,
title = {Introducing Structured Outputs in the API},
author = {{OpenAI}},
year = {2024},
url = {https://openai.com/index/introducing-structured-outputs-in-the-api/},
urldate = {2026-06-16}
}
@article{auditable-agents-2026,
title = {Auditable Agents},
author = {Nian, Yi and Yuan, Aojie and Zhang, Haiyue and Li, Jiate and Zhao, Yue},
year = {2026},
journal = {arXiv preprint arXiv:2604.05485},
url = {https://arxiv.org/abs/2604.05485},
urldate = {2026-06-16}
}
@article{audit-trails-llm-2026,
title = {Audit Trails for Accountability in Large Language Models},
author = {Ojewale, Victor and Suresh, Harini and Venkatasubramanian, Suresh},
year = {2026},
journal = {arXiv preprint arXiv:2601.20727},
url = {https://arxiv.org/abs/2601.20727},
urldate = {2026-06-16}
}
+253 -87
View File
@@ -34,6 +34,13 @@ keywords:
- MCP
- Python sources
header-includes:
- |
<style>
code {
white-space: pre-wrap;
word-break: break-word;
}
</style>
- \usepackage{graphicx}
- \usepackage{booktabs}
- \usepackage{hyperref}
@@ -41,6 +48,8 @@ header-includes:
- \usepackage[dvipsnames]{xcolor}
- \usepackage{fancyhdr}
- \pagestyle{fancy}
- \usepackage{fvextra}
- \fvset{breaklines=true, breaknonspaceingroup=true, breakanywhere=true}
- \fancyhead[L]{\small lda.chat}
- \fancyhead[R]{\small\leftmark}
- \fancyfoot[C]{\thepage}
@@ -48,7 +57,7 @@ header-includes:
- \setlength{\parindent}{0pt}
- \setkeys{Gin}{width=\linewidth,height=0.55\textheight,keepaspectratio}
- \renewcommand{\arraystretch}{1.3}
- \hypersetup{pdfauthor={lda.chat}, pdftitle={Design and Implementation of lda.chat}}
# - \hypersetup{pdfauthor={lda.chat}, pdftitle={Design and Implementation of lda.chat}}
diagram:
engine:
mermaid:
@@ -57,12 +66,12 @@ diagram:
# Introduction
External LLM agents are useful workflow authors and operators, but reusable
workspace automation needs a typed execution substrate. This report describes
the design and implementation of `lda.chat`, a prototype platform where agents
can author, validate, execute, and inspect reusable workspace workflows without
making the LLM itself responsible for runtime state, validation, source binding,
or persistence.
This report assumes a setting in which external LLM agents are used as workflow
authors and operators, and asks what platform substrate they need for reusable
workspace automation. It describes the design and implementation of `lda.chat`,
a prototype platform where agents can author, validate, execute, and inspect
reusable workspace workflows without making the LLM itself responsible for
runtime state, validation, source binding, or persistence.
The central claim is that agent-facing workflow automation should separate
planning from execution. The LLM or human author can propose and revise workflow
@@ -74,7 +83,7 @@ platform represent, validate, execute, and persist reusable workspace
automations while keeping planning separate from deterministic execution?
The short version of the thesis is: the LLM plans; the runtime executes; source
providers expose capabilities; stores preserve durable lifecycle records. The
providers expose capabilities; stores preserve persisted lifecycle records. The
implementation demonstrates this model across controlled built-in, MCP, and
Python source examples.
@@ -98,23 +107,34 @@ discuss limitations and future work. Section 11 concludes.
# Problem Statement And Requirements
Many direct tool-calling agent patterns let an LLM orchestrate side effects
through sequential tool calls. This pattern can create several practical
problems for workspace automation:
A common pattern in agent systems lets an LLM orchestrate side effects
through sequential tool calls. ReAct-style prompting demonstrates interleaved
reasoning and action, while Toolformer-style work demonstrates learned external
API/tool use [@react-2022; @toolformer-2023].
The problem statement here is narrower: when a tool loop is used as a reusable
workspace automation substrate, several practical platform concerns appear.
- **Weak validation before execution.** A planner that assembles tool-call
sequences often lacks a typed contract describing what each step expects and
produces. Invalid plans reach the runtime and fail at execution time rather
than during authoring.
than during authoring. Structured-output work supports the design assumption
that schema adherence can be treated as an API/runtime contract rather than
left entirely to planner inference [@openai-structured-outputs-2024].
- **Poor resumability after interruption.** Raw tool-call loops do not
checkpoint their progress. If the process restarts, the agent must reconstruct
its prior state from scratch or lose work.
its prior state from scratch or lose work. Durable agent frameworks expose
persistence/checkpoint layers specifically because continuation, failure
recovery, and memory across interactions are runtime concerns
[@langgraph-persistence-2026].
- **Hard-to-audit traces.** Successful tool-call chains leave logs, but the
causal structure of a multi-step procedure is not separated from the transport
or provider noise. Inspecting what happened, why a step failed, or what the
intermediate state was requires manual log parsing.
intermediate state was requires manual log parsing. Recent
agent-auditability and LLM-accountability work frames action recoverability,
lifecycle coverage, and evidence integrity as explicit requirements
[@auditable-agents-2026; @audit-trails-llm-2026].
- **Limited reuse.** A successful tool-call procedure is embedded in a
conversation transcript or script. Extracting it into a named, versioned,
@@ -132,7 +152,7 @@ arbitrary office work end-to-end. Examples include document transformation, data
collection, tool and API calls, report preparation, and monitoring checks.
Scheduled execution is a future deployment mode, not implemented in this
prototype. The thesis frames the platform as a response to these pressures: a
typed execution substrate where durable lifecycle records, validation, source
typed execution substrate where persisted lifecycle records, validation, source
binding, and trace inspection are first-class platform concerns rather than
responsibilities of the planner.
@@ -143,7 +163,7 @@ The design requirements that follow from this problem statement are:
and designed to admit future source families that can be projected into the
existing capability/source contract.
3. Server, API, and CLI surfaces that external agents can drive, backed by
durable lifecycle stores.
persisted lifecycle stores.
4. Validation and inspection mechanisms intended to reduce planner
trial-and-error.
5. Next-action guidance that points an agent toward useful lifecycle operations
@@ -163,10 +183,13 @@ qualitative positioning, not a benchmark across products.
Direct tool orchestration through an LLM is the most open-ended approach: the
planner can choose tools dynamically and adapt immediately. In this report's
framing, that flexibility becomes a problem when the tool loop is also expected
to provide durable lifecycle state, validation, audit structure, and
to provide persisted lifecycle records, validation, audit structure, and
resumability. The platform argues that reusable workspace automation benefits
from separating planning from a typed execution substrate.
This comparison is to the bare tool-loop pattern, not to a tool loop embedded
inside an additional workflow, tracing, persistence, or orchestration framework.
## Generated Scripts
Generated scripts are a serious baseline. For many tasks, a script is simpler,
@@ -174,49 +197,49 @@ more maintainable, and easier to debug than a workflow graph. The platform
argument is that reusable workspace automation benefits from lifecycle
affordances that scripts do not automatically provide: typed validation, source
binding, artifact/deployment separation, run records, resumability, trace
inspection, and repairable diagnostics.
inspection, and diagnostics with repair hints.
A script can be wrapped with these affordances, but then the comparison shifts
from "script" to a custom workflow platform assembled around the script.
## Workflow Automation Platforms
Platforms such as Zapier and RPA tools are stronger today at polished
Zapier-style automation platforms are stronger today at polished
non-programmer UIs, large integration catalogs, hosted scheduling and triggers,
and operational maturity. Zapier's own documentation describes a hosted,
stateless runtime with explicit execution-time and payload constraints, plus
published Zap limits and rate limits. The prototype does not claim feature
parity with these products. Instead, it explores a different trade-off: a
platform where external AI agents can operate the full lifecycle directly,
where local Python and MCP sources share one workflow surface, and where
artifacts, deployments, runs, and traces are first-class inspectable records.
and operational maturity. This report uses Zapier as a representative hosted
automation platform rather than surveying the full RPA/workflow market.
Zapier's own documentation describes a hosted, stateless runtime with explicit
execution-time and payload constraints, plus published Zap limits and rate
limits [@zapier-operating-constraints; @zapier-zap-limits]. The prototype does
not claim feature parity with these products. Instead, it explores a different
trade-off: a platform where external AI agents can operate the full lifecycle
directly, where local Python and MCP sources share one workflow surface, and
where artifacts, deployments, runs, and traces are first-class inspectable
records.
## Agent Graph Frameworks
LangGraph-style durable agent graphs share the idea of typed execution
substrates for agent workflows. LangGraph's official documentation positions it
as an orchestration runtime for long-running, stateful agents, with persistence,
human-in-the-loop behavior, and durable execution. The lda.chat platform focuses
on reusable workspace workflows backed by explicit source providers, where the
workflow is a deployable artifact independent of any particular agent instance.
human-in-the-loop behavior, and durable execution [@langgraph-overview-2026;
@langgraph-persistence-2026]. This is not a claim that `lda.chat` is more
durable or more general than LangGraph. The difference claimed here is the
artifact/deployment/run lifecycle and source-provider binding model for
reusable workspace automations.
## Model Context Protocol
MCP is a useful protocol for exposing tools, resources, and prompts. Its
official lifecycle is a client-server connection lifecycle: initialization,
operation, and shutdown. It is not itself the workflow artifact, deployment, and
run lifecycle. The lda.chat platform treats MCP as one source family behind a
provider boundary, not as the product identity. This distinction is important:
MCP demonstrates why source-provider correctness matters, because a source may
require persistent sessions, auth context, catalog refresh, and prompt
inventory. The platform places this complexity behind a neutral
`CapabilitySource` interface.
operation, and shutdown [@mcp-tools-2025; @mcp-lifecycle-2025]. It is not
itself the workflow artifact, deployment, and run lifecycle. The lda.chat
platform treats MCP as one source family behind a provider boundary, not as the
product identity. This distinction is important: MCP demonstrates why
source-provider correctness matters, because a source may require persistent
sessions, auth context, catalog refresh, and prompt inventory. The platform
places this complexity behind a neutral `CapabilitySource` interface.
External references used for this positioning include the MCP tools and
lifecycle specifications [@mcp-tools-2025; @mcp-lifecycle-2025], LangGraph's
overview and persistence documentation [@langgraph-overview-2026;
@langgraph-persistence-2026], Zapier's operating constraints and Zap limits
[@zapier-operating-constraints; @zapier-zap-limits], and agent evaluation
discussions such as SWE-bench Verified, NIST CAISI's examples of agent
evaluation cheating, and OpenAI's SWE-bench Verified audit
[@swebench-verified; @nist-agent-cheating-2025; @openai-swebench-audit-2026].
These sources contextualize the comparison; the implementation claims in this
report remain grounded in repository evidence.
@@ -235,6 +258,7 @@ The document uses these terms with specific meanings:
| Source family | A class of source implementations. | built-in, MCP, Python |
| Source provider | Server-side code that loads or manages sources for a source family. | Python source loading |
| Tool | A provider-native operation before projection into workflow form. | MCP tool |
| Agent-operable | A surface designed for machine clients: structured output, explicit validation, stable commands, inspectability, and bounded summaries. It does not mean independently proven agent success rates. | `wf deploy validate`, `wf run trace` |
| Outcome | A control-flow label returned by a node and consumed by graph edges. | `ok`, `error`, `submitted` |
| Output | The data payload returned by a node or workflow. | `{ "report": "..." }` |
| Reducer | A pure state-merge operation selected by state schema. | `wf.std.replace`, `wf.std.append` |
@@ -267,7 +291,7 @@ aggregation. General fork/gather parallelism is future work, so this report
does not claim complete concurrent graph semantics. Interrupts represent typed
external input points. Subgraphs compose workflows as nodes.
The graph model improves the safety posture by making automation structure
The graph model improves inspectability by making automation structure
explicit. Node contracts, source requirements, state writes, outcomes,
validation gates, and trace records are visible before and after execution. The
platform does not guarantee safe behavior from provider code, credentials, or
@@ -360,9 +384,9 @@ The resolution path for a source reference is:
4. The concrete source is looked up in the server's source inventory and
delegated to the appropriate runtime handler.
This design allows workflow portability across environments: the same artifact
can be deployed with different concrete source bindings, while the workflow
graph references logical names only.
This design provides a portability mechanism across environments: the same
artifact can be deployed with different concrete source bindings, while the
workflow graph references logical names only.
# System Architecture
@@ -570,7 +594,7 @@ The implementation is organized into focused packages with clear boundaries:
| --- | --- |
| `wf_core` | Deterministic workflow kernel: graph execution, state, outcomes, trace, resume |
| `wf_authoring` | Authoring primitives: `NodeSpec`, `WorkflowBuilder`, DSL, reducer authoring, recipes |
| `wf_platform` | Neutral source DTOs, source visibility, permissions, and policy |
| `wf_platform` | Neutral source DTOs, source visibility, permission metadata, and policy |
| `wf_artifacts` | Artifact, deployment, and run models; file-backed stores; validation |
| `wf_api` | Application surface: capabilities, drafts, artifacts, deployments, runs |
| `wf_server` | `WorkflowServer` composition from config, stores, and source providers |
@@ -656,7 +680,7 @@ flowchart LR
Runs --> Live[Optional Live Source Checker]
```
The diagram is important because it shows the design contribution at the
This diagram highlights the design contribution at the
application layer: drafts, artifacts, deployments, and runs are not independent
file operations. They share source inventory, stores, event recording, runtime
execution, validation, and live-source checks through one operation context.
@@ -712,9 +736,9 @@ diagnostics and runtime status remain the source of truth.
The `NextActions` object includes `can_continue`, `can_save_now`,
`recommended_next_tool`, `reason`, `patch_examples` with concrete request
payloads, and `warnings`. This makes the surface agent-operable: an LLM agent
can read the hint and execute the suggested operation without reconstructing
the lifecycle state.
payloads, and `warnings`. This is intended to make the surface agent-operable:
an LLM agent can read the hint and execute the suggested operation without
reconstructing the lifecycle state.
(Evidence: `src/wf_api/next_actions.py`.)
@@ -738,10 +762,10 @@ The JSON-RPC-over-HTTP transport exposes the Workflow API Surface to CLI and
future HTTP clients. The transport is protocol-neutral; it maps JSON-RPC
method calls to `WorkflowApi` operations and returns structured JSON responses.
The CLI communicates over this transport. CLI commands are designed as
agent-operable first: structured output, status and inspect commands, validation
commands, compact summaries, and guarded destructive actions make the CLI a
practical surface for external agents.
The CLI communicates over this transport. CLI commands are designed for machine
clients as well as humans: structured output, status and inspect commands,
validation commands, compact summaries, and guarded destructive actions make
the CLI a practical surface for external agents.
(Evidence: `src/wf_transport_rpc_http/`, `src/wf_cli/`, `tests/wf_cli/`.)
@@ -808,6 +832,12 @@ expressiveness. Graph features such as interrupts, foreach, subgraphs, joins,
and reducer behavior are covered by targeted tests and code evidence in the
evaluation section.
The thesis-critical automated report-workflow run is intentionally narrow: it
executes a single deterministic extraction node through the artifact,
deployment, and run lifecycle. The same Python source exposes additional
capabilities for discovery and extensibility evidence, and the supplemental
browser-click example covers a serial three-node workflow.
## Case Study Components
The example bundle lives at
@@ -988,7 +1018,7 @@ benchmark.
| Deployment/source binding | Not inherent | Manual config | Platform-specific | Yes |
| Typed validation before run | Tool-schema dependent | Custom | Varies | Yes |
| Durable run record | Not inherent | Custom | Often yes | Yes, at stopped boundaries |
| Source drift diagnostics | Not inherent | Custom | Varies | Yes, controlled examples |
| Source drift diagnostics | Not inherent | Custom | Varies | Yes, schema-hash controlled examples |
| Agent-operable repair hints | Not inherent | Custom | Usually human UI | Prototype support |
| Scheduling | Depends on agent | External scheduler | Yes | Future work |
@@ -1007,12 +1037,37 @@ The evidence supporting the thesis claims includes:
| Interrupted runs resume at explicit boundaries | `tests/wf_api/test_run_api.py` and resume-concurrency tests | Stopped run state is persisted and resumed through the run API | Pass in focused test suite |
| Python source lifecycle works | `tests/examples/test_report_workflow_example.py` | Python capability -> artifact -> deployment -> run completes | Pass in focused test suite |
| Serial multi-node workflow works | `tests/examples/test_browser_click_workflow_example.py` | `open_click_page` -> `wait_for_click` -> `collect_snapshots` completes with before/after evidence | Pass in focused test suite |
| Agent challenge harness exists | `examples/agent_challenges/browser_click_challenge/` | Agents can be prompted, classified, and manually audited against a browser-click workflow challenge | Harness tests pass; aggregate model results pending |
| Agent challenge harness implementation exists; no aggregate agent-performance claim | `examples/agent_challenges/browser_click_challenge/` | Agents can be prompted, classified, and manually audited against a browser-click workflow challenge | Harness tests pass; aggregate model results pending |
| CLI and JSON-RPC share the API surface | `tests/wf_transport_rpc_http/` and `tests/wf_cli/` | Transport and CLI delegate to the same workflow operations | Pass in focused test suite |
The table summarizes repository evidence; it is not a substitute for rerunning
the verification commands before final submission.
## Verification Snapshot
This draft includes one focused verification snapshot to make the evidence
claims auditable from the text. A final submission should regenerate this table
from the exact submitted commit.
| Field | Value |
| --- | --- |
| Date run | 2026-06-16 |
| Baseline commit | `e24f2892` before subsequent document-polish edits |
| Command | `uv run pytest tests/docs tests/examples/test_report_workflow_example.py tests/examples/test_browser_click_workflow_example.py tests/examples/test_opencode_browser_click_challenge.py tests/artifacts/test_validation.py tests/wf_api/test_run_api.py -q` |
| Result | `72 passed in 9.22s` |
| Environment | Local Windows development environment, Python via `uv` |
| Scope | Documentation links, report workflow, browser-click workflow, challenge harness, deployment validation, and run API tests |
## Implemented Scope Matrix
| Area | Implemented evidence | Not claimed | Future work |
| --- | ---- | ---- | ---- |
| Workflow lifecycle | Draft, artifact, deployment, run, trace, and list/inspect/resume surfaces | Exactly-once execution or arbitrary mid-node crash recovery | Transactional stores and richer run debugging |
| Source providers | Built-in, MCP, and Python source families | Symmetric feature depth across all providers | Provider add/update/remove/reload lifecycle |
| Execution model | Outcome-routed graph with node, condition, foreach, subgraph, join, interrupt, and end steps | General fork/gather programming model | Parallel fork/gather and aggregation |
| Agent-operable surface | CLI, JSON-RPC, validation diagnostics, next-action hints, compact output | Measured agent success or token reduction | Broader agent challenge suite and aggregate evaluation |
| Auth/security | Auth record plumbing and source diagnostics | Production security, encrypted-at-rest secrets, RBAC, sandboxing | Secret-manager integration and policy enforcement |
### Architecture And Code Walkthrough
The four-layer architecture (core, API surface, server composition, transport)
@@ -1032,8 +1087,8 @@ from plan to completed run.
Tests verify that draft validation catches schema violations, deployment
validation detects source drift, and diagnostics include repair hints. The
validation tests demonstrate that failed states are machine-readable and
repairable.
validation tests demonstrate that failed states are machine-readable and include
repair guidance.
(Evidence: `tests/artifacts/test_validation.py`, `tests/wf_api/test_source_admin_api.py`.)
@@ -1258,11 +1313,11 @@ The following areas are identified as likely future work:
# Conclusion
External LLM agents can author and operate workflows, but reusable workflow
life cycle records should live in a typed platform substrate. This report
described the design and implementation of `lda.chat`, a prototype platform
that separates planning from execution across controlled built-in, MCP, and
Python source examples.
External LLM agents can be used to author and operate workflows, but reusable
workflow lifecycle records should live in a typed platform substrate. This
report described the design and implementation of `lda.chat`, a prototype
platform that separates planning from execution across controlled built-in,
MCP, and Python source examples.
The implementation supports five claims:
@@ -1272,10 +1327,10 @@ The implementation supports five claims:
one workflow surface, and is designed to admit future source families that
can be projected into the existing capability/source contract without
core-runtime changes.
3. Validation and diagnostics produce machine-readable repairable failure
states intended to reduce planner trial-and-error.
4. The CLI and JSON-RPC transport provide an agent-operable surface that
external LLM agents can drive without direct runtime access.
3. Validation and diagnostics produce machine-readable failure states with
repair hints, intended to reduce planner trial-and-error.
4. The CLI and JSON-RPC transport provide a surface designed for external LLM
agents to drive without direct runtime access.
5. The deterministic report-workflow case study demonstrates the full lifecycle
from config validation through run execution and trace inspection.
@@ -1388,18 +1443,119 @@ run trace <run_id> --from 0 --limit 5
## Appendix B: Evidence Index
See [evidence-index.md](evidence-index.md) for the full claim-to-evidence map.
This appendix maps thesis claims to implementation evidence. It is a guardrail
against unsupported claims and complements the focused verification snapshot in
the Evaluation section.
| Claim | Evidence |
| --- | ----- |
| Workflow lifecycle models | `src/wf_artifacts/models.py`, `src/wf_artifacts/runs/` |
| Source provider boundary | `src/wf_platform/sources.py`, `src/wf_server/config.py` |
| JSON-RPC/CLI surface | `src/wf_transport_rpc_http/`, `src/wf_cli/` |
| Python source case study | `examples/report_workflow/`, `tests/examples/test_report_workflow_example.py` |
| MCP stateful runtime | `tests/wf_sources_mcp/test_runtime.py`, `tests/wf_transport_rpc_http/test_mcp_backed_server_rpc.py` |
| Validation/diagnostics | `src/wf_artifacts/validation.py`, `src/wf_api/next_actions.py`, `tests/artifacts/test_validation.py` |
| Agent-operable surface | `src/wf_cli/`, `tests/wf_cli/` |
| Platform domain | `src/wf_api/service.py`, `tests/wf_api/test_artifact_api.py`, `tests/wf_api/test_run_api.py` |
### Core Workflow Lifecycle
Claim: The platform separates mutable drafts, immutable artifacts, deployments,
runs, and traces.
Evidence:
- `src/wf_artifacts/models.py` — artifact/deployment models.
- `src/wf_artifacts/runs/` — run records and run store.
- `src/wf_api/service.py` — facade for workflow lifecycle operations.
- `tests/wf_api/test_artifact_api.py`
- `tests/wf_api/test_run_api.py`
### Source Provider Boundary
Claim: Workflow execution consumes source-provided capabilities without making
the core runtime MCP-specific.
Evidence:
- `src/wf_platform/sources.py` — neutral source DTOs and source policy.
- `src/wf_server/config.py` — server composition for configured sources.
- `src/wf_sources_mcp/` — MCP source family.
- `src/wf_sources_python/` — Python source family.
- `docs/source_architecture.md`
### Agent-Operable Surface
Claim: The workflow lifecycle is designed to be operated by external agents
through stable CLI/API surfaces.
Evidence:
- `src/wf_cli/`
- `src/wf_transport_rpc_http/`
- `tests/wf_cli/`
- `tests/wf_transport_rpc_http/`
- `docs/wf_cli.md`
- `examples/agent_challenges/browser_click_challenge/` — challenge harness for
CLI-operability trials; aggregate model results require manual audit.
### Validation And Diagnostics
Claim: Validation and diagnostics make failed workflow states machine-readable
and include repair hints.
Evidence:
- `src/wf_artifacts/validation.py`
- `src/wf_api/next_actions.py`
- `src/wf_api/source_admin.py`
- `tests/artifacts/test_validation.py`
- `tests/wf_api/test_source_admin_api.py`
### Stateful MCP Source Correctness
Claim: MCP-backed sources can preserve stateful sessions across workflow calls.
Evidence:
- `src/wf_sources_mcp/runtime/`
- `src/wf_sources_mcp/client/`
- `tests/wf_sources_mcp/test_runtime.py`
- `tests/wf_transport_rpc_http/test_mcp_backed_server_rpc.py`
### Python Source Case Study
Claim: The source-provider model is not MCP-only.
Evidence:
- `examples/report_workflow/`
- `src/wf_sources_python/`
- `tests/examples/test_report_workflow_example.py`
- `tests/wf_sources_python/test_loader.py`
- `examples/browser_click_workflow/`
- `tests/examples/test_browser_click_workflow_example.py`
- `examples/agent_challenges/browser_click_challenge/`
- `tests/examples/test_opencode_browser_click_challenge.py`
### Agent Challenge Evaluation Protocol
Claim: The project has a repeatable protocol for evaluating whether external
agents can use the product-facing CLI lifecycle, but aggregate model results are
not yet claimed.
Evidence:
- `examples/agent_challenges/browser_click_challenge/workspace_template/prompt.md`
— challenge prompt and self-report schema.
- `examples/agent_challenges/browser_click_challenge/run_opencode_trials.py`
trial runner.
- `examples/agent_challenges/browser_click_challenge/classification.py`
convenience classifier for returned reports.
- `examples/agent_challenges/browser_click_challenge/reports.py` — report
extraction and saving.
- `tests/examples/test_opencode_browser_click_challenge.py`
### Limitations
Claim: This is a prototype platform substrate, not a finished automation
product.
Evidence:
- `docs/add/thesis-outline.md`
- `docs/current_roadmap.md`
- Absence of scheduler, visual-editor, and secret-manager production packages
in the current source tree.
## Appendix C: Agent Challenge Harness
@@ -1423,6 +1579,15 @@ The current browser-click prompt allows two product-facing authoring paths:
2. **Raw-plan path.** Write a raw workflow plan and load it with
`wf artifact create-from-plan`, then deploy and run.
```mermaid
flowchart LR
Transcript[Agent transcript and files] --> YAML[YAML self-report]
YAML --> Classifier[Automatic convenience classification]
Transcript --> Audit[Manual audit]
Classifier --> Audit
Audit --> Outcome[Official outcome]
```
The challenge report is a YAML self-report with fields for product-path use,
helper-script use, workflow file, deployment id, run id, before/after booleans,
read-behavior flags, attempt counts, missed requirements, and notes. The harness
@@ -1433,14 +1598,15 @@ product source code, adjacent attempts, prior stores, or existing solutions.
This distinction is intentional. Agent benchmark literature and practice show
that automated scores and self-reports can be misleading when an agent can
inspect hidden answers, prior artifacts, source code, or evaluator state. The
harness therefore records possible invalidation flags such as helper-script
bypass, adjacent-attempt leakage, prior-store reuse, product-code dependency,
false YAML claims, timeouts, parse failures, and missing run evidence.
inspect hidden answers, prior artifacts, source code, or evaluator state
[@nist-agent-cheating-2025; @openai-swebench-audit-2026]. The harness therefore
records possible invalidation flags such as helper-script bypass,
adjacent-attempt leakage, prior-store reuse, product-code dependency, false YAML
claims, timeouts, parse failures, and missing run evidence.
At the time of this report, the harness and browser-click workflow are
implemented and unit-tested, and preliminary manual trials have already informed
CLI and prompt improvements. The report does not claim aggregate model success
implemented and unit-tested, and informal preliminary trials have informed CLI
and prompt improvements. The report does not claim aggregate model success
rates, timeout distributions, command counts, or retry-reduction results yet.
Those require repeated, clean-workspace trials across the selected free opencode
models and manual audit of saved trial reports.
+1 -1
View File
@@ -16,6 +16,6 @@ $filter = Join-Path $PSScriptRoot 'diagram.lua'
$env:MERMAID_BIN = whereis mmdc
$env:TIKZ_BIN = 'xelatex'
Write-Host @args
Write-Host $($args -join ' ')
& pandoc --lua-filter $filter --pdf-engine=xelatex @args
+4 -2
View File
@@ -11,14 +11,16 @@ def markdown_links(text: str) -> set[str]:
return set(re.findall(r"(?<!!)\[[^\]]+\]\(([^)]+)\)", text))
def test_big_doc_links_case_study_and_evidence_index() -> None:
def test_big_doc_links_case_study_and_embeds_evidence_index() -> None:
doc = (ROOT / "docs" / "add" / "system-design-implementation.md").read_text(
encoding="utf-8"
)
links = markdown_links(doc)
assert "evidence-index.md" in links
assert any(link.startswith("../../examples/report_workflow") for link in links)
assert "## Appendix B: Evidence Index" in doc
assert "### Core Workflow Lifecycle" in doc
assert "### Agent Challenge Evaluation Protocol" in doc
def test_project_map_links_big_doc() -> None: