48 KiB
Workflow Runtime Context
This context defines the core workflow runtime language used by wf_core,
wf_authoring, and the MCP-facing workflow platform.
Language
lda.chat: An AI agent platform for authoring and executing workspace workflows. Avoid: Chatbot, MCP server
Agent-Compatible Infrastructure: A platform surface that external AI agents can use to author, validate, run, and inspect workflows without bundling a specific agent brain. Avoid: Built-in autonomous agent, chatbot implementation
AI Agent Platform:
The broader product framing for lda.chat when discussing the intended user
experience and ecosystem role. Use carefully today because the platform is
currently driven by external agents rather than a bundled autonomous agent.
Avoid: Claiming lda.chat already contains the agent brain
Workspace Operator: The human beneficiary and controller of workspace automation. Avoid: End user, consumer user
Workflow Owner: The human who wants the workflow to exist, run, and produce useful output. They may review the workflow result without directly operating the CLI or API; an external LLM agent can drive those surfaces for them. Avoid: Assuming the CLI user and product beneficiary are always the same actor
Workspace: A user-controlled digital work environment where office-style work happens, including files, tools, services, credentials, and project context exposed through approved capability sources. Avoid: Filesystem directory, user account
LLM Agent: An external AI agent, such as Claude Desktop or OpenCode, that can use the platform surfaces to plan, edit, validate, and invoke workflow operations. Avoid: Built-in agent brain, chatbot
Workspace Workflow: A reusable, typed procedure for digital work in a workspace, represented as a graph with schemas, state, outcomes, source bindings, and durable run records. Avoid: Sequence of tool calls, prompt chain
Durable Workflow Contract: The persisted workflow intent across time: artifact, deployment bindings, validation against current sources, run records, and traces. Persistence keeps the objects; validation determines whether the current environment can still run them. A deployment becoming unrunnable after incompatible source changes is a contract-preserving outcome, not silent failure. Avoid: Treating durability as only file persistence or tying it to one store implementation
Source Drift: Change in a source catalog, capability contract, provider availability, or source binding after an artifact or deployment was created. Source drift is a normal lifecycle state surfaced through validation and diagnostics, not an unexpected crash. Avoid: Assuming old deployments remain runnable forever
Workflow Artifact: An immutable, versioned workflow definition saved from authoring state. It captures the workflow graph and declared requirements, but it is not itself a live execution. Avoid: Mutable draft, run record
Editable Workflow: A process-local, mutable Python authoring object used to construct or revise a workflow graph before saving an immutable workflow artifact. Editing an existing artifact creates a new editable workflow seeded from that exact version; it never mutates the saved artifact. Avoid: Draft workspace, workflow artifact, deployment
Remote Capability: A Python client object reconstructed from an inspected workflow capability. It preserves the canonical capability reference, schemas, outcomes, direct-call behavior, and graph-authoring contract while hiding transport payloads. Avoid: Raw provider tool, unvalidated JSON-RPC result, local NodeSpec
Workflow Deployment: A binding from a workflow artifact version to concrete source/runtime context. Deployment validation determines whether the artifact can currently run. Avoid: Artifact version, draft workspace, incidental runtime config
Binding Contract: The deployment-level mapping from logical workflow requirements to concrete sources or runtime context. Binding contracts make workflows portable across workspaces/accounts while allowing validation to detect missing, disabled, or incompatible sources. Avoid: Ad-hoc environment variables, hidden source lookup
Step Input Binding: A node-local assignment from one workflow source or constructed input expression to one capability input target. Avoid: Indexed local targets, output binding
Input Expression: A finite recursive value construction made only from JSON literals, workflow paths, arrays, and objects. It assembles data but does not compute behavior. Avoid: Formula, transform, inline capability call
Platform Source:
A process-provided source with a fixed identity, such as wf.std or
wf.source. Platform sources can satisfy workflow requirements without
deployment self-bindings because their logical source id is also their concrete
source id, and deployments should not override them with custom bindings.
Avoid: Configured account source, user-installed provider
Configured Source: A server/operator-selected source such as an MCP connection, trusted Python module registry, or future OpenAPI provider. Configured sources are where deployment bindings, source registry state, auth refs, catalog cache, and health diagnostics matter most. Avoid: Built-in standard library source
Source Inventory:
The listable capabilities owned by a source: workflow-callable NodeSpecs,
provider-native tools, resources, prompts, and related metadata. Inventory
inspection is not the same as invoking a tool, reading a resource, or rendering
a prompt.
Avoid: Assuming list equals execute
Source Resource Ref:
An inert pass-by-value reference to source-owned content, carrying a logical
source and provider URI. The URI is not globally meaningful by itself; runtime
must resolve the logical source through deployment/platform context before a
helper such as wf.source.read_resource can dereference it.
Avoid: Resource Ref as the canonical term, bare URI, immediate content fetch
Prompt Inventory: Source-owned prompt/template names and metadata. Listing prompts is safe inventory; rendering a prompt is an upstream operation that may be stateful and needs a concrete graph use case, argument schema, and bounded output policy before becoming a workflow helper. Avoid: Source Prompt Ref as premature symmetry, treating prompt list entries as already-rendered text
Workflow Portability: The goal that a workflow artifact can describe logical requirements separately from concrete source/account bindings. Portability is scoped: deployments bind artifacts to a specific workspace environment, and local Python code, MCP catalogs, auth records, and source stores may still differ between environments. Avoid: Universal run-anywhere portability
Draft Workspace: Mutable workflow authoring state where an agent or user can patch, validate, and iterate before saving an immutable artifact. Avoid: Deployed workflow
Workflow Run: An execution record for a deployment, including status, diagnostics, output, trace, and resumable stopped/interrupted state where applicable. Avoid: Artifact, deployment
Run Resume: A supporting durable-execution capability where stopped or interrupted workflow state can be inspected and continued later. It proves the lifecycle is more than final-output storage, but it is not yet full time-travel debugging. Avoid: Full LangGraph-style time travel
Run Trace: A bounded debug record attached to a workflow run. Current claims should be based on code-supported trace counts and caller-bounded trace slices for debugging, not broad production observability. Avoid: Distributed tracing, metrics platform, OpenTelemetry claims
Evidence-Backed Claim: A thesis claim grounded in implemented code, tests, live smoke runs, examples, or clearly marked future work. The current project stage is making the platform good and correct; making it fast is later unless performance is measured. Avoid: Unmeasured performance, reliability, security, or production-readiness claims
Evidence Package: The thesis evidence set: architecture/code walkthrough tied to the four-layer model, automated tests, live smoke runs, source-provider examples, run persistence/resume checks, stateful MCP reuse checks, Python source integration, and a small failed-attempt case study. Avoid: Large user study claims without the study
Case Study Workflow: A concrete representative workflow used in the thesis evidence package. Prefer a document/report preparation task over toy echo tools: parse or transform a small document, extract key points or action items, and produce structured report output. It should run deterministically with current sources; LLM enhancement can be future or optional. Avoid: Echo-only demos, flaky external services
Reproducible Example Environment: A small runnable example bundle for thesis evidence: fixture input files, source provider code, workflow config, store directory setup, and CLI/server commands. It should let the case study be rerun without relying on hidden local state. Avoid: Screenshot-only demo, undocumented local setup
Representative Workspace Task: A recurring digital work pattern such as document processing, project or repository monitoring, data/report preparation, or scheduled workspace operations. Avoid: Generic business process, arbitrary web task
Automation Platform Baseline: An existing workflow automation product, such as Zapier, used as a comparison point for maturity, speed, usability, integrations, and scheduling. Avoid: Direct competitor claim, full feature parity
Prototype Platform Substrate: The implemented backend, API, CLI, source-provider, and durable workflow lifecycle that external agents or future UI surfaces can drive. Avoid: Final conversational product, complete user interface
Prototype Platform: A usable but incomplete platform implementation that demonstrates the core workflow lifecycle and source-provider model. It is not yet a full product because scheduling, visual workflow editing, production auth/secret storage, general fork/gather, richer source lifecycle, and broad evaluation remain future work. Avoid: Production-ready product, finished agent app
Platform Architecture Contribution: The thesis contribution: separating external-agent planning from typed workflow execution through artifact/deployment/run lifecycle, source-provider boundaries, durable API/CLI/server surfaces, and validation/inspection mechanisms that reduce planner trial-and-error. Avoid: Framing the work as only another automation tool or AI model
Automation Baseline: A comparison point used to position the prototype. Direct LLM tool use, manual scripts, and mature platforms such as Zapier or RPA systems expose different trade-offs: durability and reuse, accessibility and adaptability, or integration and operational maturity. Avoid: Treating one baseline as the entire problem space
Workflow Automation Accessibility: The goal of making reusable workflow automation available without requiring the workspace operator to hand-code scripts or manually chain tools. Avoid: Fully automatic correctness, no-review automation
Offline Scheduling: Future execution of deployed workspace workflows without an active chat or interactive session. Avoid: Background prompt, cron job
Workflow Visualization: Future UI support for inspecting, editing, and explaining workflow graphs visually. Avoid: Current runtime feature, trace log
LLM Node: A workflow node that calls an LLM as one capability inside the graph, rather than making the whole runtime itself an LLM loop. Avoid: Built-in planner, hidden agent loop
Deterministic Execution Substrate: The workflow control layer that validates schemas, routes outcomes, commits state, records trace, and persists runs predictably even when individual capabilities call nondeterministic tools, APIs, Python code, or LLMs. Avoid: Deterministic external tools, deterministic LLM output
Runtime Scope: The state universe for one workflow invocation, containing its input and committed state. Every state lineage belongs to exactly one runtime scope. Avoid: Execution frame, branch, global run state
State Lineage: One isolated state worldview inside a runtime scope. Its ancestry describes state visibility, independently of which execution frame currently uses it. Avoid: Execution frame, scheduler task, activation token
Execution Frame: A schedulable graph cursor that executes against one state lineage. Frame ancestry describes scheduling ownership and block/wake behavior, not state ancestry. Avoid: State lineage, workflow invocation, operating-system thread
Control Region: The one static foreach-owner stack assigned to a workflow node use from graph entry. A node use cannot belong to several control regions. Avoid: Execution frame, state lineage, dynamic call stack
Frame Completion: The end of one execution frame. Its owner determines whether this completes a workflow invocation, returns from a subgraph, or completes one foreach item. Avoid: Always completing the run, domain outcome, implicit break
Foreach Return: Completion of one foreach item through a graph back-edge targeting the item's owning foreach. The owner resumes its serial controller or concurrent barrier; the item does not execute the foreach node again. Avoid: Workflow end, generic graph cycle, implicit break
Foreach Activation: One dynamic visit to a foreach controller, owning that visit's barrier and item executions. Re-entering the same node use creates another activation. Avoid: Foreach node use, control region, item frame
Fork: An explicit workflow control step that creates several concurrent branch activations. A fork is distinct from an ordinary outcome, which selects exactly one continuation. Avoid: Multiple matching edges, implicit broadcast, async handler call
Gather: An explicit workflow control step that waits for compatible activation tokens at every declared local input slot, merges their lineage-local state according to reducers and gather policy, and emits one continuation. Several alternative edges may satisfy one slot. Avoid: Pass-through join marker, arbitrary graph convergence, reference to one originating fork
Activation Token: A runtime value identifying one control-flow arrival, its state lineage, activation context, and merge provenance. Compatible tokens can rendezvous at a gather without mixing loop iterations, subgraph invocations, or repeated fork activations. The lineage determines the token's runtime scope. Avoid: Static edge, node output payload, scheduler position
Gather Slot: A named local rendezvous requirement on a gather. A gather requires all of its slots, while any one compatible incoming edge assigned to a slot can satisfy that slot. Avoid: Fork branch reference, positional input, executable edge predicate
Planner Efficiency: The degree to which the platform reduces LLM trial-and-error when creating or running workflows. Typed schemas, source catalogs, validation errors, dry checks, inspectable traces, compact CLI output, and reusable artifacts should help an external LLM agent converge without spending excessive tokens probing blindly. Representative evidence includes fewer failed attempts before a valid workflow is created and run. Avoid: Assuming an LLM planner is cheap enough to brute-force the workflow
Failed Workflow Attempt: An agent action that tries to advance workflow creation or execution but cannot be reused because it hits an avoidable structural, validation, source-binding, configuration, session, or runtime issue. Examples include failed draft validation, non-runnable deployments, source/session failures, and run failures caused by incorrect workflow structure or provider assumptions. Avoid: Counting deliberate exploration or user clarification as workflow failure
Validation-Centered Workflow Lifecycle: The product principle that LLM-authored workflows should move through explicit validation gates before and after execution. Config validation, capability contracts, draft validation, deployment validation, run inspection, and output safety reduce blind trial-and-error and make workflow automation reviewable. Avoid: Treating validation as a convenience check after implementation
Next Actions: Advisory lifecycle guidance returned to clients and LLM agents to suggest the next useful operation, such as validating a draft, patching a workspace, saving an artifact, or inspecting a run. Next actions guide planner behavior but do not replace validation authority. Avoid: Treating next actions as hard permission checks
Machine-Client UX: API and CLI response design for LLM agents and other programmatic clients. It uses compact structured payloads, stable fields, diagnostics, and next actions instead of relying on prose or oversized raw provider output. Avoid: Assuming only human-readable UI needs interaction design
Agent-Operable CLI: A CLI designed so external LLM agents can drive workflow lifecycle operations predictably through structured output, compact summaries, validation commands, status/inspect/list operations, and guarded destructive actions. Humans can use it, but it is not only a human terminal UI. Avoid: Treating the CLI as the workflow runtime
Workflow API Surface: The application contract for workflow lifecycle operations: capabilities, drafts, artifacts, deployments, runs, admin/source operations, validation, inspection, and next actions. Process-local clients, JSON-RPC, future MCP frontends, and UI backends can implement or consume this surface. Avoid: Treating JSON-RPC as the product boundary
Transport Adapter: An implementation that exposes the Workflow API Surface over a specific communication mechanism, such as process-local calls or JSON-RPC-over-HTTP. Transport adapters should not define workflow semantics. Avoid: Putting workflow behavior in the transport
Four-Layer Platform Model: The canonical architecture split: workflow core, platform domain, workflow API surface, and server/transport composition. The core owns execution semantics; the platform domain owns artifacts, deployments, runs, sources, bindings, validation, stores, and admin concepts; the API surface exposes lifecycle operations; server/transport composition wires concrete stores, sources, runtimes, and communications. Avoid: Collapsing server, API, platform, and core into one layer
Authoring Support: Helper libraries and tools that create workflow specs, node specs, drafts, and wrapper structures for API surfaces, source providers, and MCP/admin tools. Authoring support feeds the platform domain but is not a separate runtime layer and does not own workflow semantics. Avoid: Treating authoring helpers as the execution engine
Repairable Validation Failure: A validation failure that identifies what failed, where it failed, why it matters, and what the caller can try next. It should help an LLM agent patch the workflow instead of causing another blind probe. Avoid: Tracebacks, generic bad-request messages, silent rejection
Graph Safety Boundary: The safety argument that a typed workflow graph is easier to validate, inspect, reuse, and constrain than arbitrary generated code. The graph model limits where logic lives and makes source bindings, schemas, state, and outcomes explicit. This does not make external tools, trusted Python sources, or credentials safe by default. Avoid: Claiming workflow graphs eliminate all automation risk
Code Boundary: Code belongs behind source providers and capabilities. A workflow should orchestrate typed capabilities through graph structure, bindings, outcomes, state, and interrupts; it should not embed arbitrary code directly. Future escape hatches such as unsafe Python or LLM nodes should still appear as source capabilities. Avoid: Treating a workflow as an opaque script blob
LLM Node: A future source capability that calls an LLM for a bounded task such as summarization, classification, extraction, or controlled looping. It should have typed inputs, outputs, and outcomes like other capabilities. Avoid: Hidden LLM orchestration engine
Source-Provider Correctness: The requirement that a source provider preserve the behavior expected by the external system it represents. Some providers are not meaningfully stateless: they may depend on initialized sessions, created resources, authentication context, or provider-local state. Workflow execution should not silently degrade those sources into one-off calls when stateful behavior is part of correctness. Avoid: Treating all capability calls as independent stateless requests
MCP Source:
An MCP-backed source family that exposes upstream tools, resources, and prompts
through the workflow source boundary. MCP is a source integration and stress
test for source-provider correctness; it is not the identity of the whole
platform.
Avoid: Treating lda.chat as only an MCP server or MCP proxy
MCP Frontend: A future transport/frontend role where external MCP clients could operate the Workflow API Surface. This is distinct from MCP Source, which consumes upstream MCP servers as workflow capabilities. The current thesis should not claim the new platform already has a clean MCP frontend. Avoid: Confusing upstream MCP sources with client-facing MCP transport
MCP Widget Proxying: Proxying upstream MCP UI resources/widgets, iframe metadata, and related client resource behavior through the platform. This is not supported now because it requires protocol-specific UI/resource handling beyond ordinary workflow capabilities. Avoid: Claiming durable workflows preserve upstream MCP interactive widgets
First-Party Workflow UI: A future UI surface owned by the platform, such as listing, inspecting, and editing workflows or deployments. This is different from proxying upstream MCP widgets. Avoid: Treating first-party workflow UI as MCP widget passthrough
Reviewable Lifecycle: The current human-in-the-loop property: drafts, deployments, runs, diagnostics, traces, and guarded destructive commands create review points before or after execution. This is not a full approval, policy, roles, or multi-user review system. Avoid: Enterprise approval workflow
Auth Plumbing: Prototype support for auth records, source auth diagnostics, and admin surfaces. Auth matters for source readiness, but production secret management and verified end-to-end credential workflows are not thesis-core claims yet. Avoid: Production secret manager, fully verified auth system
Python Source:
A trusted project-local source family that exposes Python-authored NodeSpec
registries as workflow capabilities. Python sources provide developer
extensibility and prove the source model is not MCP-only, but they are not
sandboxed non-programmer plugins yet.
Avoid: Safe end-user extension, arbitrary untrusted code
OpenAPI Source: A future source family that could expose HTTP/OpenAPI operations as typed workflow capabilities. It would broaden integration coverage, but MCP and Python sources are already enough to demonstrate the source-provider boundary. Avoid: Claiming HTTP/OpenAPI source support exists today
Natural-Language Authoring: Using natural language as the way a workspace operator communicates intent to an external LLM agent. The durable result should be a typed workflow graph and deployment, not an opaque prompt transcript. Avoid: Treating natural language as the execution model
Reusable Workspace Procedure: A repeatable digital-work procedure that can be represented as a typed workflow: transforming documents, collecting data, calling tools or APIs, preparing reports, monitoring changes, or running scheduled checks. The platform automates these procedures rather than claiming to solve arbitrary office work end-to-end. Avoid: Arbitrary job completion, one-off tool click
Offline Scheduling: Future execution of deployed workflows without an active chat or CLI session. Persisted deployments and server-side execution make scheduling plausible, but scheduling itself is not implemented yet. Avoid: Claiming production background scheduling exists today
Scheduler Foundation: The runtime model that selects runnable frames and advances workflow execution without assuming there is only one active cursor. Avoid: Foreach feature, concurrent foreach implementation
Frame: An execution cursor for one active portion of a workflow run. Avoid: Thread, task
Frame Set: The collection of lifecycle frames owned by a run, including pending, running, interrupted, and completed frames until cleanup rules remove them. Avoid: Running frames
Runnable Frame: A pending frame that the scheduler may select for the next deterministic runtime step. Avoid: Active frame, async task
Ready Queue: The ordered list of runnable frame identifiers that defines deterministic scheduling order. Avoid: Frame scan, implicit dict order
Foreach Policy: The future runtime configuration that decides concurrent foreach admission, item failure handling, and quiescence behavior. Avoid: Parallel flag
Blocked Frame: A live frame that is waiting on child frame completion or an external resume event and is not currently runnable. Avoid: Callback, sleeping thread
Block Reason: Typed metadata describing what a blocked frame is waiting for. Avoid: Ad hoc metadata flag
Interrupt: A run-level pause requested by one frame that waits for external input before any more workflow scheduling happens. Avoid: Per-frame prompt queue, background prompt
Runtime Failure: An execution failure that stops scheduling unless a future policy explicitly handles it. Avoid: Error outcome
Trace: An append-only chronological record of executed workflow steps. Avoid: Sorted report, graph view
Reducer: A named merge rule that combines multiple writes to the same state path. Avoid: Last writer wins
State Patch: A validated set of state writes produced by one executed step before it is committed to run state. Avoid: Partial mutation
State Visibility: The committed state a frame may read based on workflow topology and completed upstream work. Avoid: Shared live object
Lineage Isolation: The rule that a concurrent child frame sees parent-visible state plus its own ancestor writes, but not sibling branch writes. Avoid: Global committed state
Barrier: A merge boundary that waits for multiple child or upstream frames and combines their lineage patches before continuation. Avoid: Noop join
Run: One execution attempt of a workflow with input, state, frames, trace, and final output. Avoid: Job, invocation
Relationships
- A Run owns one or more Frames.
- Every State Lineage belongs to exactly one Runtime Scope.
- Every Execution Frame executes against one State Lineage and derives its runtime scope through that lineage.
- Frame ancestry expresses scheduling ownership; lineage ancestry expresses state visibility.
- Every workflow node use belongs to exactly one static Control Region.
- Each dynamic visit to a foreach node use creates a distinct Foreach Activation whose item frames share that activation identity.
- A Foreach Return completes an item frame and returns control to its owning foreach without ending the workflow scope.
- A Frame Set is the source of truth for runtime cursors.
- The Scheduler Foundation selects one Runnable Frame at a time in the first pass.
- The Ready Queue defines which Runnable Frame is selected next.
- A Blocked Frame is intentionally absent from the Ready Queue until the child/event it waits on wakes it.
- A Blocked Frame should carry a Block Reason so deadlocks and wakeups are explainable.
- Concurrent foreach and native subgraphs depend on the Scheduler Foundation.
- Concurrent foreach requires a Foreach Policy before public support is enabled.
- Future concurrent foreach should be bounded by default.
max_activeandmax_outstandingprevent one workflow from spawning unbounded MCP, HTTP, browser, or external service calls. - Future concurrency-specific foreach settings should live in a nested
policy object rather than expanding
ForeachNodewith many top-level fields. - Item error handling is foreach-wide, not concurrent-only. Today, serial
foreach supports
fail; concurrent foreach supportsfail,skip, andcollect. Future serial foreach can reuse the sameskip/collectpolicy shape when its execution path is upgraded. collectitem error policy must declare an explicit destination for structured item errors. Collected errors should be ordered by item index, not async completion order.- Item error policy handles runtime failures inside an item frame, such as
thrown handler errors, schema failures, runtime-error utility nodes, or
invalid graph execution. It does not intercept normal graph outcomes like
errorwhen those outcomes are routed. - A collected or skipped item failure should leave that item frame
FAILED. The foreach parent/barrier decides whether that failed child is handled and whether the run may continue. skipitem error policy writes no hidden state. Failed frame status and trace/observability are the record; usecollectwhen structured error state is needed.collectwrites structured item error records before emitting an aggregate error outcome. Records include item index, frame id, failing node id, error type/message, and the item value when it can be represented safely.collectdestinations must be state array fields. The foreach barrier writes the full ordered error list once, so users should not expect per-item append writes. Authoring and MCP validation should surface this clearly.completed_with_errorsis the canonical foreach aggregate outcome when all items have finished but one or more item failures were handled bycollectorskip.completed_with_errorsis barrier aggregate vocabulary, not foreach-only. Future explicit barriers may emit it when handled branch failures occurred.- Aggregate control-flow outcomes use result nouns/adjectives such as
done,completed_with_errors, and optionallyfailed. Policy actions use verbs such asfail. - If a foreach policy can emit
completed_with_errors, validation should treat it as a required routable outcome. Users may route it to the same target asdoneexplicitly. collectwrites an empty list to its destination when all items succeed and emitsdone. It emitscompleted_with_errorsonly when at least one item failure was collected.skipemitscompleted_with_errorswhen one or more item failures were skipped. It emitsdoneonly when all items succeed.- Concurrent
failitem policy should stop scheduling new items, drain already started jobs to a quiescent point, capture their results safely, and then fail the run. It should not assume hard cancellation is safe. - After
failtrips, drained sibling results are for trace/observability and cleanup only. They should not commit normal state progress after the failure boundary. - Foreach modes that continue after item failure, such as
collectandskip, should buffer item state patches until the foreach barrier completes. Future concurrent foreach should always use barrier-buffered commits. - Serial
failmay keep immediate commits for compatibility. Commit strategy should be extracted into runtime helpers instead of being smeared through foreach execution code. - Barrier-buffered commits should store pending state patches/results, not full state snapshots. Patch creation and commit must extract/reuse the existing node output validation, output binding, and reducer logic rather than creating a second write system.
- Foreach barrier commits should merge item results in item index order, not completion order. Reducer behavior such as list appends should therefore be deterministic.
- Successful item results may exist as pending barrier results before commit, but they are not visible as workflow state until the barrier commits.
- Pending barrier results must live in resumable
RunState/frame metadata, not only in trace. Trace records history; runtime state is what resume/checkpoint uses. - Concurrent item frames need lineage-local pending state: later nodes in the same item lineage can read earlier pending patches from that item, while sibling items cannot. The exact aggregate output API for committing item results to parent state is deferred.
- Lineage-local state should be represented as patch overlays over parent-visible state, not deep-copied full state snapshots.
- Patch overlays conceptually belong to lineage tokens, not execution frames. A scoped foreach implementation may start by storing them in item/parent metadata, but future Fork/Gather needs first-class lineage ownership.
RunState.stateremains committed parent/global state. Future frame execution should resolve reads against a frame-specific visible state view built from committed state plus visible lineage overlays.- At a barrier, missing reducer means default replace only for single-writer paths. Multiple sibling lineages writing the same path require an explicit reducer; otherwise barrier commit raises a runtime error with writer details.
- Barrier conflict detection should use the same write-overlap rules as normal state writes. Ancestor/descendant writes from different lineages are conflicts unless an explicit merge strategy covers them.
- Reducers at barriers apply incrementally in deterministic lineage order by default: item index order for foreach, declared branch token order for future gather. Completion-order merging is a possible explicit barrier policy, not reducer behavior.
- Collected error records follow the same barrier merge order policy. The default is deterministic lineage order.
- Trace
state_changesshould mean committed state changes only. Do not encode pending barrier patches there unless a future typed trace field is added. - Foreach may keep implicit barrier/iteration state on the foreach parent frame because it owns item spawning and refill. General branch convergence should become an explicit Gather Node later. Shared barrier merge/result helpers should prevent foreach and future barriers from duplicating patch ordering, conflict checks, and failure aggregation.
- Future Fork/Gather needs Lineage Tokens, not raw edge ids. A fork produces
branch lineage tokens; a gather consumes a declared set of tokens and produces
a new merged token. This supports partial gathers such as merging branches
a+bbefore later merging withc. - Explicit Fork/Gather is deferred until lineage tokens are designed. Parallel foreach remains the nearer target because its implicit lineage tokens are item indexes owned by one foreach activation.
- Future foreach metadata should evolve into inherited structured lineage
context. Alias lookup such as
context.documentcan remain authoring sugar, but runtime metadata should preserve nested foreach lineage without flat key collisions. - Nested active foreach aliases should not shadow each other. Alias collisions in inherited context scope should be validation errors.
- Core runtime context should be structural, not alias-first. Foreach lineage
should be addressable through paths such as
context.foreach.<id>.index;wf_authoringcan provide ergonomic alias helpers such ascontext_path(foreach_ref("docs").index). - Python
RuntimeContextshould eventually expose typed structured foreach context, such asctx.foreach["docs"].index, while serialized frame metadata remains JSON-compatible dictionaries. - Structured foreach context keys should be foreach node ids. The
as_alias is authoring sugar for current item access, not the canonical runtime key. wf_authoring.WorkflowBuilder.foreach(...)should eventually return a richer ref exposing context selectors such as.itemand.index, so users do not hand-writecontext.foreach.<id>...paths.- Normal node authors should receive foreach values through mapped input.
Inspecting
RuntimeContext.foreachis an advanced escape hatch for nodes that genuinely need index/frame/lineage context. as_remains useful ergonomic sugar and human-readable trace/docs context, but structured foreach refs should be preferred for non-trivial authoring.- Future foreach policy shape should validate cross-field rules:
collectrequirescollect_to, non-collect actions forbidcollect_to,mode="concurrent"requires a concurrent policy, andmode="serial"forbids a concurrent policy. Deprecated top-levelon_item_errormay parse into the nested item error policy, but canonical dumps should use the nested shape. ForeachConcurrentPolicyshould split limits intomax_activeandmax_outstanding. Defaults aremax_active=4andmax_outstanding=20; validation requiresmax_outstanding >= max_active. Ready or running item frames consume active capacity; blocked item frames consume outstanding capacity but not active capacity.- A blocked non-interrupt item frame frees active capacity and may let foreach
start another item when
max_outstandingalso has room. A run-level interrupt still stops scheduling. - The foreach parent frame owns refill decisions. Scheduler wakes/schedules frames; foreach-specific policy decides whether to start more children, finish, or fail.
- When an item frame becomes blocked, it frees active capacity only for its nearest foreach capacity owner. Future item metadata should identify that owner rather than waking arbitrary ancestors.
- Foreach capacity is local correctness policy, not total process protection.
Future runtime should also have basic global run limits in
wf_core; source- or tool-specific limits belong in the platform layer. - A future global
wf_coreruntime limit should count active node handler calls, not all active frames. Control-flow frames are scheduler work; node calls are the expensive external/user-code execution boundary. - Global node-call limits do not replace foreach caps. Foreach caps bound local scheduling fairness, outstanding frame count, memory, and pending results; global node-call limits bound expensive handler execution across the run.
- Platform source/tool/account limits should be enforced at the node-handler boundary, before invoking the external call. They layer on top of core global node-call admission instead of replacing it.
- Node calls waiting on platform source/tool/account limits still count as active node calls. Core global node-call admission should happen before platform-specific semaphore acquisition.
- Waiting on a source/tool/account semaphore keeps the frame
RUNNING; it is backpressure inside the admitted node-call boundary, not workflow-levelBLOCKEDstate that frees foreach active capacity. - An Interrupt pauses the whole Run, even if the interrupted frame is a child of future concurrent work.
- A Runtime Failure is distinct from a node returning an
erroroutcome. Outcomes are graph control flow; runtime failures are scheduler stops unless a policy handles them. - A node
erroroutcome remains normal graph control flow. A runtime-error utility node or engine exception is what turns control flow into a Runtime Failure. - Missing edges for declared outcomes are validation errors, not normal runtime branch policy.
- Unsupported concurrent foreach semantics should be rejected by validation before runtime. Runtime may stay defensive, but validation owns the user-facing gate.
on_item_error="collect"and"skip"are supported for concurrent foreach. Serial foreach still behaves as fail-only until its execution path explicitly adopts barrier-buffered item error handling.- A Trace records actual scheduler execution order; grouping or sorting by foreach index is a presentation concern.
- Concurrent child frames may write to the same state path only through a Reducer.
- Future concurrent frames should produce State Patches that the scheduler commits atomically.
- A frame's State Visibility excludes uncommitted sibling writes.
- Future concurrent branches require Lineage Isolation: sibling branch writes are invisible unless the graph explicitly joins or merges them.
- A Barrier is the explicit merge boundary for concurrent lineage patches.
- Foreach may own an implicit Barrier; future graph-level convergence may use an explicit barrier node.
Example Dialogue
Dev: "Are we implementing concurrent foreach now?" Domain expert: "No. First we are implementing the Scheduler Foundation so a Run can eventually manage multiple runnable Frames safely."
Flagged Ambiguities
- "concurrent foreach" is the future workflow mode. The resolved current scope is Scheduler Foundation: deterministic internal runtime prep before public concurrent foreach support.
current_frame_id/current_node_idare compatibility fields for the selected cursor, not the source of truth for all runnable work.current_frame_idremains persisted for compatibility, but means "the frame currently selected by the scheduler", not "the only live frame".- When the Ready Queue is empty, the scheduler must explicitly resolve the
run as completed, interrupted, failed, or deadlocked.
current_node_idalone is not a terminal-state rule. - In the first scheduler pass, any Runtime Failure fails the whole run and stops scheduling.
RUNNINGcurrently means "selected cursor during a deterministic step", not "async task already in flight".BLOCKEDshould be a first-class frame status for live frames waiting on child frame completion or external resume.INTERRUPTEDremains a distinct frame status, not justBLOCKEDwith an interrupt reason, because it carries externally visible resume semantics.- Scheduling order should be explicit through a Ready Queue, not inferred from frame dictionary iteration.
mode="concurrent"stays unsupported until Foreach Policy and parent completion semantics are explicit.- A Run has at most one outstanding Interrupt. Resume wakes the frame referenced by that interrupt and scheduling continues from the Ready Queue.
- When a child frame interrupts, only that frame becomes
INTERRUPTED. Ancestors waiting on it remainBLOCKED; the Run status carries the global pause. - Ready sibling frames are preserved during an Interrupt, but the resumed frame is placed at the front of the Ready Queue before scheduling continues.
- Future concurrent execution should not assume in-flight node calls can be safely cancelled. When an Interrupt occurs, scheduling should stop; already started jobs should drain to pending results before resume/commit policy decides what becomes visible.
- Future concurrent interrupts should return control to the caller only at a quiescent pause point: no new work is scheduled, already-started jobs have drained, and their results are captured without unsafe commits.
- Future async execution must protect append-only Trace writes, but should not change trace semantics. Async runtime is an execution capability for simultaneous async node handlers, not a separate foreach workflow mode.
- Scheduler block/wake events should not be added to the public Trace in the first pass. Future observability, including OpenTelemetry-style spans, is a separate concern.
- Frame identifiers remain strings in the first pass, but construction should be centralized so a future structural frame id can replace the string format.
- Typed runtime metadata helpers should return
Nonefor the wrong frame kind, but raise for malformed metadata on the right kind. Corrupt runtime metadata is a runtime invariant failure. - First-pass Block Reason support only needs child-frame blocking; future interrupt, subgraph, and concurrent barrier reasons can extend the same shape.
- Frame creation must reject duplicate frame identifiers. The Ready Queue must contain only existing pending frames and must never contain duplicate frame identifiers.
- Enqueue wakeups may be idempotent: enqueueing an already-ready frame should not duplicate it, and priority enqueue may move it to the front.
- Completed frames should remain in the Frame Set for the lifetime of an in-memory run. Later persistence can add explicit compaction, but scheduler correctness should not depend on deleting completed frames.
- Stack-style frame collapse should be demoted in favor of explicit frame completion and parent wakeup helpers. Parallel scheduling cannot rely on "collapse to parent" semantics.
- The engine/scheduler selects a frame before step preparation.
prepare_stepremains a readability helper for resolving the selected frame's current node, not for choosing runnable work. - Sync and async engine loops should migrate to scheduler selection together so their runtime semantics do not diverge.
- The Ready Queue is part of Run state and should be serialized with
RunStateso resume/checkpoint behavior preserves scheduler order. - A new Run starts with the root frame pending in the Ready Queue. Compatibility cursor fields may still point at root initially.
- First-pass serialized
RunStatechanges should be additive.ready_frame_idsmay be added, but existing run fields and trace shape should not churn. - Existing valid serial workflows should keep the same final output/state, node execution order, foreach item count, interrupt resume behavior, and meaningful trace order after the first scheduler refactor.
- Scheduler foundation is not complete without tests for ready queue selection, duplicate protection, block/wake behavior, deadlock detection, serial foreach regression, and interrupt resume priority.
- Selecting a Runnable Frame removes it from the Ready Queue immediately
and marks it
RUNNING. - A normal non-terminal step advance marks the same frame
PENDINGand re-enqueues it. Terminal, blocked, interrupted, or failed frames are not re-enqueued. - Ready scheduling is FIFO and one-step-at-a-time. A still-runnable frame goes to the back of the Ready Queue after each step.
- Serial foreach blocks the parent on one iteration child at a time, so the child runs until completion/interruption/failure before the parent wakes.
- Future async execution must protect shared state writes. Last-writer-wins is not an acceptable merge policy for concurrent child frames.
- Serial execution may continue mutating state immediately until concurrent execution needs scheduler-controlled State Patch commits.
- Concurrent sibling frames must not depend on observing each other's writes. Use serial foreach or explicit graph structure for ordered dependencies.
- Missing reducers mean replace only within a single serial lineage. At a Barrier, multiple writes to the same path without a reducer are conflicts, not last-writer-wins replacements.
- The pass-through
JoinNodewas removed. Future Gather behavior starts from an explicit barrier contract rather than inheriting placeholder semantics. - The first Scheduler Foundation implementation should create ready/block/wake seams only. Lineage Isolation and Barrier merge behavior are future concurrent semantics, not first-pass behavior.
- First-pass scheduler code should add durable seams such as ready queue, block/wake helpers, and typed foreach metadata. It should not add public concurrent policy, barrier, snapshot, or lineage-patch fields before those semantics are enforced.
- Scheduler rules are shared by sync and async runtime paths; only node handler execution differs.
- The first scheduler helpers should live in a scheduler module, not a concurrency module. Actual async task orchestration can get separate concurrency code later.
- Scheduler helpers are internal runtime infrastructure in the first pass and
should not be exported as public
wf_coreAPI yet.