20 KiB
Native Subgraphs Design
Status: prepared-child execution, saved-artifact resolution, and durable stopped-run interrupt resume implemented
Native subgraphs should make a workflow usable as a workflow step without
collapsing the child run into one opaque Python node call. The current
wf_authoring.subgraph_node and async_subgraph_node helpers are useful
compatibility wrappers, but they hide the child trace, child frames, and child
interrupt lifecycle from wf_core.
This design defines the core runtime shape. The boundary model, prepared-child execution, routed child interrupt resume, and platform loading of immutable saved-child artifact versions are implemented. Durable persisted resume is the remaining run-lifecycle concern and is specified separately.
Goals
- Add a first-class core step for running a child workflow inside a parent workflow.
- Preserve child trace information in a way that can be inspected without pretending child nodes are parent nodes.
- Bubble child interrupts to the parent run and resume back into the child.
- Reuse existing input/output binding, state write, reducer, and schema validation machinery.
- Keep saved workflow / deployment resolution outside
wf_core. - Leave room for future fork/gather and saved-workflow-as-node execution.
Non-Goals
- Do not make arbitrary graph convergence or fork/gather in this pass.
- Do not dynamically load saved artifacts inside
wf_core. - Do not support multiple simultaneous child workflow activations from one subgraph node until the frame identity model is explicit.
- Do not hide async behavior behind sync helpers or
asyncio.run(). - Do not expose child internal state as parent state except through explicit output bindings.
Current Wrapper Problem
The wrapper helpers convert a child workflow into a NodeSpec by calling
execute_workflow or execute_workflow_async from inside a node handler. That
means:
- the parent trace sees one node call
- the child trace is not embedded in the parent run state
- child interrupts cannot bubble cleanly into the parent
- resume cannot re-enter the child workflow
- child workflow identity/version is not part of the core graph
That is acceptable as a temporary compatibility path, but it is not native subgraph execution.
Model Shape
Add a core step model:
class SubgraphNode(BaseModel):
id: str
type: Literal["subgraph"]
workflow: WorkflowRef
input_schema: SchemaRef
output_schema: SchemaRef
input: list[InputBinding] = Field(default_factory=list)
output: list[OutputBinding] = Field(default_factory=list)
outcomes: list[str] = Field(default_factory=lambda: ["ok"])
Current implementation status: the boundary scaffolding is implemented.
wf_core has SubgraphNode; its workflow field is a structural
WorkflowRef: local compiled workflows use {"name": "child"}, while saved
artifacts can use {"artifact_id": "child", "version": 1}. Legacy strings
still parse as input, but saved graphs persist the structural shape. The
placeholder carries input/output schemas and bindings so validation can check
the parent boundary before native execution exists. Core workflows also
declare terminal outcomes through Workflow.outcomes and EndNode.
wf_authoring.subgraph_ref(...) and WorkflowBuilder.subgraph(...) build the
native boundary, while artifact helpers convert saved/capability workflow
references into core WorkflowRef values.
Runtime execution now accepts caller-supplied PreparedSubgraph dependencies
for local workflow refs. It creates a child scope/lineage, schedules child
frames in the parent run, retains child trace entries, and applies mapped
output only at boundary completion. Saved artifact refs are not loaded by
wf_core. Prepared child interrupts bubble through a typed internal route and
resume inside their original child scope.
WorkflowRef should be structural, not a dotted string parser:
class WorkflowRef(BaseModel):
name: str | None = None
artifact_id: str | None = None
version: int | None = None
The reference has two valid forms: local compiled {"name": ...} or saved
artifact {"artifact_id": ..., "version": ...}. It must not derive meaning
from formatted display names. Higher layers may resolve saved artifacts,
deployments, or local builders into an executable child workflow before the
core runtime starts.
SubgraphNode.workflow identifies the child workflow. It does not carry Python
handlers. Handler registries remain runtime dependencies, not graph schema.
Runtime Dependencies
wf_core should execute only already-resolved child workflows. The platform or
authoring layer should prepare a runtime dependency object such as:
SubgraphRuntime(
workflows: Mapping[WorkflowRef, Workflow],
registries: Mapping[WorkflowRef, Mapping[str, NodeHandler]],
reducers: Mapping[WorkflowRef, Mapping[str, ReducerDefinition]],
)
The exact type can evolve, but the boundary matters:
wf_coreowns execution semantics.wf_artifactsowns saved workflow artifact models.wf_mcp/ platform layers own source binding, auth, deployment resolution, and capability availability checks.
Frame Model
A subgraph step creates a child frame tree owned by the parent subgraph frame. The parent frame blocks until the child workflow completes, interrupts, or fails.
Recommended frame metadata:
SubgraphFrameMetadata(
parent_step_id: str,
workflow_ref: WorkflowRef,
child_root_frame_id: str,
child_run_id: str | None = None,
)
The child root frame should have:
kind="subgraph_root"or another typed kindparent_frame_idset to the parent subgraph framenode_idset to the child workflow start node- metadata identifying the child workflow
Child frames created by foreach inside the child workflow remain descendants of the child root, not siblings of the parent graph.
Frame ids should be centrally constructed. The display format may be stringy for now, but logic should not parse frame ids for workflow semantics.
Child State
A child workflow has its own input, state, output, frames, ready queue, and trace semantics.
For v1, the parent RunState can store child runtime data in typed subgraph
metadata rather than a fully nested RunState object. However, the design
should preserve this invariant:
Child workflow state is not parent state.
Parent state changes only happen when the subgraph node completes and applies
its explicit output bindings.
This prevents child internal keys from leaking into the parent and keeps reducer behavior local to the parent output boundary.
Input and Output Mapping
Subgraph input uses the same InputBinding model as NodeUse and
InterruptNode.request:
- read from parent
input,state, orcontext - build the child workflow input payload
- validate against child workflow
input_schema
Subgraph output uses the same OutputBinding model as NodeUse:
- read from child workflow output
- write to parent workflow state
- validate against parent state schema
- apply parent reducers only at the parent write boundary
Child workflow internal reducers are applied only inside child execution.
Trace Shape
Do not flatten child trace entries into parent trace as if they were parent nodes. That loses ownership and makes frame ids misleading.
Recommended trace representation:
- Parent trace gets a
subgraphstep entry for the parent step. - Child trace entries keep their child frame ids and child node ids.
- Each child trace entry should be inspectable through parent run state with structural ownership fields, not a generic metadata bag.
Potential future shape:
TraceEntry(
scope_id="subgraph:run_child",
lineage_id="subgraph:run_child:root",
parent_trace_id="trace:root:run_child",
frame_id="root:child_demo",
node_id="classify",
step_type="node",
...
)
Current TraceEntry has no scope, lineage, or parent-trace fields. The first
implementation can either add explicit optional fields or store child traces in
a separate typed child-trace structure and expose an inspection helper. The
design preference is structural fields or typed containers, not metadata
dictionaries and not overloaded node_id strings.
Interrupt Bubbling
If a child frame reaches an InterruptNode:
- The child frame becomes
INTERRUPTED. - The parent subgraph frame remains
BLOCKED. - The whole parent run status becomes
INTERRUPTED. RunState.interruptpoints to the child interrupt, with enough route data to resume into the child.
The client-facing interrupt identity should remain the parent subgraph
boundary: its frame/step identify the reusable graph node that asked for
interaction, while kind and payload describe what the caller must supply.
The runtime must separately retain the actual child route required for resume:
- parent frame id and subgraph parent step id for the public request identity
- child workflow reference
- child scope and lineage ids
- interrupted child frame id and interrupt node id
- interrupt kind and payload
Current InterruptRequest only has id, frame_id, node_id, kind,
payload, and resumable. Native subgraphs need either:
- explicit structural route fields on
InterruptRequest, such asscope_id,lineage_id,parent_frame_id, andworkflow_ref, or - a typed nested route object, such as
InterruptRoute.
The preferred direction is a typed InterruptRoute stored by
InterruptRequest. A generic metadata field would recreate the ad hoc frame
metadata problem, while string parsing is exactly what the project has been
moving away from. Root-workflow interrupts may omit the route and retain their
existing direct frame/node identity.
Resume Semantics
Resume should target the interrupted child frame, not the parent subgraph step.
On resume:
- Validate that the outstanding interrupt belongs to a live child frame.
- Apply the child interrupt
resumebindings to child state. - Advance the child frame through its resume outcome.
- Put the child frame at the front of the ready queue.
- Continue scheduling.
The parent subgraph frame wakes only when the child workflow reaches a terminal workflow output state.
This mirrors the current scheduler rule: ancestors blocked on child work do not become runnable until the child boundary is actually done.
Completion Semantics
When the child workflow completes:
- Validate child workflow output against child output schema.
- Apply the subgraph step
outputbindings from child output into parent state. - Record a parent
subgraphtrace entry with committed parent state changes. - Advance the parent subgraph frame through the child workflow outcome.
Core workflows now declare Workflow.outcomes, and explicit EndNode steps
set RunState.outcome. The legacy __end__ token remains compatibility
shorthand for workflow outcome ok. Native subgraph execution should use that
workflow-level outcome as the parent-visible subgraph outcome, instead of
guessing from the child node that happened to route to a terminal.
Failure Semantics
Runtime failure inside a child workflow fails the parent run unless future policy explicitly handles child failures.
Do not turn child runtime failures into normal graph outcomes by default. Normal outcomes are graph control flow; runtime failures are execution failures.
If a child node returns an error outcome and the child graph routes it, that is
ordinary child workflow behavior. If the child graph reaches a runtime error,
that is a run failure.
Validation
validate_workflow should add subgraph checks:
- subgraph step has a resolvable child workflow reference in the runtime environment or validation context
- subgraph input bindings target child input paths
- subgraph output bindings source child output paths and target parent state paths
outcomesare declared and outgoing edges match them- native subgraphs cannot be recursive unless explicit cycle detection exists
- child workflow structural validation runs before parent execution
Pure model validation should still avoid executing or resolving external artifacts. Runtime/deployment validation can perform stronger checks with a resolved dependency set.
Authoring Layer
wf_authoring should expose native subgraph use separately from wrapper-node
composition.
Current helper:
child = parent.subgraph(
id="run_child",
workflow=child_builder.compile(),
input=[input_from(state_path("request"), "request")],
output=[output_to("summary", state_path("child_summary"))],
)
This copies the compiled child workflow contract into a core SubgraphNode,
appends it to the builder, and returns the step for normal routing. For a local
child builder, parent.prepare_subgraph(child_builder) registers the compiled
graph, handlers, and reducers required by parent.execute(...) and
parent.resume(...). Higher layers still need dependency resolution before
saved/deployed workflow refs can run. The lower-level subgraph_ref(...)
helper exists for code that wants only the core step object.
Possible API:
child = parent.subgraph(
workflow=child_builder.compile(),
id="run_child",
input=[input_from(state_path("request"), "request")],
output=[output_to("summary", state_path("child_summary"))],
)
parent.connect(child, "ok", END)
For saved artifacts, use the lower-level helper with a structural core ref:
child = subgraph_ref(
id="run_child",
workflow=child_builder.compile(),
workflow_ref=WorkflowRef(artifact_id="demo_child", version=1),
...
)
subgraph_node and async_subgraph_node should remain compatibility helpers
until native subgraphs cover the same use cases. They should keep warning in
docs that they are wrapper nodes.
MCP and Artifact Layer
Saved workflows should be reusable through the same native subgraph boundary,
but wf_core should not know how to load them.
The platform layer should:
- resolve workflow artifact refs to concrete workflows
- resolve capability/source bindings for that artifact
- provide node registries and reducers for the child workflow
- validate dependency availability before run
- expose clear diagnostics when a child workflow is unrunnable
This keeps auth, source availability, deployment binding, and MCP account
selection out of wf_core.
Saved Child Deployment Environment
For the first saved-subgraph slice, one deployment defines the runnable
environment for the full graph tree. A parent deployment that loads a
structural child ref such as {"artifact_id": "child", "version": 2} must:
- load exactly that immutable child artifact version
- resolve the child's logical node and reducer dependencies using the same deployment bindings as the parent
- recursively prepare non-interrupting descendant child artifacts using that same environment
- detect missing artifacts, missing bound capabilities, and saved-child reference cycles before execution
This intentionally does not add a child deployment id to WorkflowRef.
Supporting an explicit per-child deployment override later remains compatible
with the structural child reference and preparation boundary. Such an override
must be keyed by the parent subgraph dependency/use site, not only by child
artifact id: one parent graph may intentionally invoke the same immutable
child artifact twice against different accounts or capability bindings.
Saved Child Interrupt Status
Native prepared children can interrupt and resume in core. The workflow surface
now exposes durable run_deployment / resume_run support for saved
interrupting artifacts, including interrupts raised by saved descendants.
Stopped checkpoints pin the deployment and exact root/child artifact
definitions; dependency drift can block resume without mutating the stopped
execution. Durable run persistence is specified in
2026-05-26-durable-workflow-runs-and-resume-design.md.
Implementation Slices
Completed Scaffold: Typed Native Boundary
SubgraphNodeis part of the coreStepunion and validates its declared parent-side boundary.WorkflowRefis structural and supports local compiled or saved artifact references without requiring runtime string parsing.Workflow.outcomes,EndNode, andRunState.outcomedefine child terminal outcome semantics before child execution exists.subgraph_ref(...)andWorkflowBuilder.subgraph(...)produce native boundaries; wrapper-node helpers remain compatibility APIs.- Artifact conversion helpers bridge saved workflow identities to core
WorkflowRefvalues.
Completed Slice 1: Prepared Subgraph Runtime
- Local/prepared child
WorkflowRefdependencies resolve throughPreparedSubgraph;wf_coredoes not load saved artifacts. - Child workflows execute through child frames in the parent scheduler.
- Each activation owns a child runtime scope/lineage so child state is isolated from parent state until boundary completion.
- Child trace entries remain in the parent run with child frame ids; completion
records the parent
subgraphtrace entry. - Child output maps to parent state through existing output binding machinery.
- The parent step routes through the child's terminal workflow outcome.
- Child interrupts route structurally and preserve the blocked parent boundary until the child resumes and completes.
Completed Slice 2: Interrupt Bubbling and Resume
- Extend
InterruptRequestwith explicit route structure. - Bubble child interrupts to the parent run.
- Resume into the child frame.
- Tests: child interrupt pauses parent, resume continues child, parent completes, wrong resume target fails clearly.
Completed Slice 3: Saved Workflow References
- Structural saved-workflow references and conversion helpers identify exact saved child versions without runtime string parsing.
- Platform-level resolution prepares saved workflow artifacts before runtime.
- The parent deployment binding environment applies transitively to each exact saved child artifact version.
- Dependencies, missing artifacts, and saved-child cycles validate before execution.
- Tests cover saved child execution through deployment bindings, nested saved child dependencies, missing/cyclic diagnostics, and durable interrupt/resume for saved children.
Slice 4: Optional Policy Expansion
- Workflow outcome propagation is settled: child
RunState.outcomeis the parent-visible subgraph outcome; legacy__end__meansok, while explicitEndNodecarries other declared outcomes. - Keep child runtime failures as parent runtime failures by default.
- Only add configurable child-failure policy or richer boundary result semantics when an actual use case requires it.
Risks
- Trace shape can become confusing if child entries are flattened too early.
- Storing nested run state directly may bloat persisted runs unless inspection APIs paginate trace/state detail.
- Multiple child workflow dependency registries can make runtime dependencies complex; keep the boundary typed early.
Open Questions
- Should child runtime state be stored as a nested
RunState, or as typed child frame metadata plus shared parentRunState.frames? - Should
TraceEntrygain explicitscope_id,lineage_id, and parent-trace fields, or should child traces live in a separate inspectable structure?
Recommendation
The typed boundary scaffold, prepared-child runtime, saved-child resolution, and durable routed interrupt resume are complete. Do not delete wrapper-node helpers yet; they remain compatibility APIs while saved native execution matures. Next policy work is optional per-use-site child deployment overrides and richer protocol-native long-running progress/reporting.