Files
lda-wf/docs/deployment_scheduling.md
T

8.9 KiB
Raw Permalink Blame History

Deployment Scheduling Operations

Schedules start ordinary deployment runs without a connected client. The scheduler is opt-in, lives in the workflow server, and introduces no workflow node type. It is not the core runtime's frame scheduler.

Current contract: deployment scheduling spec.

Enablement

Scheduling is disabled by default. Enable it for a local/static server with the config section or the CLI flag (MCP-backed servers reject it):

{"server": {"scheduler": {"enabled": true}}}
wf-rpc-server --store-root .wf_store --enable-scheduler

Tuning (server.scheduler): poll_interval_s (default 1.0), max_concurrent_runs (default 4, the execution-slot bound), drain_grace_s (default 30.0). Schedule data lives at <store_root>/schedules next to run data; one lock file at <store_root>/scheduler.lock proves exclusive ownership. A second scheduler over the same stores is rejected; shut the first down before starting another.

Mental model: schedule vs occurrence vs run

  • A schedule is a durable definition: deployment, trigger (one-shot or cron), input bindings, overlap/misfire policies, and a revision. Edits bump the revision and affect only future admissions.
  • An occurrence is one resolved calendar instant, identified by (schedule_id, resolved UTC instant). An occurrence is immutable once admitted and is never replayed.
  • A run is the execution of one admitted occurrence, with a store-backed run-000001-style id, a pinned input/artifact snapshot, and a stopped checkpoint when it stops.

Triggers and time zones

One-shot timestamps must include an offset. Cron uses five-field Unix expressions (0 and 7 both mean Sunday; day-of-month/day-of-week match with OR) plus an explicit IANA time zone, default UTC. Occurrence instants persist as UTC; the definition keeps the zone name. croniter owns calendar resolution including DST gaps and folds; there is no custom calendar filtering. Reject invalid zones and naive times.

Overlap

Overlap is per schedule, not per deployment. With default overlap="skip", an unfinished scheduled run — including an interrupted run awaiting input — blocks another occurrence from that schedule. Manual runs and other schedules never participate. overlap="parallel" admits independent runs up to a required positive max_active_runs; admitted, running, and interrupted runs all count, and lowering the limit blocks new admission without cancelling existing runs. A known completed or failed run releases overlap.

Misfire

Default misfire="skip" drops missed starts (past the 60-second configurable lateness allowance). misfire="latest" retains at most one latest unadmitted candidate per schedule and runs it as soon as overlap and capacity allow; it never replays a burst. Creation never backfills time before the revision: the consumed watermark starts at creation.

Pause

Pausing stops future admission, not an active run. Resume selects the next future occurrence; paused times are not replayed. Pause is not downtime: the paused span is consumed, so downtime catch-up never resurrects it.

Deletion

Deleting a schedule stops future admission. Existing runs and occurrence history survive, including their schedule identity; active runs continue and may still complete or resume against the retained history. Schedule ids are never reused.

Failure and restart

An abandoned in-flight execution (the process died mid-run) becomes failed without retrying its occurrence; the failure discloses that external effects may already have occurred. A durably interrupted run is not abandoned: it stays resumable and keeps occupying its schedule's overlap slot across restarts. Every resume marks a store-backed attempt first, so recovery can tell a fresh result from a stale checkpoint and fail the ambiguous case closed instead of retrying it. Corrupt or contradictory records fail closed with diagnostics and block the schedule rather than clearing overlap.

On shutdown the server stops admission first and drains active tasks within the grace period; new scheduled resumes are rejected for the duration of the drain, and a resume still running past the deadline is cancelled and joined before ownership is released, keeping its marks for the next startup recovery. Anything else still running keeps its executing mark, and the next startup recovery abandons it truthfully.

Occurrence inspection

list_schedule_occurrences pages stored history (pending, coalesced/superseded, skipped-overlap, skipped-misfire, preflight-rejected, admitted/running, interrupted, completed, failed, exhausted) oldest-first with a next_cursor. A currently held (unadmitted) candidate is synthesized as a pending row at the top of the first page; once admitted, the durable admitted entry replaces it. While a candidate is held, the first page still honors limit; with limit=1, its continuation cursor is an opaque marker that resumes the stored history from its first row.

Administration surface

Python API (server.api.schedules), JSON-RPC (workflow.schedules.*), and the Python client (App schedule methods) share these operations:

  • create_schedule (workflow.schedules.create): caller-chosen id; trigger, deployment, binding, and sample-schema checks.
  • get_schedule (workflow.schedules.get): full definition payload.
  • list_schedules (workflow.schedules.list): deleted excluded unless asked.
  • update_schedule (workflow.schedules.update): expected_revision required; provided fields only; None means unpatched.
  • pause_schedule (workflow.schedules.pause): idempotent; clears candidates, consumes the span.
  • resume_schedule (workflow.schedules.resume): idempotent; resumes from the next future instant.
  • delete_schedule (workflow.schedules.delete): soft delete; runs and history survive.
  • list_schedule_occurrences (workflow.schedules.occurrences.list): cursor pages, limit 1100, live pending synthesis.

inspect_run also reads admitted (checkpoint-less) runs: status admitted with no fabricated trace, output, or checkpoint.

Hypothetically used as follows

The complete runnable setup is covered by tests/examples/test_scheduled_deployment_example.py (run it with uv run pytest -q tests/examples/test_scheduled_deployment_example.py). The abbreviated sequence below is illustrative rather than executable as written because it elides setup details:

ILLUSTRATIVE example (the runnable test above contains the complete setup):

server = build_local_static_workflow_server(root, schedules=True)
await server.api.create_artifact_from_plan(
    artifact_id="scheduled_hello", version=1, ...,
)
await server.api.save_deployment({...})
due = datetime.now(UTC) + timedelta(seconds=0.5)
await server.api.schedules.create_schedule(
    schedule_id="hello-once",
    deployment_id="scheduled_hello.default",
    trigger={"kind": "oneshot", "at": due.isoformat()},
)
service = build_scheduler_service(server, SchedulerServiceConfig(...))
await service.start()
# ... the service admits the occurrence and completes the run ...
inspected = await server.api.inspect_run(run_id=run_id)
assert inspected["output"]["result"] == "hello on a schedule"
page = await server.api.schedules.list_schedule_occurrences(
    schedule_id="hello-once"
)
await service.stop()

Longer-horizon operator example (also illustrative and not executed):

# Monday: an hourly report, latest-wins catch-up, at most two at once.
await schedules.create_schedule(
    schedule_id="hourly-report",
    deployment_id="report.default",
    trigger={"kind": "cron", "expression": "0 * * * *", "timezone": "UTC"},
    misfire="latest",
    overlap="parallel",
    max_active_runs=2,
)
# Friday: pause for maintenance; Monday: resume from the next hour.
await schedules.pause_schedule(schedule_id="hourly-report")
await schedules.resume_schedule(schedule_id="hourly-report")
# A bad edit is rejected without touching the running definition:
await schedules.update_schedule(
    schedule_id="hourly-report", expected_revision=1, ...
)

Known limitations

  • One composition's stores nested inside another live composition's store subtree (without sharing its identical roots) is unsupported operator error; shared-store cross layouts are rejected outright.
  • Manual runs bypass scheduler capacity by design; capacity governs scheduled dispatch plus scheduled resumes. Resuming a scheduled interrupted run acquires a server execution slot first (rejected while saturated or draining, without dispatching) and releases it when the resumed result is persisted; a drain-cancelled resume keeps its marks for recovery instead.
  • A set max_steps budget cannot be cleared back to unset through update (recreate the schedule for an unbounded budget).
  • MCP-backed servers reject scheduler enablement for now.