Files
lda-wf/docs/deployment_scheduling.md
T

207 lines
8.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Deployment Scheduling Operations
Schedules start ordinary deployment runs without a connected client. The
scheduler is opt-in, lives in the workflow server, and introduces no
workflow node type. It is not the core runtime's frame scheduler.
Current contract:
[`deployment scheduling spec`](superpowers/specs/2026-09-08-deployment-scheduling-design.md).
## Enablement
Scheduling is disabled by default. Enable it for a local/static server
with the config section or the CLI flag (MCP-backed servers reject it):
```json
{"server": {"scheduler": {"enabled": true}}}
```
```bash
wf-rpc-server --store-root .wf_store --enable-scheduler
```
Tuning (`server.scheduler`): `poll_interval_s` (default 1.0),
`max_concurrent_runs` (default 4, the execution-slot bound),
`drain_grace_s` (default 30.0). Schedule data lives at
`<store_root>/schedules` next to run data; one lock file at
`<store_root>/scheduler.lock` proves exclusive ownership. A second
scheduler over the same stores is rejected; shut the first down before
starting another.
## Mental model: schedule vs occurrence vs run
- A **schedule** is a durable definition: deployment, trigger
(one-shot or cron), input bindings, overlap/misfire policies, and a
revision. Edits bump the revision and affect only future admissions.
- An **occurrence** is one resolved calendar instant, identified by
`(schedule_id, resolved UTC instant)`. An occurrence is immutable once
admitted and is never replayed.
- A **run** is the execution of one admitted occurrence, with a
store-backed `run-000001`-style id, a pinned input/artifact snapshot,
and a stopped checkpoint when it stops.
## Triggers and time zones
One-shot timestamps must include an offset. Cron uses five-field Unix
expressions (`0` and `7` both mean Sunday; day-of-month/day-of-week
match with OR) plus an explicit IANA time zone, default UTC.
Occurrence instants persist as UTC; the definition keeps the zone name.
`croniter` owns calendar resolution including DST gaps and folds; there
is no custom calendar filtering. Reject invalid zones and naive times.
## Overlap
Overlap is per schedule, not per deployment. With default
`overlap="skip"`, an unfinished scheduled run — including an
interrupted run awaiting input — blocks another occurrence from that
schedule. Manual runs and other schedules never participate.
`overlap="parallel"` admits independent runs up to a required positive
`max_active_runs`; admitted, running, and interrupted runs all count,
and lowering the limit blocks new admission without cancelling existing
runs. A known completed or failed run releases overlap.
## Misfire
Default `misfire="skip"` drops missed starts (past the 60-second
configurable lateness allowance). `misfire="latest"` retains at most
one latest unadmitted candidate per schedule and runs it as soon as
overlap and capacity allow; it never replays a burst. Creation never
backfills time before the revision: the consumed watermark starts at
creation.
## Pause
Pausing stops future admission, not an active run. Resume selects the
next future occurrence; paused times are not replayed. Pause is not
downtime: the paused span is consumed, so downtime catch-up never
resurrects it.
## Deletion
Deleting a schedule stops future admission. Existing runs and
occurrence history survive, including their schedule identity; active
runs continue and may still complete or resume against the retained
history. Schedule ids are never reused.
## Failure and restart
An abandoned in-flight execution (the process died mid-run) becomes
failed without retrying its occurrence; the failure discloses that
external effects may already have occurred. A durably interrupted run
is not abandoned: it stays resumable and keeps occupying its
schedule's overlap slot across restarts. Every resume marks a
store-backed attempt first, so recovery can tell a fresh result from a
stale checkpoint and fail the ambiguous case closed instead of
retrying it. Corrupt or contradictory records fail closed with
diagnostics and block the schedule rather than clearing overlap.
On shutdown the server stops admission first and drains active tasks
within the grace period; new scheduled resumes are rejected for the
duration of the drain, and a resume still running past the deadline is
cancelled and joined before ownership is released, keeping its marks
for the next startup recovery. Anything else still running keeps its
executing mark, and the next startup recovery abandons it truthfully.
## Occurrence inspection
`list_schedule_occurrences` pages stored history (`pending`,
`coalesced`/`superseded`, `skipped-overlap`, `skipped-misfire`,
`preflight-rejected`, `admitted`/`running`, `interrupted`,
`completed`, `failed`, `exhausted`) oldest-first with a `next_cursor`.
A currently held (unadmitted) candidate is synthesized as a `pending`
row at the top of the first page; once admitted, the durable
`admitted` entry replaces it. While a candidate is held, the first
page still honors `limit`; with `limit=1`, its continuation cursor is an
opaque marker that resumes the stored history from its first row.
## Administration surface
Python API (`server.api.schedules`), JSON-RPC (`workflow.schedules.*`),
and the Python client (`App` schedule methods) share these operations:
- `create_schedule` (`workflow.schedules.create`): caller-chosen id;
trigger, deployment, binding, and sample-schema checks.
- `get_schedule` (`workflow.schedules.get`): full definition payload.
- `list_schedules` (`workflow.schedules.list`): deleted excluded
unless asked.
- `update_schedule` (`workflow.schedules.update`): `expected_revision`
required; provided fields only; `None` means unpatched.
- `pause_schedule` (`workflow.schedules.pause`): idempotent; clears
candidates, consumes the span.
- `resume_schedule` (`workflow.schedules.resume`): idempotent; resumes
from the next future instant.
- `delete_schedule` (`workflow.schedules.delete`): soft delete; runs
and history survive.
- `list_schedule_occurrences`
(`workflow.schedules.occurrences.list`): cursor pages, `limit` 1100,
live pending synthesis.
`inspect_run` also reads admitted (checkpoint-less) runs: status
`admitted` with no fabricated trace, output, or checkpoint.
## Hypothetically used as follows
EXECUTABLE example (runs in CI as
`tests/examples/test_scheduled_deployment_example.py`; run it with
`uv run pytest -q tests/examples/test_scheduled_deployment_example.py`):
```python
server = build_local_static_workflow_server(root, schedules=True)
await server.api.create_artifact_from_plan(
artifact_id="scheduled_hello", version=1, ...,
)
await server.api.save_deployment({...})
due = datetime.now(UTC) + timedelta(seconds=0.5)
await server.api.schedules.create_schedule(
schedule_id="hello-once",
deployment_id="scheduled_hello.default",
trigger={"kind": "oneshot", "at": due.isoformat()},
)
service = build_scheduler_service(server, SchedulerServiceConfig(...))
await service.start()
# ... the service admits the occurrence and completes the run ...
inspected = await server.api.inspect_run(run_id=run_id)
assert inspected["output"]["result"] == "hello on a schedule"
page = await server.api.schedules.list_schedule_occurrences(
schedule_id="hello-once"
)
await service.stop()
```
ILLUSTRATIVE example (not executed; shows a cron week with an operator
pause — same calls, longer horizons):
```python
# Monday: an hourly report, latest-wins catch-up, at most two at once.
await schedules.create_schedule(
schedule_id="hourly-report",
deployment_id="report.default",
trigger={"kind": "cron", "expression": "0 * * * *", "timezone": "UTC"},
misfire="latest",
overlap="parallel",
max_active_runs=2,
)
# Friday: pause for maintenance; Monday: resume from the next hour.
await schedules.pause_schedule(schedule_id="hourly-report")
await schedules.resume_schedule(schedule_id="hourly-report")
# A bad edit is rejected without touching the running definition:
await schedules.update_schedule(
schedule_id="hourly-report", expected_revision=1, ...
)
```
## Known limitations
- One composition's stores nested inside another live composition's
store subtree (without sharing its identical roots) is unsupported
operator error; shared-store cross layouts are rejected outright.
- Manual runs bypass scheduler capacity by design; capacity governs
scheduled dispatch plus scheduled resumes. Resuming a scheduled
interrupted run acquires a server execution slot first (rejected
while saturated or draining, without dispatching) and releases it when the
resumed result is persisted; a drain-cancelled resume keeps its marks
for recovery instead.
- A set `max_steps` budget cannot be cleared back to unset through
update (recreate the schedule for an unbounded budget).
- MCP-backed servers reject scheduler enablement for now.