210 lines
8.9 KiB
Markdown
210 lines
8.9 KiB
Markdown
# Deployment Scheduling Operations
|
||
|
||
Schedules start ordinary deployment runs without a connected client. The
|
||
scheduler is opt-in, lives in the workflow server, and introduces no
|
||
workflow node type. It is not the core runtime's frame scheduler.
|
||
|
||
Current contract:
|
||
[`deployment scheduling spec`](superpowers/specs/2026-09-08-deployment-scheduling-design.md).
|
||
|
||
## Enablement
|
||
|
||
Scheduling is disabled by default. Enable it for a local/static server
|
||
with the config section or the CLI flag (MCP-backed servers reject it):
|
||
|
||
```json
|
||
{"server": {"scheduler": {"enabled": true}}}
|
||
```
|
||
|
||
```bash
|
||
wf-rpc-server --store-root .wf_store --enable-scheduler
|
||
```
|
||
|
||
Tuning (`server.scheduler`): `poll_interval_s` (default 1.0),
|
||
`max_concurrent_runs` (default 4, the execution-slot bound),
|
||
`drain_grace_s` (default 30.0). Schedule data lives at
|
||
`<store_root>/schedules` next to run data; one lock file at
|
||
`<store_root>/scheduler.lock` proves exclusive ownership. A second
|
||
scheduler over the same stores is rejected; shut the first down before
|
||
starting another.
|
||
|
||
## Mental model: schedule vs occurrence vs run
|
||
|
||
- A **schedule** is a durable definition: deployment, trigger
|
||
(one-shot or cron), input bindings, overlap/misfire policies, and a
|
||
revision. Edits bump the revision and affect only future admissions.
|
||
- An **occurrence** is one resolved calendar instant, identified by
|
||
`(schedule_id, resolved UTC instant)`. An occurrence is immutable once
|
||
admitted and is never replayed.
|
||
- A **run** is the execution of one admitted occurrence, with a
|
||
store-backed `run-000001`-style id, a pinned input/artifact snapshot,
|
||
and a stopped checkpoint when it stops.
|
||
|
||
## Triggers and time zones
|
||
|
||
One-shot timestamps must include an offset. Cron uses five-field Unix
|
||
expressions (`0` and `7` both mean Sunday; day-of-month/day-of-week
|
||
match with OR) plus an explicit IANA time zone, default UTC.
|
||
Occurrence instants persist as UTC; the definition keeps the zone name.
|
||
`croniter` owns calendar resolution including DST gaps and folds; there
|
||
is no custom calendar filtering. Reject invalid zones and naive times.
|
||
|
||
## Overlap
|
||
|
||
Overlap is per schedule, not per deployment. With default
|
||
`overlap="skip"`, an unfinished scheduled run — including an
|
||
interrupted run awaiting input — blocks another occurrence from that
|
||
schedule. Manual runs and other schedules never participate.
|
||
`overlap="parallel"` admits independent runs up to a required positive
|
||
`max_active_runs`; admitted, running, and interrupted runs all count,
|
||
and lowering the limit blocks new admission without cancelling existing
|
||
runs. A known completed or failed run releases overlap.
|
||
|
||
## Misfire
|
||
|
||
Default `misfire="skip"` drops missed starts (past the 60-second
|
||
configurable lateness allowance). `misfire="latest"` retains at most
|
||
one latest unadmitted candidate per schedule and runs it as soon as
|
||
overlap and capacity allow; it never replays a burst. Creation never
|
||
backfills time before the revision: the consumed watermark starts at
|
||
creation.
|
||
|
||
## Pause
|
||
|
||
Pausing stops future admission, not an active run. Resume selects the
|
||
next future occurrence; paused times are not replayed. Pause is not
|
||
downtime: the paused span is consumed, so downtime catch-up never
|
||
resurrects it.
|
||
|
||
## Deletion
|
||
|
||
Deleting a schedule stops future admission. Existing runs and
|
||
occurrence history survive, including their schedule identity; active
|
||
runs continue and may still complete or resume against the retained
|
||
history. Schedule ids are never reused.
|
||
|
||
## Failure and restart
|
||
|
||
An abandoned in-flight execution (the process died mid-run) becomes
|
||
failed without retrying its occurrence; the failure discloses that
|
||
external effects may already have occurred. A durably interrupted run
|
||
is not abandoned: it stays resumable and keeps occupying its
|
||
schedule's overlap slot across restarts. Every resume marks a
|
||
store-backed attempt first, so recovery can tell a fresh result from a
|
||
stale checkpoint and fail the ambiguous case closed instead of
|
||
retrying it. Corrupt or contradictory records fail closed with
|
||
diagnostics and block the schedule rather than clearing overlap.
|
||
|
||
On shutdown the server stops admission first and drains active tasks
|
||
within the grace period; new scheduled resumes are rejected for the
|
||
duration of the drain, and a resume still running past the deadline is
|
||
cancelled and joined before ownership is released, keeping its marks
|
||
for the next startup recovery. Anything else still running keeps its
|
||
executing mark, and the next startup recovery abandons it truthfully.
|
||
|
||
## Occurrence inspection
|
||
|
||
`list_schedule_occurrences` pages stored history (`pending`,
|
||
`coalesced`/`superseded`, `skipped-overlap`, `skipped-misfire`,
|
||
`preflight-rejected`, `admitted`/`running`, `interrupted`,
|
||
`completed`, `failed`, `exhausted`) oldest-first with a `next_cursor`.
|
||
A currently held (unadmitted) candidate is synthesized as a `pending`
|
||
row at the top of the first page; once admitted, the durable
|
||
`admitted` entry replaces it. While a candidate is held, the first
|
||
page still honors `limit`; with `limit=1`, its continuation cursor is an
|
||
opaque marker that resumes the stored history from its first row.
|
||
|
||
## Administration surface
|
||
|
||
Python API (`server.api.schedules`), JSON-RPC (`workflow.schedules.*`),
|
||
and the Python client (`App` schedule methods) share these operations:
|
||
|
||
- `create_schedule` (`workflow.schedules.create`): caller-chosen id;
|
||
trigger, deployment, binding, and sample-schema checks.
|
||
- `get_schedule` (`workflow.schedules.get`): full definition payload.
|
||
- `list_schedules` (`workflow.schedules.list`): deleted excluded
|
||
unless asked.
|
||
- `update_schedule` (`workflow.schedules.update`): `expected_revision`
|
||
required; provided fields only; `None` means unpatched.
|
||
- `pause_schedule` (`workflow.schedules.pause`): idempotent; clears
|
||
candidates, consumes the span.
|
||
- `resume_schedule` (`workflow.schedules.resume`): idempotent; resumes
|
||
from the next future instant.
|
||
- `delete_schedule` (`workflow.schedules.delete`): soft delete; runs
|
||
and history survive.
|
||
- `list_schedule_occurrences`
|
||
(`workflow.schedules.occurrences.list`): cursor pages, `limit` 1–100,
|
||
live pending synthesis.
|
||
|
||
`inspect_run` also reads admitted (checkpoint-less) runs: status
|
||
`admitted` with no fabricated trace, output, or checkpoint.
|
||
|
||
## Hypothetically used as follows
|
||
|
||
The complete runnable setup is covered by
|
||
[`tests/examples/test_scheduled_deployment_example.py`](../tests/examples/test_scheduled_deployment_example.py)
|
||
(run it with `uv run pytest -q tests/examples/test_scheduled_deployment_example.py`).
|
||
The abbreviated sequence below is illustrative rather than executable as
|
||
written because it elides setup details:
|
||
|
||
ILLUSTRATIVE example (the runnable test above contains the complete setup):
|
||
|
||
```python
|
||
server = build_local_static_workflow_server(root, schedules=True)
|
||
await server.api.create_artifact_from_plan(
|
||
artifact_id="scheduled_hello", version=1, ...,
|
||
)
|
||
await server.api.save_deployment({...})
|
||
due = datetime.now(UTC) + timedelta(seconds=0.5)
|
||
await server.api.schedules.create_schedule(
|
||
schedule_id="hello-once",
|
||
deployment_id="scheduled_hello.default",
|
||
trigger={"kind": "oneshot", "at": due.isoformat()},
|
||
)
|
||
service = build_scheduler_service(server, SchedulerServiceConfig(...))
|
||
await service.start()
|
||
# ... the service admits the occurrence and completes the run ...
|
||
inspected = await server.api.inspect_run(run_id=run_id)
|
||
assert inspected["output"]["result"] == "hello on a schedule"
|
||
page = await server.api.schedules.list_schedule_occurrences(
|
||
schedule_id="hello-once"
|
||
)
|
||
await service.stop()
|
||
```
|
||
|
||
Longer-horizon operator example (also illustrative and not executed):
|
||
|
||
```python
|
||
# Monday: an hourly report, latest-wins catch-up, at most two at once.
|
||
await schedules.create_schedule(
|
||
schedule_id="hourly-report",
|
||
deployment_id="report.default",
|
||
trigger={"kind": "cron", "expression": "0 * * * *", "timezone": "UTC"},
|
||
misfire="latest",
|
||
overlap="parallel",
|
||
max_active_runs=2,
|
||
)
|
||
# Friday: pause for maintenance; Monday: resume from the next hour.
|
||
await schedules.pause_schedule(schedule_id="hourly-report")
|
||
await schedules.resume_schedule(schedule_id="hourly-report")
|
||
# A bad edit is rejected without touching the running definition:
|
||
await schedules.update_schedule(
|
||
schedule_id="hourly-report", expected_revision=1, ...
|
||
)
|
||
```
|
||
|
||
## Known limitations
|
||
|
||
- One composition's stores nested inside another live composition's
|
||
store subtree (without sharing its identical roots) is unsupported
|
||
operator error; shared-store cross layouts are rejected outright.
|
||
- Manual runs bypass scheduler capacity by design; capacity governs
|
||
scheduled dispatch plus scheduled resumes. Resuming a scheduled
|
||
interrupted run acquires a server execution slot first (rejected
|
||
while saturated or draining, without dispatching) and releases it when the
|
||
resumed result is persisted; a drain-cancelled resume keeps its marks
|
||
for recovery instead.
|
||
- A set `max_steps` budget cannot be cleared back to unset through
|
||
update (recreate the schedule for an unbounded budget).
|
||
- MCP-backed servers reject scheduler enablement for now.
|