Files
lda-wf/docs/historical/superpowers/plans/2026-09-09-deployment-scheduling-implementation-plan.md
T

35 KiB
Raw Blame History

Deployment Scheduling: Verification Report And Sequenced Implementation Plan

Status: verification complete; planning only. No production scheduling code exists on this branch. No blocking product decisions remain (the calendar policy is decided: croniter owns DST resolution — see Gate 2; the former custom skip/filter approach is retired, not pending).

  • Branch: opencode/sched-verify-plan
  • Worktree: C:\Users\Admin\Documents\lda.chat\lda-workflow-as-struct-sched-verify
  • Spec under review: docs/superpowers/specs/2026-09-08-deployment-scheduling-design.md (updated in this worktree; the original untracked main-worktree copy is unchanged and is no longer identical)
  • Probes (disposable, not production): probes/deployment_scheduling_verify/

Scheduling policies remain settled. Calendar behavior follows the subsequent approved simplification: croniter owns DST resolution. Findings below record the missing implementation seams and the completed verification work.

Completion status (2026-09-09; plan retired to historical/)

All gates and tasks completed on branch opencode/sched-verify-plan:

  • R0, R1, R2, R3: passed during phased implementation.
  • R4 (fault-injection review): passed in three waves — wave 2 bound ownership to store composition and stabilized recovery failure; wave 3 added one canonical lock identity plus checkpoint coherence and ordering authority; wave 4 closed the shared-store overlap hole (sibling-distinct identity, overlapping layouts rejected).
  • T01T11: implemented per phase (calendar, expressions, admission, scheduler core, resume safety, recovery, ownership).
  • T12: opt-in server lifecycle with bounded real execution, drain, and subprocess-tested death paths; independent review passed.
  • T13: administration API + JSON-RPC surface + Python client; an independent review gated it on torn-admin ordering and admin/poll staleness, both fixed (crash-safe admin ordering, creation watermark, per-schedule poll freshness) and re-reviewed to pass.
  • T14: this completion record; spec updated in place; roadmap and user docs updated; disposable probes deleted after production-test equivalence was verified item by item (calendar Part A mirrored in test_calendar_adapter.py; APScheduler rejected-candidate evidence preserved in the design spec; expression pins mirrored or superseded by T03/T04 plus core tests; all 31 state-model behaviors mapped to production tests, adding overlap-independence and store-backed run-identity pins where no equivalent existed).

Probe retirement map: probes/deployment_scheduling_verify/ deleted. test_calendar_probe.py Part A is subsumed by tests/scheduling/test_calendar_adapter.py; Part B (APScheduler gap phantom + fold replay strict xfails) is preserved as narrative evidence in the design spec, not as runnable tests (it needs an isolated env with apscheduler installed). test_expression_contract_probe.py is subsumed by tests/scheduling/test_schedule_expressions.py, tests/core/test_input_sources.py, and core strict-JSON tests, except the pre-T03 no-seam observation (deliberately superseded) and one unowned InputBinding micro-pin of untouched core code (noted, not ported). test_schedule_state_model.py is subsumed by tests/scheduling/ (matrix, coalescing, slots, pause/edit/delete, restart, downtime, fairness, fault injection, ownership, corrupt handling) plus the two pins added at retirement.

Gate 1 — Spec audit against actual code

Three parallel audit sweeps (expression bindings, deployment invocation, runs/resume/server/stores) agree. Headline: the spec's description of the current code is accurate, and every scheduling seam it names is genuinely missing. In particular: the existing run API supports stopped-run persistence only; durable admission does not exist.

1.1 Input authoring and serialization (spec lines 164-207)

  • Reuse literal/object/array/target-path/strict-JSON/budget semantics: src/wf_core/models/input_bindings.py:11-12 (budget consts), :15-35 InputPathBinding, :38-56 InputValueBinding + strict JSON, :59-97 literal/path/array/object expressions, src/wf_core/models/json_values.py:9-31, src/wf_core/local_paths.py:38-74 overlap checks. Supported as vocabulary.
  • Closed 4-kind expression union; no date/format/arithmetic machinery: input_bindings.py:100-103 discriminator union; only graph-path resolution in src/wf_core/runtime/input_bindings.py:23-114. Supported.
  • Shared composition traversal behind a typed source-resolver seam: resolve_input_expression() takes concrete state/workflow_input/context mappings; grep SourceResolver in src/ has no hits; two more hardcoded recursions exist (src/wf_core/validation/steps.py:242-284, src/wf_api/input_expressions.py:99-253). Missing seam — schedule evaluation can only copy the evaluator or fake context.
  • Typed occurrence reference; graph refs invalid in schedules: no occurrence/scheduled_at/schedule_id model anywhere in src/. Missing (expected; this is the work).
  • Do not extend GraphSourcePath with schedule-only roots: src/wf_core/paths.py:17 (GraphRoot = input/state/context), :232-265 closed roots/parse/factories, :423-446 resolution. Supported as constraint; any occurrence-as-root edit contradicts it.
  • Expression budget enforced on the schedule path: budget checked once at InputExpressionBinding.check_limits (input_bindings.py:176-183); the runtime resolver has no limit check. Missing on any path that bypasses StepInputBinding.

1.2 Deployment invocation and pinned environments (spec 22-24, 198-206)

  • Deployment-run API receives resolved input data: WorkflowRunApi.run_deployment(deployment_id, workflow_input: dict, ...) (src/wf_api/runs.py:77-84); same shape in service.py:1045, surface.py:514, protocols.py:118, transport/models.py:434, wf_client/deployments.py:182. Supported — the scheduler needs only a binding→resolved composer, no API change.
  • Pinned artifact tree captured at run start: PinnedRunEnvironment{deployment, root_artifact, child_artifacts} (src/wf_artifacts/runs/models.py:51), resolve_saved_subgraph_tree (src/wf_api/saved_subgraphs.py:65), resume reuses record.environment without re-reading the deployment (runs.py:192-201). Supported for resume; missing for admission (frozen only in memory, persisted only after stop).
  • Deployment revision for edit-race detection: WorkflowDeployment (src/wf_artifacts/models.py:215) has NO revision; save_deployment overwrites (src/wf_artifacts/store.py:113); only DraftWorkspace has revisions. Missing — "captured invocation stays unchanged" needs a deployment revision or content hash.
  • Durable admission (preassigned run id, persist-before-dispatch, idempotent reconcile): runs.py:101-110 runs raw_plan_from_artifact → full in-memory run_workflow_from_planpersist_stopped_run; id allocated post-execution in run_lifecycle.py:63. Missing entirely (see 1.3).

1.3 Durable admission, resume, server, stores (spec 184-260)

  • wf_api/runs.py executes before persisting: runs.py:102-110 (execute, then persist_stopped_run). Supported.
  • Run models permit only stopped summaries + required checkpoint id: StoredRunStatus = interrupted/completed/failed (runs/models.py:27-32); WorkflowRunRecord.latest_checkpoint_id required (:70); persist_stopped_run rejects active states (run_lifecycle.py:46-61); restore/load assume a checkpoint exists (:101-134). Supported — inspection of an admitted/in-flight run is impossible today.
  • Resume marks the active attempt durably before re-executing: resume_run (src/wf_api/runs.py:194-208): load → validate → resume_workflow_from_plan → persist. No store write between load and execute (the RUNNING flip in src/wf_core/runtime/preparation.py:87 is in-memory only). Missing — crash-during-resume re-presents the old interrupted checkpoint as safe to retry.
  • persist_stopped_run is a transaction: two separate writes, save_checkpoint then save_run (run_lifecycle.py:96-97); each file atomic via tmp+rename (runs/store.py:90-94). Supported as single-file atomicity; cross-file atomicity missing (recovery authority needed).
  • Server has an enable flag, poll loop, capacity, drain: WorkflowServer is a frozen dataclass with no lifecycle methods (wf_server/context.py:279-341); cli.py:112-118 bare uvicorn.run; no lifespan/startup/shutdown/background task in wf_server or wf_transport_rpc_http/app.py:27-68. All missing (expected; this is the work).
  • File-store single-process limits: per-process RLock around individual writes (runs/store.py:42-54); per-run async lock is process-local (runs.py:147, run_locks.py:15-55); zero flock/fcntl/msvcrt/portalocker/FileLock hits in src/; no PID/lease/ownership concept. Supported — matches 2026-06-09-store-transaction-boundary.md. Scheduler ownership needs a Windows-tested held-lock design (that spec forbids ad-hoc lock files without one).
  • Stopped-run persistence implies durable admission: rejected. run_deployment proves execute-then-persist; there is no pre-dispatch record, no preassigned id, no admission lock, and no reconcile.

Gate 2 — Calendar-library probe

Isolated env (NOT the repo env): Python 3.14.7, apscheduler==3.11.0, croniter==6.2.4, tzdata==2026.3, pytest==9.1.1, no pytz (APScheduler used its zoneinfo path).

Reproduce:

$probe = "C:\Users\Admin\AppData\Local\Temp\opencode\sched-cal-probe"
$file = "<worktree>\probes\deployment_scheduling_verify\test_calendar_probe.py"
uv run --project $probe python -m pytest $file -q -p no:cacheprovider `
  -o addopts=""

Result: 16 passed, 2 xfailed (Part A pins the croniter contract; Part B's 2 xfails are APScheduler rejected-candidate evidence — gap phantom and fold replay — marked strict so re-opening that candidacy fails loudly).

Policy: croniter owns calendar calculation, including DST resolution. Part A tests describe observed 6.2.4 behavior the thin adapter consumes (convert now into the schedule zone, ask for next/previous, convert back to UTC). No custom skip/filter machinery exists anymore.

  • UTC daily chain (30 9 * * *): strictly increasing, unique UTC instants over 10 occurrences. Pass.
  • Asia/Ho_Chi_Minh: 09:30 local == 02:30 UTC; 48/48 hourly hits across March and November windows (no DST). Pass.
  • Expression forms on an ordinary day: * → next minute, 5/15 → :05, 0,30 1-2 → 01:00, */20 1-3 from 01:50 → 02:00. Pass.
  • Weekday dialect: 0 AND 7 mean Sunday (weekday()=6), 1/mon mean Monday (0 12 * * 0 == 0 12 * * sun). Pass.
  • DOM/DOW: day_or=True selected (Unix OR — 0 12 13 * fri hits 09-04, 09-11, 09-13, 09-18); day_or=False pinned as the non-selected AND reference. Pass.
  • Zone handling: a UTC start yields UTC results with no conversion — the caller supplies the schedule-zone instant. Pass (adapter rule).
  • DST gap forward (America/New_York 2026-03-08): daily 02:30 resolves to 03:00-04:00 the same day; the following occurrence is 03-09 02:30. Per-minute streams jump 01:59 EST straight to 03:00 EDT with no 02:xx wall times, strictly increasing unique UTC. Pass (observed; the resolution IS the occurrence).
  • DST gap backward: get_prev on the gap day returns the resolved 03:00-04:00; from 03-09 00:00 the same; pre-gap queries return the prior valid occurrence. Pass (observed; latest-missed on a gap day is the resolved instant).
  • DST fold (2026-11-01 01:30): both repeats occur as distinct UTC instants (05:30Z then 06:30Z); per-minute iteration across the fold is 560/560 strictly increasing unique. Pass.
  • Long downtime, per-minute, ~3 years: get_prev answers the latest missed occurrence in ~0.0001s, tz-aware, within a minute of now (UTC and named-zone variants). Pass.
  • Boundary exclusivity: get_next/get_prev from exactly a due instant return the neighboring occurrences; a tick just before due admits it. Pass (drives the T01 watermark rule below).
  • Impossible date (Feb 30): raises documented CroniterBadDateError in ~0.001s. Pass (adapter maps it to exhausted).
  • Malformed expression: CroniterBadCronError at construction. Naive datetimes pass straight through unrejected — the adapter rejects naive itself (spec requires rejection; no DST logic involved). Pass.

Rejected-candidate evidence (APScheduler 3.11.0; NOT acceptance):

  • DST gap: APScheduler fabricates a phantom 02:30-05:00 (= 07:30Z, actually 03:30 EDT) — a wall time that never existed. FAIL (strict xfail).
  • DST fold per-minute: after 06:00Z, APScheduler flips back to -04:00 and replays 05:01Z06:00Z (~59 duplicate UTC identities, UTC goes backward). FAIL (strict xfail).
  • No bounded latest-missed seam exists in its documented trigger API surface (forward-only get_next_fire_time); ~1.6M iterations would be needed for 3 years of per-minute misses.

Thin-adapter contract (replaces the retired skip rule)

T01 implements exactly this and nothing more:

  • Convert the query instant into the schedule's named zone; ask croniter for the next (next_after) or previous (prev_before) occurrence; convert the result to UTC.
  • Iteration is exclusive both directions (pinned above), so forward queries start from the last-consumed instant and catch-up results are compared against the same watermark: a due occurrence is admitted exactly once, never missed by exclusivity nor doubled.
  • Use the library's bounded-search controls and documented exceptions: CroniterBadDateError maps to exhausted; CroniterBadCronError at construction is a definition rejection. Search exhaustion or a backward/non-progressing result fails visibly — it must never look like successful progress, and the adapter must not "repair" results with its own calendar engine.
  • Admitted identity stays (schedule_id, resolved UTC), which always exists. The definition's cron expression, time-zone name, and admission snapshots are retained; no finer intended-wall-time provenance is claimed. Minute precision (this slice).

Recommendation

  1. APScheduler has no bounded latest-missed seam in its documented trigger API surface (forward-only get_next_fire_time). The spec's "never enumerate years of missed occurrences" gate cannot be met with it; croniter.get_prev answers in ~0.0004s.
  2. APScheduler violates the occurrence-identity gate: it replays ~59 past UTC minutes after every fall fold (duplicates + backward UTC), which would double-admit per-minute schedules. croniter is clean (pinned).
  3. Across DST gaps APScheduler fabricates a nonexistent wall time; croniter resolves gap days forward to an existing wall time and emits both fold hours as distinct UTC occurrences (all pinned).
  4. The spec needs only iteration, not APScheduler's job store/executor (which it already excludes). croniter is the smaller dependency with the needed get_prev seam and standard DOM/DOW day_or=True (pinned).
  5. Pin croniter==6.2.4 + tzdata floor (python-dateutil stays transitive-only); the probe pins gap/fold/dialect/zone behavior so upgrades re-probe (both strict xfails fail loudly if a future APScheduler fix tempts re-opening that candidacy, and the croniter pins — including the exact version assertion — catch drift).

Gate 3 — Scheduling state-model probe

probes/deployment_scheduling_verify/test_schedule_state_model.py: pure-stdlib reference model (injected clock, scripted executor outcomes, fault injection, ownership lock) + 31 tests, all passing in the repo env:

uv run pytest -q `
  probes/deployment_scheduling_verify/test_schedule_state_model.py `
  -p no:cacheprovider -o addopts=""

The calendar is abstract (next_after/prev_before only — deliberately no enumeration seam), so this pressures the state rules, not date math. Each test is a short timeline, no sleeps. Coverage maps to the assignment bullets:

  • overlap=skip|parallel × misfire=skip|latest — all four combinations (incl. parallel limits + limit lowering + parallel×latest catch-up)
  • latest-means-ONE-candidate + supersession + no-double-admit at the boundary tick + newer-due-never-touches-admitted-runs + timely admission supersedes older held candidates
  • manual runs and other schedules' runs excluded from per-schedule overlap
  • run identities are store-backed and survive restart (no _seq-style in-memory counter reuse — this was a real reference-model bug, fixed with a regression test)
  • terminal overlap skips never reappear after restart/catch-up
  • terminal overlap skips never reappear after restart/catch-up
  • per-schedule limits count interrupted runs; waiting interruptions hold no task slot; resume requires a task slot
  • explicit pause is not downtime (pause excludes the interval; enabled downtime catches up) via resume_schedule/edit_schedule semantics
  • edit discards candidates without backfill; delete clears pending, keeps history, never cancels active runs, ids not reusable
  • restart persists candidate + consumed progress; rollback never re-admits
  • 3-year per-minute downtime: ≤ SCAN_CAP + 2 source calls, one interval-summary row, ≤1 admission (capacity) or exactly 1 held candidate
  • round-robin fairness across schedules; capacity-deadline expiry for skip vs hold-one for latest; capacity-wait within allowance (undecided, not consumed) then admit-or-expire
  • fault before AND after admission (no dispatch; recovery materializes the view and flags pending-dispatch, the next poll dispatches through capacity checks exactly once — recovery NEVER executes), fault after materialize, fault after complete and after interrupt (recovery reconciles the missing terminal record), fault after resume-mark / resume-complete / resume-attempt-clear / resume-interrupt before AND after (attempt identity distinguishes fresh results from stale checkpoints), crash after dispatch (abandoned FAILED with external-effects disclosure, no replay), resumed runs re-interrupting durably
  • crash during resume: ACTIVE-attempt marker distinguishes ambiguous (FAILED, never retried) from merely waiting (resumable)
  • preflight rejection invents no run; admitted runs freeze invocation
  • second-owner rejection; no-expiry held lock; unsupported locking rejects startup; corrupt view-without-admission fails closed and blocks the schedule

Key state findings for implementation (all encoded in the reference model):

  • F1 — scan cap + jump rule. Per poll/schedule: at most SCAN_CAP (100) next_after calls, then jump via one prev_before (latest) or an interval-summary + consumed = now (skip). This is the executable form of "never enumerate"; only prev_before makes latest bounded (Gate 2).
  • F2 — consumed + candidate are the whole restart story. consumed_through (latest decided instant) plus at most one pending candidate reconstruct everything; terminal skips/admissions are never rebuilt because their instants are ≤ consumed.
  • F3 — ordering inside admission: recheck → overlap → capacity → allocate/freeze → persist admission → materialize view → dispatch; the admitted-history entry belongs to the admission persist, not the view.
  • F4 — resume marker first. resume_attempts[run] = ACTIVE is durable before re-execution; recovery reconciles a matching stopped result and fails unmatched ACTIVE attempts closed. Pre-existing waiting interruptions (no marker) stay resumable — no migration problem.
  • F5 — dispatched RUNNING work at recovery is abandoned (in-memory tasks die with the process). Proven undispatched admissions remain pending; a run without an admission is corrupt and must block its schedule.
  • F6 — one-shot exhaustion is a flag, set on admission or on skip-expiry; without it the expiry branch refires every poll.
  • F7 — pause/edit/delete/disable are baseline operations (consumed = max(...) to now, clear candidates), not filters inside the poll loop. Resolved edge: administrative disable behaves like pause for catch-up (the spec defines no separate disable semantics).
  • F8 — timely admission supersedes older held candidates. A poll that admits instant N while a candidate for older instant C is held records C superseded and clears it; the admitted-history entry belongs to the admission persist (before the view materialization), so an admission-after crash still reconciles exactly once.
  • F9 — recovery reconciles missing terminal records. A COMPLETED or waiting-INTERRUPTED run with no terminal history entry (crash between state persist and record append) gets exactly one reconciled entry; re-polling never re-admits because consumed already advanced. A COMPLETED run with a still-ACTIVE attempt crashed between completion and attempt-clearing: reconcile to DONE, never re-execute.
  • F10 — run identities are store-backed. The id counter lives in the store, not the scheduler instance; restart allocates fresh identities and both occurrences keep their runs.
  • F11 — recovery never executes. Runs materialized by recovery are flagged pending-dispatch: they hold a schedule slot but no task slot, and the next poll dispatches them through capacity checks (waiting while full). Pending runs of blocked schedules stay pending.
  • F12 — resume completion is three persists, re-interruption is durable, and stopped results carry the attempt identity. Completion persist, attempt-clear, and history append each have a fault boundary; a resumed run may interrupt again (own persist, stays resumable). Every resume attempt takes a store-backed identity at mark time, and every stopped result it produces echoes that identity back: recovery matches result to ACTIVE attempt (fresh → DONE + resumable) versus stale (ambiguous → failed, never retried). Crashes inside completion windows still fail closed, but a persisted re-interruption is never mistaken for the previous checkpoint.

Sequenced implementation plan

Conventions: each task lists goal, exact seams/files, test-first source (port the named probe tests into tests/scheduling/ — do not copy probe logic into src/), and done criteria. Dependencies noted per phase. Assumes the Gate 2 calendar policy (croniter owns DST resolution; thin adapter, no custom skip/filter). New production code lives in a focused package src/wf_scheduling/ (new area; per AGENTS.md prefer packages over flat files) plus the listed seam edits; no second evaluator, no fake workflow context, no GraphSourcePath extension.

R0 — Review checkpoint (before any production code)

Reviewers confirm the thin-adapter contract, the croniter pin, and the src/wf_scheduling/ package boundary. Gate: this document + green probes.

Phase 0 — Calendar adapter (no scheduler yet)

  • T01 — Occurrence-source seam + croniter adapter. New src/wf_scheduling/calendar.py: OccurrenceSource Protocol (next_after, prev_before over aware datetimes), CronSource (5-field cron + IANA zone; converts now→zone, queries, returns UTC) and OneShotSource (offset-required, naive rejected). Deliberate dialect: croniter Unix convention (numeric 0 AND 7 = Sunday; day_or=True selected — NOT APScheduler's Monday-first). Thin by construction: no calendar correction of its own — forward queries start from the last-consumed instant, catch-up results compare against the same watermark (iteration is exclusive both directions, pinned), CroniterBadDateError maps to exhausted, CroniterBadCronError at construction is a definition rejection, and any backward/non-progressing result fails visibly. Add croniter==6.2.4 and a tzdata floor to dependencies (python-dateutil remains transitive-only). Port: all Part A test_calendar_probe.py tests into adapter-level tests (the two Part B strict xfails stay as APScheduler rejected-candidate pins). Done: adapter suite green; impossible schedules map to exhausted (no CroniterBadDateError leaks).
  • T02 — Occurrence identity + UTC persistence helpers. src/wf_scheduling/occurrences.py: identity = (schedule_id, resolved_utc); occurrence_id derived deterministically from that pair (document the derivation; spec exposes schedule_id/occurrence_id/scheduled_at); monotonic-clock sleep vs wall-clock eligibility split; rollback guard (resolved <= consumed never re-admitted). Port: calendar uniqueness/ increasing tests plus the reference-model rollback test. Depends: T01.

Phase 1 — Expression seam (review checkpoint R1)

  • T03 — Extract shared traversal behind a typed source resolver. New src/wf_core/runtime/input_sources.py hosting the SourceResolver Protocol; refactor src/wf_core/runtime/input_bindings.py:23-114 so path resolution goes through it (graph root stays a resolver argument — no GraphSourcePath change); unify the sibling recursions (src/wf_core/validation/steps.py:242-284, src/wf_api/input_expressions.py:99-253) onto one traversal + budget check. Existing node/subgraph/interrupt call sites (src/wf_core/runtime/ops/nodes.py:48-66, src/wf_core/runtime/subgraphs.py:139-145, src/wf_core/runtime/ops/interrupts.py:26-43) pass the graph resolver. Done: no behavior change; full existing suite green. New tests: resolver unit tests (graph vs occurrence resolvers over one traversal).
  • T04 — Typed occurrence expression kind. New OccurrenceExpression{field: schedule_id|occurrence_id|scheduled_at} in src/wf_core/models/input_bindings.py WITHOUT touching src/wf_core/paths.py roots; schedule-side binding list type (ScheduleInputBindings) reusing validate_input_expression_limits; resolution via an occurrence resolver (never via faked graph context). Strict-JSON/target-conflict/schema validation unchanged. Tests: JSON round-trip, missing-field, graph-only-path, over-budget, conflicting target, invalid-resolved-input — port the contract pins from probes/.../test_expression_contract_probe.py (8 tests documenting the current union/budget/strict-JSON/overlap/roots behavior) and extend them to the new kind. Depends: T03.
  • R1 — review: resolver Protocol, occurrence kind, budget parity.

Phase 2 — Durable admission representation (review checkpoint R2)

  • T05 — Admission record + in-flight run view. Extend src/wf_artifacts/runs/models.py: admitted/in-flight status alongside stopped statuses (never fabricate checkpoint/trace/output for unknown outcomes); preassigned run identity from a store-backed counter (never an instance counter — restart must not reuse identities, per F10); admission persists the pinned environment (PinnedRunEnvironment: deployment + root/child artifacts), resolved input, limits (max_steps), deployment revision, and the occurrence's resolved UTC instant; run inspection distinguishes admitted vs stopped-with-checkpoint. Atomic-per-file writes stay; multi-file authority order per F3 (the occurrence is decided at the admission persist: history entry, candidate clearing, consumed progress, and one-shot exhaustion all belong to it). Port: fault-before/after-admission and admission-vs-view recovery tests (reference model, minus executor). Depends: T02.
  • T06 — Admission path through the run API. Rework WorkflowRunApi.run_deployment (src/wf_api/runs.py:77-134) into recheck → allocate/freeze → persist admission → materialize → dispatch-captured → persist stopped → reconcile, reusing the resolved-input signature (no second API). Manual runs keep the current call shape; existing manual run/resume tests must pass unchanged (regression gate). Depends: T05. R2 — review: admission ordering + inspection contract.

Phase 3 — Scheduler core (review checkpoint R3)

  • T07 — Schedule/deployment-revision models + file store. src/wf_scheduling/models.py (schedule, revision, overlap/misfire, max_active_runs, lateness allowance, zone, pause/delete/disable flags, exhausted flag per F6, no id reuse) and src/wf_scheduling/store.py (FileScheduleStore: schedules, at-most-one candidate, consumed progress, interval summaries, occurrence history with cursor pagination over (resolved_utc, occurrence_id) + limit). Add deployment revision (or content hash) to WorkflowDeployment (src/wf_artifacts/models.py:215, store.py:113) for edit-race rechecks. Port: edit/delete/pause/disable/no-backfill/id-reuse tests. Inspection payload carries resolved-UTC/admission/actual-start times, revision, run id, and failure/skip reason (spec lines 263-272).
  • T08 — Poll loop: overlap/misfire/candidates/fairness/capacity. src/wf_scheduling/poll.py: F1 scan-cap rule, latest-only coalescing + supersession records, overlap-before-capacity precedence, terminal skips, capacity-deadline expiry (skip) vs hold-one (latest), round-robin fairness, per-poll batch bounds, task-slot accounting (interrupted and pending-dispatch runs hold schedule slots only), plus the pending-dispatch sweep (dispatch recovery-materialized runs through capacity checks; never inside recovery — F11). Port: matrix/coalescing/fairness/capacity/pending-dispatch reference tests. Depends: T02, T07. R3 — review: poll semantics vs reference model.

Phase 4 — Resume safety, recovery, ownership (review checkpoint R4)

Safety and recovery land BEFORE anything enables dispatch.

  • T09 — Durable resume-attempt marker + granular completion. resume_run (src/wf_api/runs.py:136-224) persists ACTIVE with a store-backed attempt identity before re-executing; completion, attempt-clearing, and history recording are separate persists with a fault boundary between each pair; every stopped result the attempt produces echoes the attempt identity back; a resumed run may interrupt again (durable re-interruption, stays resumable); waiting interruptions (no marker) stay resumable across restart. Port: crash-during-resume, resume-complete/attempt-clear/interrupt before+after fault, attempt-identity match/mismatch, and re-interruption tests. Depends: T06.
  • T10 — Startup recovery + reconciliation. Recovery under exclusive ownership, and recovery NEVER executes: it materializes missing views as pending-dispatch (dispatched later only via the poll sweep), fails abandoned/ambiguous runs with external-effects disclosure (no replay), matches stopped results to the ACTIVE attempt by identity (fresh → DONE + resumable; stale → ambiguous FAILED, never retried), reconciles missing terminal records and COMPLETED-with-ACTIVE attempts (to DONE, never re-run), fails corrupt views closed + blocks the schedule, and preserves stopped interruptions in their slots. Port: all recovery reference tests. Depends: T06, T08, T09.
  • T11 — Exclusive file-store ownership. Held cross-process lock (no PID file, no expiring lease), second-owner rejection, release on death, startup rejection where locking is unsupported — consistent with 2026-06-09-store-transaction-boundary.md (no ad-hoc lock files without a tested design). Start with a locking-design spike: candidates are a dedicated lock file held via msvcrt.locking (Windows) / fcntl.flock (POSIX) in one new module, vs a portalocker-style dependency; pick after a Windows crash-hold-release test. Port: ownership reference tests. Depends: T07. R4 — review: fault-injection suite (every boundary before/after).

Phase 5 — Lifecycle enablement and administration surface

  • T12 — Shutdown drain + lifecycle (GATED: only after R4 passes). Opt-in scheduler composition on WorkflowServer (src/wf_server/context.py:279-341; startup in src/wf_server/cli.py; transport hooks in src/wf_transport_rpc_http/app.py:27-68): explicitly-enabled flag, stop-admission-first + grace-period drain (scheduler_drain_grace_s config) with cancellation/failure recording. Dispatch must not be enabled before resume safety (T09), recovery (T10), and ownership (T11) land. Port: paused/deleted-schedule completion/resume reference tests; drain and bounded-concurrency tests are NEW (no probe source — the reference model has no server). Depends: T08, R4.

  • T13 — API + Python client. New WorkflowApi schedule methods (create/get/list/update/pause/resume/ delete_schedule — public names finalized in this task, not assumed to exist) + paginated occurrence inspection (pending, coalesced/superseded, skipped-overlap, skipped-misfire, preflight-rejection, admitted/running, interrupted, completed, failed) through wf_api/service.py + wf_api/surface.py, JSON-RPC transport (wf_transport_rpc_http), and wf_client (protocols.py, deployments.py, app.py, runs.py); revision-checked edits; run limits incl. schedule max_steps reusing manual-run validation; inspection without a checkpoint must work (admitted runs have none). Client round-trip/pagination tests are NEW (no probe source — the reference model has no transport). Depends: T08, T10.

  • T14 — Docs + seam comments + probe retirement. Update live docs per docs/AGENTS.md (roadmap pointers, no narrative bloat), add code-seam comments where docs describe behavior, delete probes/deployment_scheduling_verify/ once production tests subsume it. Markdown lint on changed files only. Depends: everything.

Non-goals (unchanged): wait nodes, distributed workers, automatic failed-run retries, replay-all bursts, generic Runtime[UserContext], exactly-once external effects, multi-writer file stores beyond scheduler ownership.

Remaining concerns

  1. No product decisions block Phase 0. The 60-second default lateness allowance is approved (per-schedule configurable).
  2. CroniterBadDateError from impossible-schedule search and naive-time passthrough are adapter-level (map to exhausted / reject; no spec change needed).
  3. Locking mechanism (T11) needs a Windows-tested design before code, per the store transaction boundary — flagged as a task, not a decision.
  4. Probes use scripted/abstract occurrence sources; production cron parsing arrives via T01's adapter, already pinned by the calendar probe.

Planning bundle files

  • docs/superpowers/specs/2026-09-08-deployment-scheduling-design.md — updated worktree specification incorporating the approved calendar policy
  • probes/deployment_scheduling_verify/README.md — new (disposable label; isolated-env + repo-env run commands)
  • probes/deployment_scheduling_verify/test_calendar_probe.py — new (16 passed, 2 strict xfailed in the isolated env: Part A pins the croniter contract, Part B keeps APScheduler rejected-candidate evidence; importorskip keeps default repo collection green)
  • probes/deployment_scheduling_verify/test_schedule_state_model.py — new (31 passed in the repo env)
  • probes/deployment_scheduling_verify/test_expression_contract_probe.py — new (8 passed in the repo env; pins the current expression contract for Phase 1)
  • docs/superpowers/plans/2026-09-09-deployment-scheduling-implementation-plan.md — this file (new; reviewed three times — first pass B1B8/A1A3, second pass (identity/ordering fixes), third pass: skip rule corrected to expanded sets with work-bounded forward/backward forms, attempt identity for stopped results, lifecycle enablement gated after safety/recovery/ownership; fourth pass: custom skip/filter machinery retired — croniter owns DST resolution, thin-adapter contract with last-consumed watermark, day_or=True selected, resolved-UTC identity throughout)