Checkpoints
Persist pipeline state mid-execution so long pipelines survive wall-clock / runner crashes.
v0.4.0concept
Checkpoints (v0.3+)
Persist pipeline state mid-execution so long pipelines survive wall-clock / runner crashes.
Use-case priority
| Use case | Status |
|---|---|
| 30s Cloudflare Workers wall-clock | MUST in v0.3 (otherwise long pipelines don’t run server-side at all) |
| Server / CI runner crash mid-execute | MUST (same mechanism) |
| User pauses run, resumes later | COULD (CLI ergonomic, can ship later) |
| Cross-execution continuity (multi-day) | COULD (separate scope) |
Storage model
- id: long-phase
checkpoint:
on: stage_complete
storage:
type: r2
path: checkpoints/{pipeline_id}/{exec_id}.json
retention: 24h
Stage serialization shape: { stages: [progress], pending_stages: [...], resolved_inputs: {...], timestamp }.
On resume: deserialize → skip completed stages → continue from pending.
What gets checkpointed
NOT stage outputs (recomputable from inputs). Only:
- Which stages completed (id + start timestamp)
- Resolved template variables (avoid re-running expensive interpolation)
- Current DAG position (UI progress display)
- Operator signature so far (resume preserves audit trail)
Behavioral oracle
packages/pipeline-runtime/tests/checkpoints.test.ts— R10.2 checkpoints: serializecompletedStageIds+ pending DAG after each stage, restore onresume: true, skip already-completed stagespackages/server/tests/checkpoint-r10-4.test.ts— server-side R2 storage backend (checkpoint write/read/delete lifecycle)packages/pipeline-runtime/tests/integration-r10-x.test.ts— end-to-end: long pipeline interrupted, resumed via checkpoint storage