Problem shape
The problem beneath them
Write each step as a plain typed function, derive the graph from parameter names, and store every run as a folder where each step writes its outputs before its metadata, so a crash loses only the steps that were running and any run can be resumed, forked or reviewed later.
Why I built it
I built hone-flow for a local pipeline that writes song ideas, lyrics, music and video on one 8 GB GPU. Each run took hours. The hand-written runner collected the usual bugs: a state.json "done" flag that didn't notice new lyrics, a step list repeated in three places that drifted, --from and --until flags rewritten for every pipeline, review notes that no step ever read, and overnight runs that failed silently. Rerunning from scratch cost hours of GPU time and gave different samples, so "just run it again" wasn't an option.
The underlying problem is that an AI pipeline needs three things most scripts don't give it: a durable record of each step's inputs and outputs, a clear rule for what may be reused, and a way to pause for a person without keeping a process alive for days.
hone-flow answers with run folders. Every run is <storage>/<workflow>/runs/<run_id>/: a manifest, one folder per step and item with its inputs, outputs and metadata.json, a spans.jsonl of what happened, and reports on timing and output size. There is no server, scheduler or database. The folder is the state.
Resume finishes the same run and refuses to mix step versions. It does not replay code inside a step or protect external side effects from running twice.
A step that crashes halfway runs again from its start; a long step keeps its own checkpoints. Payments, emails and other external effects need their own idempotency (see the durable execution reference design). There is no automatic cross-run cache: reuse happens only through an explicit fork of a run you chose.
Different problems, same shape
The surface changes. The decision structure persists.
The model call failed on item 37 of 50. Must items 1 to 36 run again?
A script keeps progress in memory or in a hand-written state file.
I changed one prompt. Which outputs are now stale?
A 'done' flag says complete although an input changed.
An editor rejected a draft with a note. Does the next attempt see it?
Review happens in chat and the pipeline never reads it.
Which code, model and parameters made this file three weeks ago?
Outputs are scattered paths with no record of what produced them.
All four questions converge into Resumable AI Workflow Runs.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Step definition
- Question
- What does each step need, and where does it come from?
- Responsibility
- Read typed function signatures; a parameter named after a step receives that step's output, `fk.Param[...]` marks a run setting, any other name is an item input.
- Input
- Decorated plain functions (`@wf.step`, `@wf.gate`, `@wf.final_step`, `@wf.select_step`).
- Output
- One workflow graph, the only definition every command reads.
- Stops when
- The graph is valid or a `WorkflowDefinitionError` names the problem.
-
02
Run folder
- Question
- Where does everything about this run live?
- Responsibility
- Create a self-contained folder, local or on S3, with a manifest of items, params, steps and their states.
- Input
- Items, params, storage location, optional label.
- Output
- `runs/<run_id>/manifest.json` and one folder per step and item.
- Stops when
- The run is registered and every step has a state.
-
03
Commit per step
- Question
- When is a step's result safe to reuse?
- Responsibility
- Write outputs and input copies first, `metadata.json` last, so a step folder with metadata is complete.
- Input
- The step's return value, inputs, code version, timing and measurements.
- Output
- A complete step folder or none.
- Stops when
- Metadata is written or the step is marked failed with its error.
-
04
Resume
- Question
- What work is left after a crash, failure or partial run?
- Responsibility
- Run pending, failed, interrupted and skipped work in the same run, and refuse with `IncompatibleRun` when a step with work left changed its version.
- Input
- An existing run folder and the current workflow code.
- Output
- The same run, advanced.
- Stops when
- Nothing is left, a gate waits, or a step fails again.
-
05
Fork
- Question
- I changed something. What must rerun?
- Responsibility
- Make a new run beside the old one; copy what didn't change, rerun changed steps and their downstream, and explain every rerun (version, source, params, input files). `dry_run=True` shows the plan.
- Input
- A finished run and either `refresh=(...)` or detected changes.
- Output
- A new run with reused results labelled `reused`.
- Stops when
- The forked run completes, pauses or fails like any run.
-
06
Gate
- Question
- Where must a person decide before more work is spent?
- Responsibility
- Pause each item at a gate; a person approves, edits or rejects with a note, possibly days later from another process.
- Input
- Outputs of the steps the gate reviews.
- Output
- An approved or edited value downstream, or a rerun of the producer with `review_note` and `previous`.
- Stops when
- Every gated item has a decision.
-
07
Evidence
- Question
- What did this run use, and what happened?
- Responsibility
- Record spans, timing, output sizes, optional CPU, memory and GPU measurements, and every attempt including failed and rejected ones.
- Input
- Step executions and their trace context.
- Output
- `spans.jsonl`, `reports/` and per-step metadata readable without the workflow code.
- Stops when
- The run ends; the evidence stays with the folder.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Step folder has `metadata.json` | Treat the step as complete | Metadata is written last, after outputs | The folder was edited by hand |
| Step with work left changed version | Refuse to resume (`IncompatibleRun`) | Mixing versions in one run hides what produced what | The user needs the change: fork instead |
| One step's prompt changed | Fork with that step refreshed | Only it and its downstream need new samples | The fork plan reruns more than expected |
| Gate rejects an output | Rerun the producer with the note and previous output | The note is information the step needs | The same item is rejected repeatedly |
| Step raises | Mark failed, keep other items going | One bad item shouldn't cost the batch | Failures exceed what the person tolerates |
| Notification endpoint down | Record pending delivery, don't fail the run | An outage of chat isn't a pipeline failure | `notify-retry` is never run |
Operating paths