Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer
On this pattern

System Pattern 01 / Durable Execution & Integration

A run is a folder you can open.

Resumable AI Workflow Runs

Write each step of an AI pipeline as a plain function, and keep every run as a self-contained folder that can be resumed after a crash, forked to rerun only what changed, and paused for a person.

Maturity
Tested / Prototype
Critical uncertainty
Resume finishes the same run and refuses to mix step versions. It does not replay code inside a step or protect external side effects from running twice.
Human boundary
Approve, edit or reject gated outputs; choose what to fork
Applications
Batch LLM generation · Document pipelines · Media generation on a local GPU · Evaluation runs
Terminal output, captured 2026-09-30: hone-flow lists two runs of a local story-room pipeline, one completed run shows 131 steps done and 682 not selected, its run folder holds one folder per step, and its report lists the slowest steps.

Problem shape

The problem beneath them

Write each step as a plain typed function, derive the graph from parameter names, and store every run as a folder where each step writes its outputs before its metadata, so a crash loses only the steps that were running and any run can be resumed, forked or reviewed later.

Why I built it

I built hone-flow for a local pipeline that writes song ideas, lyrics, music and video on one 8 GB GPU. Each run took hours. The hand-written runner collected the usual bugs: a state.json "done" flag that didn't notice new lyrics, a step list repeated in three places that drifted, --from and --until flags rewritten for every pipeline, review notes that no step ever read, and overnight runs that failed silently. Rerunning from scratch cost hours of GPU time and gave different samples, so "just run it again" wasn't an option.

The underlying problem is that an AI pipeline needs three things most scripts don't give it: a durable record of each step's inputs and outputs, a clear rule for what may be reused, and a way to pause for a person without keeping a process alive for days.

hone-flow answers with run folders. Every run is <storage>/<workflow>/runs/<run_id>/: a manifest, one folder per step and item with its inputs, outputs and metadata.json, a spans.jsonl of what happened, and reports on timing and output size. There is no server, scheduler or database. The folder is the state.

Resume finishes the same run and refuses to mix step versions. It does not replay code inside a step or protect external side effects from running twice.

A step that crashes halfway runs again from its start; a long step keeps its own checkpoints. Payments, emails and other external effects need their own idempotency (see the durable execution reference design). There is no automatic cross-run cache: reuse happens only through an explicit fork of a run you chose.

Different problems, same shape

The surface changes. The decision structure persists.

01 / Batch generation

The model call failed on item 37 of 50. Must items 1 to 36 run again?

A script keeps progress in memory or in a hand-written state file.

02 / Prompt changes

I changed one prompt. Which outputs are now stale?

A 'done' flag says complete although an input changed.

03 / Human review

An editor rejected a draft with a note. Does the next attempt see it?

Review happens in chat and the pipeline never reads it.

04 / Local GPU work

Which code, model and parameters made this file three weeks ago?

Outputs are scattered paths with no record of what produced them.

All four questions converge into Resumable AI Workflow Runs.

Exploded pattern

Open the mechanism at every decision boundary.

  1. 01

    Step definition

    Question
    What does each step need, and where does it come from?
    Responsibility
    Read typed function signatures; a parameter named after a step receives that step's output, `fk.Param[...]` marks a run setting, any other name is an item input.
    Input
    Decorated plain functions (`@wf.step`, `@wf.gate`, `@wf.final_step`, `@wf.select_step`).
    Output
    One workflow graph, the only definition every command reads.
    Stops when
    The graph is valid or a `WorkflowDefinitionError` names the problem.
  2. 02

    Run folder

    Question
    Where does everything about this run live?
    Responsibility
    Create a self-contained folder, local or on S3, with a manifest of items, params, steps and their states.
    Input
    Items, params, storage location, optional label.
    Output
    `runs/<run_id>/manifest.json` and one folder per step and item.
    Stops when
    The run is registered and every step has a state.
  3. 03

    Commit per step

    Question
    When is a step's result safe to reuse?
    Responsibility
    Write outputs and input copies first, `metadata.json` last, so a step folder with metadata is complete.
    Input
    The step's return value, inputs, code version, timing and measurements.
    Output
    A complete step folder or none.
    Stops when
    Metadata is written or the step is marked failed with its error.
  4. 04

    Resume

    Question
    What work is left after a crash, failure or partial run?
    Responsibility
    Run pending, failed, interrupted and skipped work in the same run, and refuse with `IncompatibleRun` when a step with work left changed its version.
    Input
    An existing run folder and the current workflow code.
    Output
    The same run, advanced.
    Stops when
    Nothing is left, a gate waits, or a step fails again.
  5. 05

    Fork

    Question
    I changed something. What must rerun?
    Responsibility
    Make a new run beside the old one; copy what didn't change, rerun changed steps and their downstream, and explain every rerun (version, source, params, input files). `dry_run=True` shows the plan.
    Input
    A finished run and either `refresh=(...)` or detected changes.
    Output
    A new run with reused results labelled `reused`.
    Stops when
    The forked run completes, pauses or fails like any run.
  6. 06

    Gate

    Question
    Where must a person decide before more work is spent?
    Responsibility
    Pause each item at a gate; a person approves, edits or rejects with a note, possibly days later from another process.
    Input
    Outputs of the steps the gate reviews.
    Output
    An approved or edited value downstream, or a rerun of the producer with `review_note` and `previous`.
    Stops when
    Every gated item has a decision.
  7. 07

    Evidence

    Question
    What did this run use, and what happened?
    Responsibility
    Record spans, timing, output sizes, optional CPU, memory and GPU measurements, and every attempt including failed and rejected ones.
    Input
    Step executions and their trace context.
    Output
    `spans.jsonl`, `reports/` and per-step metadata readable without the workflow code.
    Stops when
    The run ends; the evidence stays with the folder.

Decision forks

Every branch states why it exists and when it escalates.

Signals, decisions, reasons, and escalation conditions
Signal Decision Reason Escalates when
Step folder has `metadata.json` Treat the step as complete Metadata is written last, after outputs The folder was edited by hand
Step with work left changed version Refuse to resume (`IncompatibleRun`) Mixing versions in one run hides what produced what The user needs the change: fork instead
One step's prompt changed Fork with that step refreshed Only it and its downstream need new samples The fork plan reruns more than expected
Gate rejects an output Rerun the producer with the note and previous output The note is information the step needs The same item is rejected repeatedly
Step raises Mark failed, keep other items going One bad item shouldn't cost the batch Failures exceed what the person tolerates
Notification endpoint down Record pending delivery, don't fail the run An outage of chat isn't a pipeline failure `notify-retry` is never run

Operating paths

Clear, ambiguous, and failed work all reach explicit states.

Clear path

Define steps, run over items, each step commits, the final step combines results.

Final state
Completed run.
Owner
The person who started the run.
Evidence
Manifest, step folders, `spans.jsonl`, timing and size reports.
Recovery
Fork to rerun any step later without touching this run.
Ambiguous path

A gate pauses items; a reviewer approves, edits or rejects with a note.

Final state
Paused at a gate, then completed after resume.
Owner
The reviewer named by the team, working through the CLI (`approve`, `edit`, `reject`) or code.
Evidence
Review decisions stored in the step records with the note and the previous output.
Recovery
`resume` continues downstream or reruns the producer.
Failure path

A step fails or the process crashes; completed steps keep their folders.

Final state
Failed run, resumable.
Owner
The pipeline author.
Evidence
Error, attempts and the last complete step per item.
Recovery
Fix the cause and `resume`; if the fix changes a step version, fork.

Authority map

Capability does not grant authority.

RULE

May decide
Which steps are complete, which may be reused, whether versions are compatible
May not decide
That an incomplete step folder is complete
Required evidence
Metadata presence, recorded versions

MODEL

May decide
Nothing in hone-flow; models run inside user steps
May not decide
Step completion, reuse or approval
Required evidence
Recorded as spans from the step's own calls

SYSTEM

May decide
Order steps, write folders, resume, fork and explain reruns
May not decide
Skip a gate or rewrite a finished run
Required evidence
Manifest, step records, fork explanation

HUMAN

May decide
Approve, edit or reject gated outputs; choose what to fork
May not decide
Edit history in place without a trace
Required evidence
Review record with actor and note

EXCEPTION

May decide
Hold failed items and waiting gates as explicit states
May not decide
Count as completed
Required evidence
Error text, attempts, gate status

Failure modes and recovery

A failed path remains owned, evidenced, and recoverable.

Failures, detection, containment, recovery, and owners
Failure Detection Containment Recovery Owner
Process crash mid-run Steps without metadata Complete steps untouched `resume` reruns only incomplete steps Pipeline author
Model call fails for one item Step raises Other items continue `resume` after the cause is fixed Pipeline author
Step code changed between runs Version compare Resume refused Fork with explained reruns Pipeline author
Reviewer never answers Gate stays waiting; notification sent Downstream steps don't spend GPU Reviewer decides; `resume` Reviewer
Chat notification fails Delivery recorded as pending Run unaffected `notify-retry` Operator
Storage becomes unavailable Write error Step not marked complete Rerun or resume when storage returns Operator

Invariants and guarantees

Properties the structure is designed to preserve.

  • A step folder with metadata.json is complete; one without it is not.
  • A resumed run never mixes step versions.
  • Reused results are labelled reused, never passed off as fresh samples.
  • Every attempt, including failed and rejected ones, stays in the run's records.
  • A rejection note reaches the step that produced the rejected output.
  • Nothing runs unless someone runs it: no server, scheduler or background worker.

What changes between implementations

The constraints determine the final mechanism.

Storage changes: a local folder for one machine, s3:// or any fsspec URL when runs must outlive the laptop. Gates change with the team: a CLI approval for one person, a small web page reading the run through fk.open_runs for a group. Measurements change with the hardware: timing and sizes by default, GPU memory when steps share a card (hone-flow can take GPU leases from hone-models).

Open-source alternatives I compared

Tool What it does well Why it didn't fit this job
Airflow, Prefect, Dagster, Flyte, ZenML Scheduling, deployments, many integrations Built around servers and schedulers; heavy for a local tool or a script
Hamilton Dataflow of plain functions No steps over items with human pauses
Snakemake, DVC pipelines Good rerun triggers Think in files and shell commands
Kedro, Metaflow Pipeline slicing; artifacts and resume Tie you to their classes and infrastructure
Temporal, DBOS Durable execution with replay Right for services, more than a local AI pipeline needs
Burr, LangGraph Persisted state, human pauses Aimed at agents, not batches of items

None combined plain-function steps, fan-out over items, a readable record per run on local disk or S3, explicit resume and fork, human checkpoints and GPU-aware batching without a server. That combination is the reason hone-flow exists; if you already run one of the platforms above, the pattern still applies and the platform is often the better home for it.

Evidence chain

Follow the pattern into systems and software.

Implemented in hone-flow (Apache-2.0, alpha 0.1.0) with acceptance tests in CI. I use it in a local generation pipeline; no production deployment or measured result is published.

The invoice-intake reference design processes batches of documents where one bad page must not force a rerun of the rest. The on-premises call review design runs transcription and checking as long batch steps on a private GPU. Both are design relationships: those cases are reference designs, and hone-flow is one way to build their batch parts.

Implemented in hone-flow, Apache-2.0, version 0.1.0 (alpha). The design, its guarantees and every design change with its reasons are in the repository's design/ folder; the acceptance cases run as end-to-end tests in CI, and examples/ holds 27 small runnable scripts, one per concept. The run-folder format is versioned (format_version "1") and documented.

The screenshot above is real terminal output from my local pipeline on 2026-09-30: a completed run of a story-writing workflow with 131 steps done and 682 items not selected by a mid-run selection step. It shows the folder layout and the slowest steps from the run's own report. It is one personal pipeline, not a production deployment, and no performance numbers are claimed from it.

hone-flow was specified by me and built by AI coding agents working against written specifications and acceptance tests, with me reviewing their decisions; the repository says so in its README.

Known boundaries

Limitations and non-fit

  • No concurrency between steps, no dynamic fan-out of a step's list output into new items during a run, and no replay of code inside a step, in this version.
  • Not an orchestration platform: no always-on server, remote workers or scheduler. If you need those, use one of the platforms above.
  • Not an agent framework: no loops driven by model decisions.
  • External side effects are not protected from running twice; steps that pay, email or write to other systems need their own idempotency keys.
  • Alpha: the Python API may change before 1.0; the run-folder format is versioned.

Related patterns

Continue through the adjacent decision structures.

The move that makes the rest possible is small: write outputs first, metadata last, and treat the folder as the truth. Resume, fork, review and analysis all become reading and writing folders that a person can open and understand.

Adapt the pattern

Bring the problem, the boundary, and the consequence of being wrong.

Let's build something real