Five levels. Most companies should stop at two.
Everything I build fits on one ladder, and each rung genuinely requires the one below it. This page exists as much to tell you where to stop as where to start — because the expensive mistake in this field isn't going too slowly, it's buying level four while standing on level zero.
This is a ladder, not a treadmill. Two thirds of the companies I speak to get their entire return from levels one and two. The upper rungs are real, and they are for a specific kind of business rather than a more ambitious one.
—
——
The ordering isn't a sales sequence, it's a dependency chain. Level three needs agreed definitions from level two. Level four needs the history that levels one to three generate. Skipping a rung doesn't make you faster — it makes the rung above unreliable in ways nobody notices for months.
Why the order matters more than the ambition
Every failed AI programme I've looked at made the same structural error: it bought a capability whose prerequisite it didn't have.
An advisor that flags "churn risk is rising" is worthless if three departments define churn differently — it will be confidently wrong every morning, and the first person to check will stop trusting all of it. Definitions are not paperwork. They're the input the higher rungs run on.
Agreeing metrics in the abstract is a workshop nobody finishes. Agreeing them because a working system gave two teams different numbers last week takes an afternoon. A level-one project is what makes the level-two conversation possible.
You cannot automate a process that three people describe differently. Mapping it is the first task of any level-one project, and it regularly produces a fix that needs no AI at all. That's a good outcome, and I'll tell you when it happens.
Letting a system act without review requires evidence it was right for months while a human checked. That evidence is a by-product of running levels one to three properly. Nobody should sell you level four on day one, and I won't.
Six good reasons to stop climbing
Each of these has come up with real prospects, and in each case stopping was the right answer.
- Level two already covers your questions. If the answers people need are now available in seconds, the remaining rungs solve problems you don't have. Spend the money on the business instead.
- Your processes change faster than they can be encoded. Early-stage companies often shouldn't automate anything — the process you'd automate this quarter won't exist next quarter, and you'd be buying rigidity.
- Nobody will own the next rung. Every level needs someone internal whose job includes it. Without a name, the thing decays and you've bought a liability with a maintenance bill.
- The remaining processes are low volume. Once you've done the two or three that run constantly, what's left often runs weekly. Annoying isn't the same as expensive.
- The data for the next level genuinely isn't there. Level three needs twelve months of consistent history. If your systems changed nine months ago, wait — building on a short baseline produces confident nonsense.
- You've got what you came for. The point was never to reach level four. It was to stop losing hours to work a machine could do, and that goal can be fully met at rung one.
I take two clients at a time, which means I'd rather do one good level-one project than a mediocre programme. A client who stops at two and tells someone why is worth more to me than one who climbs reluctantly.
What each rung costs, and what it typically returns
Fixed fees, published. Return figures are typical ranges for a mid-market company, not promises — the level-one business case is worked out in full, assumption by assumption, in a separate document.
| Level | What gets built | Investment | Typical annual return | Payback |
|---|---|---|---|---|
| 0 | Process mappingMeasure what you actually do, including the waiting | $4,500 | — | n/a |
| 1 | One process compressedDocuments read, work sorted, drafts prepared, handoffs removed | $18k – $40k | $45k – $95k | 11–19 mo |
| 2 | Definitions & answersGoverned metric layer, ask-anything over your own data and documents | $26k – $65k | $40k – $120k | 9–18 mo |
| 3 | Daily briefObservations from your systems, with evidence and a scorecard | $30k – $55k | hard to quantify | judgement |
| 4 | Bounded autonomyClear cases handled end to end; everything else escalated | scoped | step change | varies |
Note level three has no payback figure. Its value is avoided surprises, which cannot honestly be modelled in advance — you'd be pricing things that didn't happen. It's a judgement call, and the useful-rate threshold in that product exists so you can stop paying if the judgement was wrong.
One rung at a time, with a decision point between each
There is no programme, no roadmap and no multi-year commitment. Each level is a separate decision made with the evidence the previous one produced.
A short paid assessment
One to two weeks, fixed fee, credited against the build. It measures whether the prerequisite is genuinely in place — and it can conclude "not yet," which happens often enough that it's priced separately.
Fixed scope, agreed bar
Price fixed after the assessment, accuracy bar set by you on your own test cases, and everything built in your infrastructure and repositories.
An explicit stop point
We look at whether the return arrived and whether the next rung is worth it. Stopping is a normal answer and I'd rather you stop satisfied than continue reluctantly.
What people ask about the ladder
Isn't this just a way to sell us five projects instead of one?
It would be, if the rungs weren't genuinely dependent — and if the page didn't spend a whole section telling you when to stop. Two thirds of companies get their full return from levels one and two, and I'd rather say that here than discover it with you in month eighteen.
The practical protection is structural: each level is separately scoped, separately priced, and preceded by an assessment that can conclude no. There's no programme to sign and nothing that penalises you for stopping.
Can we skip straight to level three or four?
You can buy it. It won't work, and the failure is delayed rather than immediate, which is what makes it expensive. A daily brief without agreed definitions produces confident wrong statements every morning; autonomy without a measured track record is a governance problem waiting to be discovered by someone senior.
If you already have the prerequisites from other work — a real semantic layer, a mature data function, documented processes — then you're further up the ladder than you think and we start there. The assessment establishes which.
We've already done some of this ourselves. Where does that leave us?
Usually further along than the org chart suggests, and that's a good position. Companies with a working dbt metrics layer, or an existing automation practice, frequently sit at level two already without calling it that.
It also occasionally means the opposite: a lot of level-one automation built without measurement, which looks like progress and isn't. The assessment is honest about both.
How long does the whole thing take?
If you did all four rungs consecutively, roughly eighteen months to two years with the gaps between them — and I'd be surprised and slightly suspicious if a company moved faster than that.
But that's not the useful question. The useful question is what level one costs and returns, because that decision stands entirely on its own.
What if the models change and this is all obsolete?
The models will change, which is exactly why the durable assets are the ones on the lower rungs: mapped processes, agreed definitions, evaluation sets, and a documented track record. Those survive a model swap entirely.
What gets replaced is the layer that's designed to be replaceable. Swapping a model on an established evaluation harness is an afternoon; doing it without one is a rebuild.
You're one person. Can you deliver a whole ladder?
Over eighteen months, for two clients at a time, yes — that's the shape of the practice. But it's a fair challenge and the honest answer is that you shouldn't plan around it. Each rung is built to be owned by your team afterwards, in your infrastructure, with the documentation to run it without me.
If the answer to "what if he's unavailable" needs to be an SLA and a bench of engineers, an agency is the right choice and I'd say so early.
Eight years of production systems, which is where the ordering comes from
I'm Bahman Shadmehr. Automated trading platforms on the Texas wholesale power market and US equities, a national payment gateway I helped decompose while it was live, and the pipelines, warehouses and monitoring underneath all of it.
The ladder isn't a framework I invented for a website. It's the order those systems had to be built in: measure before you change, define before you compute, observe before you alert, and prove before you let anything run unattended. Skipping a step never made a project faster — it made the failure arrive later and cost more.
You deal with me from first call to handover, on every rung.
Python, Go, Kubernetes, Docker, AWS, GCP, Postgres, MongoDB, Redis, Prometheus. Remote, UTC+3. 76 technical posts · LinkedIn
Tell me what you've already tried, including what didn't work
That places you on the ladder faster than any assessment questionnaire. I'll come back with which rung you're actually on and what the next one would cost — including when the answer is that you're already where you should be.
info@bshadmehr.meUseful in a first message
- What you've already automated or tried.
- Whether your core metrics are defined anywhere.
- How documented your processes are, honestly.
- The systems you run and how much history is in them.
- Who would own any of this internally.