Bahman Shadmehr AI systems, independently Book a call

Independent · Remote, UTC+3 · Taking work now

You don't need an AI strategy. You need one process to actually work.

I'm an engineer, not an agency. I take one process that's costing you real hours, build the system that fixes it, prove it with numbers, and hand it to your team. Then we decide whether there's a second one.

Eight years building production systems — trading pipelines, payment infrastructure, Kubernetes — before any of this was called AI engineering.

The part nobody puts on their homepage

Almost every company that has tried this already has a graveyard of pilots. The research is unusually consistent about it, and unusually consistent about the cause.

95%

of enterprise generative-AI pilots delivered zero measurable return. Not low return — zero.

MIT NANDA, 300+ deployments

4 of 33

proofs of concept that made it to production, across the enterprises studied.

IDC with Lenovo

42%

of companies now abandon most AI initiatives — up from 17% a year earlier.

S&P Global

70%

of what makes AI work is process and people. Only 10% is the model itself.

BCG's 10–20–70 rule

Read the post-mortems and the same causes come back every time: no agreed definition of success before the build started, data nobody had prepared, no named owner, and nothing measuring quality afterwards. The model is almost never the problem. Which is lucky, because the model is the only part I can't help you with.

Why that keeps happening

A demo proves the model can do it. It proves nothing about whether you can run it.

Four gaps turn a good demo into a dead pilot. All four are ordinary engineering work, which is why they get skipped — they're not the interesting part.

01

Nobody wrote down what "correct" means

The pilot got judged on whether answers felt good in a room. With no test set, every later change becomes an argument instead of a measurement, and the project stops being able to improve.

02

The content was never prepared

Documents loaded whole, or split at arbitrary lengths. Retrieval returns something plausible from the wrong page, the model writes confidently around it, and after two bad answers in front of a customer nobody trusts it again.

03

It was built to demo, not to run

No permissions, no cost ceiling, no latency budget, no record of what was retrieved and why. Perfectly reasonable for a prototype, and impossible to put in front of a real user.

04

Nobody owns it on Monday

Documents change, models get deprecated, the questions people ask drift. Without an owner and monitoring, quality decays quietly and the first person to notice is the one it embarrassed.


Two minutes, five questions

Should you be doing this at all?

Genuinely — a lot of companies shouldn't, yet. This gives you my honest read, including the answers where I'd tell you not to hire me.

Fit check Question 1 of 5


What the work actually is

Four kinds of system. Yours is probably one of them.

Stripped of vocabulary, the job is narrow: turn a messy real process into something that reads your documents and data, gets the answer right, and plugs into the tools you already run.

Retrieval / RAG

Ask your own company a question

Policies, contracts, runbooks, closed tickets and wiki pages become searchable by meaning instead of keyword, and answers come back citing the paragraph they came from.

For example: a support team where the answer to most tickets is already written down, and the cost is finding it.

Classification

Sort what arrives, before a human does

Tickets, emails, applications, invoices — read, categorised, prioritised and routed, with a confidence score deciding what still needs a person.

For example: an ops team that spends the first hour of every day working out what the day is.

Extraction

Turn documents into data

Contracts, invoices, forms and PDFs into validated records in your database, with anything the model was unsure about flagged rather than quietly guessed.

For example: a finance process where someone retypes what already exists in a file.

Assisted workflow

Multi-step work with judgement in it

Draft this, check it against that, look up the account, escalate when the numbers disagree — with hard limits on what runs unsupervised and an audit trail of every step.

For example: a repeated internal job that takes a skilled person forty minutes and is mostly the same each time.

Whichever it is, an evaluation set comes first — a hundred real cases with agreed right answers, written with your team. It's the first thing I deliver, before any of the system exists, because it's the only thing that makes the rest measurable.


Prices, on the website, like a normal business

You shouldn't have to sit through a call to find out what this costs

Four shapes of engagement. Start at the left; most people should.

Start here

Assessment week

$3,500 fixed
  • One week, remote
  • Which process to do first
  • Whether your data can support it
  • Written report, risks ranked, costs estimated
  • A call to walk through it

Fee credited in full if you go ahead with a pilot.

Prove it

Pilot

from $7,500
  • 2–4 weeks
  • One process, end to end, on your real data
  • Evaluation set built with your team
  • A number, not a demo

Credited toward a build if you continue.

Ship it

Fixed-scope build

$15k – $45k
  • Written spec and acceptance criteria
  • Permissions, guardrails, cost and latency limits
  • Monitoring and deployment pipeline
  • Runbook and handover to your engineers

You know the number before we start.

Keep going

Fractional

from $5,000/mo
  • Or $800 per day
  • In your Slack, your tracker, your standups
  • Two clients at a time, maximum
  • 30-day rolling after the first month

For when the roadmap keeps producing this kind of work.

For comparison: a full-time senior AI engineer runs well past $150k a year fully loaded, plus several months of hiring risk and the real chance the role outlives the work. Prices exclude model and infrastructure costs, which you pay directly to your provider — I'd rather you see that bill than have me mark it up.


Saving us both a call

When I'm the wrong person

  • You want a strategy deck or a transformation roadmap. I build working systems. For board-level AI strategy, hire a firm that does that properly.
  • You want the cheapest possible demo. There are much faster and cheaper ways to get an impressive prototype. I'm the expensive option precisely because the demo isn't the deliverable.
  • Nothing is written down and nobody has time to change that. If the knowledge lives entirely in people's heads, retrieval has nothing to retrieve. That's a documentation problem first, and it isn't one I can solve for you.
  • You need ten people. I'm one engineer. Large multi-team programmes need an agency, and I'll happily say so.
  • You need someone on site. I work remotely, on UTC+3, overlapping European hours and US mornings.
  • You want someone to own it indefinitely. I build it, document it and hand it over. If nobody on your side will take it, it will rot, and we'd both rather know that now.

Why an independent

The person you meet is the person who writes the code

This is the whole argument for hiring one experienced engineer instead of a firm, and the whole argument against it. Both are on the table.

Working with me
  • One senior engineer, named, doing the work
  • Eight years of production systems that had to stay up
  • Small enough to turn down a bad fit
  • Prices published; no procurement theatre
  • Built to be handed over and left behind
  • Two clients at a time
What an agency gives you instead
  • A team, and cover when someone is ill
  • Capacity for several workstreams at once
  • Procurement, insurance and compliance paperwork
  • A support contract for the long term
  • Design, product and change management alongside
  • Someone to escalate to that isn't the engineer

If the right-hand column is what you need, that's a real answer and you should go and get it.


Where the confidence comes from

I spent eight years on systems that broke expensively when nobody was watching

An automated trading platform placing orders on the Texas wholesale power market. A US equities trading platform. A national payment gateway I helped break out of its monolith. Kubernetes clusters and data pipelines under all of it.

None of that work was really about writing code. It was about knowing what to do when something drifts: metrics on every stage, alerts tuned to catch an anomalous run rather than a merely slow one, and enough logging to explain afterwards what happened.

An LLM inside a business process is that same problem wearing new clothes. It runs on a schedule, it makes decisions, it fails without raising an exception, and its quality decays silently as the world moves. The people who find that obvious are the ones who have operated production systems.

2024 — 2025Automated energy trading, ERCOT — pipeline on GCP batch jobs, health metrics and alerting across every stage
2023 — 2024US equities trading platform — AWS services on the trading path, anomaly alerting on daily jobs
2022 — 2023Vgang — built and ran the Kubernetes cluster, integration microservices in Python and Go
2020 — 2021Zibal payment gateway — led the monolith-to-microservices migration, added Prometheus and Grafana
2017 — 2020Rahpa — async backends for real-time tracking, plus dashboards and mobile apps

Python, Go, Kubernetes, Docker, AWS, GCP, PostgreSQL, MongoDB, Redis, RabbitMQ, Celery, Prometheus, Grafana. I also write — 76 technical posts on DEV, a fair few on databases and retrieval.


Asked and answered

The questions that come up on every first call

Isn't this just n8n or Zapier with extra steps?

No, and if n8n solves your problem you should use n8n — it's cheaper than me and it'll be running this afternoon. Those tools move structured data between apps on rules you write in advance, and they're good at it.

My work starts where the rules run out: unstructured content, judgement in the middle of the process, and answers that have to be checked rather than assumed. That needs retrieval over your own material, a way to measure whether it's right, permissions so it can't surface what someone shouldn't see, and monitoring for when quality drifts. None of that is a node on a canvas.

How long until we see something real?

Week one you get an honest recommendation. Two to four weeks after that you get a working system on your own data with a measured accuracy number attached. Not a demo on curated examples — the real thing, on the real mess.

Full production hardening is usually another four to six weeks on top, and that's the part most projects never do.

What happens to our data?

Decided deliberately in week one and written into the agreement. Depending on what you need: enterprise API tiers with no-training guarantees, a specified cloud region, self-hosted embeddings and vector storage inside your own infrastructure, or fully local models where the data genuinely cannot leave.

Access control matters as much as storage. Retrieval has to respect the permissions your documents already carry, or you've quietly built a way to read things people couldn't otherwise open.

Do we need to fine-tune something?

Almost certainly not, and definitely not first. In most business processes the model doesn't lack capability, it lacks your information — and that's retrieval, which is faster, cheaper and far easier to keep current than training.

Fine-tuning earns its place later, for consistent formatting or a narrow high-volume task. It's an optimisation, not a starting point.

Won't this be obsolete when the next model ships?

The models will change, which is exactly why I build so they can be swapped in an afternoon. What doesn't go obsolete is the rest: knowing which process was worth doing, the prepared data, the definition of correct, and the evaluation harness that lets you change a model on Tuesday and know by Wednesday whether it helped.

That harness is the asset. Everything else is deliberately replaceable.

You're one person. What if you get hit by a bus?

Fair, and it's why handover is a phase rather than an afterthought. Everything runs in your infrastructure, in your repositories, with a runbook and an evaluation harness your engineers can run without me. No proprietary layer, no hosted black box, nothing that stops working if I disappear.

If that isn't reassurance enough — and for some organisations it reasonably isn't — an agency is the right call.

One email is enough to start

Name one process that costs more hours than it should

That's the whole first conversation. Thirty minutes, and I'll tell you straight whether it's worth doing — I lose nothing by saying no, and you lose a lot by hearing yes from someone who shouldn't have said it.

Put this in the email

  1. The process, and roughly what it costs you a week.
  2. Where the knowledge lives — wiki, tickets, drive, someone's head.
  3. Whether your data can leave your infrastructure.
  4. Who would own it afterwards.