Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer
On this pattern

System Pattern 03 / Operations, Privacy & Security

Refuse the call that can't work.

Capability-Checked Model Calls

Put every model call, local or hosted, behind one small interface that checks the request fits the model before sending it, reports exactly how structured output was obtained, shares one GPU without crashes, and records every call.

Maturity
Tested / Prototype
Critical uncertainty
A capability registry can refuse calls that cannot work; it cannot make a model's answer correct. Answer quality still needs evaluation.
Human boundary
Choose models, registry overrides, whether repaired output is acceptable
Applications
Local LLMs on one GPU · Mixed local and hosted models · Structured output from small models · Vision, speech and embedding calls
Terminal output of hone-models calls stats, captured 2026-09-30 in my local pipeline: calls, errors, input and output tokens and mean duration for eleven local models.

Problem shape

The problem beneath them

Route every model call through one interface backed by a registry of model capabilities, so impossible requests are refused before sending, structured answers say how they were obtained, answer problems come back as values and call problems as typed errors, GPU memory is leased, and each call is recorded for replay.

Why I built it

hone-models came out of the same local pipeline as the other honeworks packages: song ideas, lyrics, music and video from Ollama models on an 8 GB GPU. The failures repeated. Ollama's default context window (2,048 tokens) silently cut a long JSON prompt, so the model answered a question it never fully saw. JSON came back in markdown fences, cut off at the output limit, or ignored the schema because format had been put in the wrong place. A thinking model returned empty content with all the text in its thinking field, and an empty string looked like a valid answer. A fixed 120-second timeout failed the slow models and made the fast ones wait. Ollama, ComfyUI and in-process models fought over the GPU. And nobody could say afterwards which prompt, model and parameters had produced a given output.

The fix isn't a smarter model. It's a small layer that knows each model's facts (vision, thinking, JSON schema support, logprobs, context size, VRAM, licence) and uses them before and after every call.

A capability registry can refuse calls that cannot work; it cannot make a model's answer correct. Answer quality still needs evaluation.

hone-models is not an agent framework, a prompt-management UI, a hosted gateway or an evaluation tool. It wraps LiteLLM for hosted coverage rather than competing with it.

Different problems, same shape

The surface changes. The decision structure persists.

01 / Long prompts

The model answered confidently. Did it ever see the end of the prompt?

The server's default context window silently cut the input.

02 / Structured output

Was this JSON constrained, parsed, repaired or guessed?

Each call site has its own `_extract_json` helper.

03 / Model choice

Why did the vision judge fail deep inside the run?

A text-only model was configured where images are sent.

04 / Shared GPU

Why does the pipeline crash only when the image model runs first?

Several model servers load onto one card without coordination.

All four questions converge into Capability-Checked Model Calls.

Exploded pattern

Open the mechanism at every decision boundary.

  1. 01

    Registry

    Question
    What can this model actually do?
    Responsibility
    Hold one entry per model: provider, tag, context size, vision, thinking, schema support, logprobs, VRAM, licence; overrides per machine in a TOML file.
    Input
    The built-in catalog plus `models.toml` overrides and environment settings.
    Output
    A resolved model entry for a registry ID such as `gemma4-12b`.
    Stops when
    The ID resolves, or the call fails with the unknown ID named.
  2. 02

    Pre-flight check

    Question
    Can this request work on this model?
    Responsibility
    Count the prompt against the context budget (images included) and check capabilities.
    Input
    Messages, attachments, schema, model entry.
    Output
    A request to send, or `ContextOverflow` with the numbers, or `CapabilityError`.
    Stops when
    The request is known to fit or is refused.
  3. 03

    Structured output

    Question
    How do we get valid structured data, and how did we get it?
    Responsibility
    Constrain where the server supports it, parse, validate, retry with the validation error, repair as a last resort; never hide a reply cut at the output limit.
    Input
    A pydantic schema and the model's reply.
    Output
    `parsed` value and `structured_path` (constrained, parsed, retried, repaired or failed), or `result.error`.
    Stops when
    A valid object exists or the attempts are spent.
  4. 04

    Decision questions

    Question
    Is this a yes/no, choice or score question?
    Responsibility
    Send it to a native decision model or emulate it on any LLM; mark probabilities calibrated only when they come from token logprobs.
    Input
    Question, options, model.
    Output
    Answer with probability and whether it is calibrated.
    Stops when
    Answered, or the answer problem is returned as a value.
  5. 05

    GPU lease

    Question
    Who may use the GPU now?
    Responsibility
    Grant leases from a cross-process ledger guarded by a file lock; unload other models when allowed; let non-LLM GPU work (Whisper, diffusion) take turns.
    Input
    Requested VRAM and the machine state.
    Output
    A lease, a wait with a stall limit, or an error for a lease that can never be granted.
    Stops when
    The lease is granted, released or refused.
  6. 06

    Record and replay

    Question
    What exactly happened in this call, and what if one thing changes?
    Responsibility
    Write an OpenTelemetry-shaped span to local SQLite with model, parameters, usage, timing, structured path and named prompt sections; replay any recorded call with one change.
    Input
    Call, result, trace context.
    Output
    A span in `.hone/models/spans.db`; replay results as new spans.
    Stops when
    The span is written.

Decision forks

Every branch states why it exists and when it escalates.

Signals, decisions, reasons, and escalation conditions
Signal Decision Reason Escalates when
Prompt exceeds context budget Raise `ContextOverflow` with numbers A truncated prompt produces confident nonsense The caller needs to shorten or switch model
Image sent to a text-only model Raise `CapabilityError` before sending A 400 deep in a run is worse Registry entry is wrong
Reply cut at output limit Report it, never return as complete Truncated JSON can still parse Output limit is too small for the task
Schema fails validation Retry with the error, then repair Small models often fix their own mistakes `structured_path` is `failed`
Thinking model returns empty content Treat as an answer problem Empty is not an answer Thinking flags misconfigured
GPU memory short Wait, or unload other models if allowed Crashing is worse than waiting Stall limit reached

Operating paths

Clear, ambiguous, and failed work all reach explicit states.

Clear path

Resolve the model, check fit, send, get a constrained answer, record.

Final state
Answer with `structured_path = constrained`.
Owner
The calling code.
Evidence
Span with model, usage, timing and prompt sections.
Recovery
Replay with a changed model or parameter.
Ambiguous path

Parse fails, retry with the validation error, repair if still invalid.

Final state
Answer with `structured_path` retried or repaired.
Owner
The calling code, which decides whether a repaired answer is acceptable.
Evidence
Attempts and the path taken.
Recovery
Tighten the schema, raise the output limit, or pick a model with native schema support.
Failure path

Request refused before sending, or answer unusable after all attempts.

Final state
Typed error (call problem) or `result.error` (answer problem).
Owner
The calling code.
Evidence
Error with the numbers or the failed attempts.
Recovery
Change model, shorten input, or route the item to a person.

Authority map

Capability does not grant authority.

RULE

May decide
Refuse requests that don't fit the registry facts
May not decide
Accept a request by guessing capabilities
Required evidence
Registry entry and computed budget

MODEL

May decide
Produce the answer
May not decide
Report its own answer as valid structure
Required evidence
Raw reply, structured path

SYSTEM

May decide
Validate, retry, repair, lease the GPU, record
May not decide
Hide truncation or a repair
Required evidence
Span per call and per attempt

HUMAN

May decide
Choose models, registry overrides, whether repaired output is acceptable
May not decide
Be surprised by silent truncation
Required evidence
Registry file, call records

EXCEPTION

May decide
Return typed errors and answer errors
May not decide
Return an empty string as success
Required evidence
Error type and message

Failure modes and recovery

A failed path remains owned, evidenced, and recoverable.

Failures, detection, containment, recovery, and owners
Failure Detection Containment Recovery Owner
Context overflow Pre-flight count Call not sent Shorten or switch model Caller
Truncated reply Finish reason Not returned as complete Raise output limit Caller
Invalid JSON Validation Retry with error Repair or fail visibly Caller
Model server down Connection error Typed error Retry later or fail over explicitly Operator
GPU contention Ledger wait No second load Lease granted or stall error Operator
Slow model timeout Speed-aware timeout Call bounded Adjust timeout from measured speed Operator

Invariants and guarantees

Properties the structure is designed to preserve.

  • A request that doesn't fit the model is refused before it's sent.
  • A reply cut at the output limit is always reported.
  • structured_path says how every structured answer was obtained.
  • Answer problems are values; call problems are exceptions.
  • Probabilities are marked calibrated only when they are.
  • Every call leaves a span that can be replayed.

What changes between implementations

The constraints determine the final mechanism.

The providers change: Ollama and any OpenAI-compatible server (llama.cpp, vLLM, LM Studio, OpenAI) are built in; LiteLLM adds hosted providers. The machine changes: a single GPU needs leases and unloads; a server with many GPUs needs fewer. The media change: the same interface covers embeddings, text to speech, transcription with word timestamps, and image, music and video generation through ComfyUI workflows, each with its own registry entry and guide.

Open-source alternatives I compared

Tool What it does well Why it didn't fit this job
LiteLLM One API for many hosted providers; large model registry No local model lifecycle, GPU memory or decision questions; hone-models uses it as an optional provider
Instructor, PydanticAI Validation and retries No context budgets, truncation detection or local-model quirks
llama-swap Swaps local model servers Knows nothing about the calls going through them
LangChain Very broad abstraction surface I wanted a small, explicit one

Evidence chain

Follow the pattern into systems and software.

Implemented in hone-models (Apache-2.0, alpha) with acceptance tests in CI and real-model tests run through a GPU lock. The screenshot shows its call records from my local pipeline; no production result is claimed.

The on-premises call review design needs exactly this layer: local models that must refuse overflowing transcripts instead of truncating them, and GPU sharing between transcription and a language model. The small-model marketplace design needs structured output from a small model with a record of how it was obtained. Both are reference designs.

Implemented in hone-models, Apache-2.0, alpha. The core needs only pydantic and httpx. The repository holds the design, its acceptance cases and seventeen recorded design changes, from replay and GPU leases to speech, generation models and machine state. Real-model tests run through a machine-wide GPU lock and are separate from the default offline suite.

The screenshot is real output of hone-models calls stats from my local pipeline on 2026-09-30: eleven local models, their call counts, errors, tokens and mean durations. It shows what the records make visible (one model with 7 errors in 89 calls, for instance). It is a personal machine, and these durations are not benchmarks.

Known boundaries

Limitations and non-fit

  • No streaming or async clients in this version.
  • The registry is only as good as its entries; a wrong capability means wrong refusals or missed ones.
  • It records and checks calls; it doesn't evaluate answer quality.
  • Some model entries carry licences with restrictions; the registry records them, it doesn't clear them for you.
  • If you only call one hosted API and never run local models, a provider SDK plus a validation library may be enough.

Related patterns

Continue through the adjacent decision structures.

Most model bugs in a pipeline are not model bugs. They're missing facts: how long the context really is, whether the answer was cut, which model got the image. Writing those facts down once and checking them on every call turns silent failures into loud ones.

Adapt the pattern

Bring the problem, the boundary, and the consequence of being wrong.

Let's build something real