Problem shape
The problem beneath them
Route every model call through one interface backed by a registry of model capabilities, so impossible requests are refused before sending, structured answers say how they were obtained, answer problems come back as values and call problems as typed errors, GPU memory is leased, and each call is recorded for replay.
Why I built it
hone-models came out of the same local pipeline as the other honeworks packages: song ideas, lyrics, music and video from Ollama models on an 8 GB GPU. The failures repeated. Ollama's default context window (2,048 tokens) silently cut a long JSON prompt, so the model answered a question it never fully saw. JSON came back in markdown fences, cut off at the output limit, or ignored the schema because format had been put in the wrong place. A thinking model returned empty content with all the text in its thinking field, and an empty string looked like a valid answer. A fixed 120-second timeout failed the slow models and made the fast ones wait. Ollama, ComfyUI and in-process models fought over the GPU. And nobody could say afterwards which prompt, model and parameters had produced a given output.
The fix isn't a smarter model. It's a small layer that knows each model's facts (vision, thinking, JSON schema support, logprobs, context size, VRAM, licence) and uses them before and after every call.
A capability registry can refuse calls that cannot work; it cannot make a model's answer correct. Answer quality still needs evaluation.
hone-models is not an agent framework, a prompt-management UI, a hosted gateway or an evaluation tool. It wraps LiteLLM for hosted coverage rather than competing with it.
Different problems, same shape
The surface changes. The decision structure persists.
The model answered confidently. Did it ever see the end of the prompt?
The server's default context window silently cut the input.
Was this JSON constrained, parsed, repaired or guessed?
Each call site has its own `_extract_json` helper.
Why did the vision judge fail deep inside the run?
A text-only model was configured where images are sent.
Why does the pipeline crash only when the image model runs first?
Several model servers load onto one card without coordination.
All four questions converge into Capability-Checked Model Calls.
Exploded pattern
Open the mechanism at every decision boundary.
-
01
Registry
- Question
- What can this model actually do?
- Responsibility
- Hold one entry per model: provider, tag, context size, vision, thinking, schema support, logprobs, VRAM, licence; overrides per machine in a TOML file.
- Input
- The built-in catalog plus `models.toml` overrides and environment settings.
- Output
- A resolved model entry for a registry ID such as `gemma4-12b`.
- Stops when
- The ID resolves, or the call fails with the unknown ID named.
-
02
Pre-flight check
- Question
- Can this request work on this model?
- Responsibility
- Count the prompt against the context budget (images included) and check capabilities.
- Input
- Messages, attachments, schema, model entry.
- Output
- A request to send, or `ContextOverflow` with the numbers, or `CapabilityError`.
- Stops when
- The request is known to fit or is refused.
-
03
Structured output
- Question
- How do we get valid structured data, and how did we get it?
- Responsibility
- Constrain where the server supports it, parse, validate, retry with the validation error, repair as a last resort; never hide a reply cut at the output limit.
- Input
- A pydantic schema and the model's reply.
- Output
- `parsed` value and `structured_path` (constrained, parsed, retried, repaired or failed), or `result.error`.
- Stops when
- A valid object exists or the attempts are spent.
-
04
Decision questions
- Question
- Is this a yes/no, choice or score question?
- Responsibility
- Send it to a native decision model or emulate it on any LLM; mark probabilities calibrated only when they come from token logprobs.
- Input
- Question, options, model.
- Output
- Answer with probability and whether it is calibrated.
- Stops when
- Answered, or the answer problem is returned as a value.
-
05
GPU lease
- Question
- Who may use the GPU now?
- Responsibility
- Grant leases from a cross-process ledger guarded by a file lock; unload other models when allowed; let non-LLM GPU work (Whisper, diffusion) take turns.
- Input
- Requested VRAM and the machine state.
- Output
- A lease, a wait with a stall limit, or an error for a lease that can never be granted.
- Stops when
- The lease is granted, released or refused.
-
06
Record and replay
- Question
- What exactly happened in this call, and what if one thing changes?
- Responsibility
- Write an OpenTelemetry-shaped span to local SQLite with model, parameters, usage, timing, structured path and named prompt sections; replay any recorded call with one change.
- Input
- Call, result, trace context.
- Output
- A span in `.hone/models/spans.db`; replay results as new spans.
- Stops when
- The span is written.
Decision forks
Every branch states why it exists and when it escalates.
| Signal | Decision | Reason | Escalates when |
|---|---|---|---|
| Prompt exceeds context budget | Raise `ContextOverflow` with numbers | A truncated prompt produces confident nonsense | The caller needs to shorten or switch model |
| Image sent to a text-only model | Raise `CapabilityError` before sending | A 400 deep in a run is worse | Registry entry is wrong |
| Reply cut at output limit | Report it, never return as complete | Truncated JSON can still parse | Output limit is too small for the task |
| Schema fails validation | Retry with the error, then repair | Small models often fix their own mistakes | `structured_path` is `failed` |
| Thinking model returns empty content | Treat as an answer problem | Empty is not an answer | Thinking flags misconfigured |
| GPU memory short | Wait, or unload other models if allowed | Crashing is worse than waiting | Stall limit reached |
Operating paths