Skip to content
Bahman Shadmehr Independent AI Systems & Automation Engineer

Model access

Refuse the model call that can't work, and record the ones that do.

One interface for local and hosted AI models, with structured output, GPU scheduling and full call records.

Prevents: Prompts silently cut at the context limit, JSON rescued in different ways at every call site, empty replies from thinking models, and models crashing each other on one GPU.

Why it exists

The failure appears after the happy path.

Calling models reliably from ordinary software, especially on your own hardware.

Scope. A client library and CLI; not an agent framework, gateway or evaluation tool.

Five-minute orientation

See the smallest complete path.

pip install "git+https://github.com/honeworks/hone-models"

Real output

What it looks like running.

Terminal: hone-models calls stats lists eleven local models with calls, errors, tokens and mean duration.
Real output of hone-models calls stats from my local pipeline on 2026-09-30. Durations are from a personal machine, not benchmarks.

Architecture

The architecture is a set of promises.

A registry of model facts drives checks before each call; providers for Ollama, OpenAI-compatible servers and LiteLLM share one result shape; a file-locked ledger grants GPU leases; spans go to local SQLite.

  • A request that doesn't fit the model is refused before it's sent. Inspect test
  • A reply cut at the output limit is always reported. Design note
  • structured_path says how each structured answer was obtained. Inspect test
  • Every call leaves a span that can be replayed. Inspect test

Core concepts

The few concepts you need before reading the code.

Registry

One entry per model with its capabilities, context size, VRAM and licence.

Why. Decisions about a call need facts about the model.

Boundary. Wrong entries mean wrong refusals.

GPU lease

Permission to use VRAM, granted from a ledger shared by processes.

Why. Several model servers on one card crash without coordination.

Boundary. Callers wrap GPU work in a lease; it isn't automatic for every call.

Operations

What happens after install.

`hone-models models` and `hone-models calls list / show / stats`.

Current boundary

What this project does not solve.

Alpha. No streaming or async clients. Checks can't make an answer correct.

Near-term roadmap. Follow the design history in design/changes/.

What it implements

  • capability registry and model catalog
  • pre-flight context and capability checks
  • structured output with a reported path
  • decision questions
  • GPU leases across processes
  • embeddings, speech, transcription and generation models
  • call records and replay

Engineering checklist

What the repository ships.

Quickstart
Present
Success tests
Present
Failure tests
Present
Architecture notes
Present
Security notes
Not published
Runbook
Not published
Limitations
Present
Changelog
Present

Recent changes

What changed, with the commits.

  • Generation models, model guides and catalog, machine state (0015, 0016) Commit
  • Model endpoints from the environment and a .env file (0017) Commit
  • Two flaky tests fixed (SQLite crash test, AC-22 span order) Commit

Need the control, not just the component?

Fit it to the system that has to survive.

The repository exposes the mechanism. Production work is defining the permissions, data, failure costs, evidence, and owners around it.

Let's build something real