Registry
One entry per model with its capabilities, context size, VRAM and licence.
Why. Decisions about a call need facts about the model.
Boundary. Wrong entries mean wrong refusals.
Model access
One interface for local and hosted AI models, with structured output, GPU scheduling and full call records.
Prevents: Prompts silently cut at the context limit, JSON rescued in different ways at every call site, empty replies from thinking models, and models crashing each other on one GPU.
Why it exists
Calling models reliably from ordinary software, especially on your own hardware.
Scope. A client library and CLI; not an agent framework, gateway or evaluation tool.
Five-minute orientation
pip install "git+https://github.com/honeworks/hone-models"
Real output
Architecture
A registry of model facts drives checks before each call; providers for Ollama, OpenAI-compatible servers and LiteLLM share one result shape; a file-locked ledger grants GPU leases; spans go to local SQLite.
Core concepts
Operations
`hone-models models` and `hone-models calls list / show / stats`.
Current boundary
Alpha. No streaming or async clients. Checks can't make an answer correct.
Near-term roadmap. Follow the design history in design/changes/.
What it implements
Engineering checklist
Need the control, not just the component?
The repository exposes the mechanism. Production work is defining the permissions, data, failure costs, evidence, and owners around it.