AI Providers
Hefty supports three providers out of the box. You choose your provider in the AI Settings dialog — accessible on first launch or any time via the admin menu. Multiple provider entries can be configured at once (the first is the primary, the rest act as ordered fallbacks); see AI Settings for details.
Free, local inference. No API key needed. Supports model listing, embeddings, and automatic model downloads. Recommended for privacy-first setups.
Any service that speaks the OpenAI chat completions protocol — OpenAI, OpenRouter, Groq, or a local server such as LM Studio or vLLM. Requires an API key for cloud services. Supports embeddings where the service offers them.
The managed gateway operated by Hefty — usable with a Hefty platform API key. The gateway address is fixed and cannot be changed; you configure the key and go. Ships with a default text model and a matching embedding model.
Earlier Hefty releases offered additional provider choices (OpenAI, LM Studio, NVIDIA, Hugging Face, SGLang, vLLM, Gemini, and others). Existing entries are migrated automatically when your configuration loads: anything that isn't Ollama or the Hefty platform becomes OpenAI compatible, keeping its base URL, model, and API key. Note that Anthropic's native API is no longer supported — use a service or proxy that exposes an OpenAI-compatible endpoint instead.
Embedding Provider
Hefty uses a separate embedding model for its knowledge system (vector search). All three providers support embeddings, and the embedding provider can be the same as or different from your text model provider — for example, a cloud text model paired with a local Ollama instance running nomic-embed-text, or the Hefty platform gateway for both.