AI Providers

Hefty supports three providers out of the box. You choose your provider in the AI Settings dialog — accessible on first launch or any time via the admin menu. Multiple provider entries can be configured at once (the first is the primary, the rest act as ordered fallbacks); see AI Settings for details.

Ollama Local

Free, local inference. No API key needed. Supports model listing, embeddings, and automatic model downloads. Recommended for privacy-first setups.

Default URL: localhost:11434
Recommended models: gemma4:e4b (9.6 GB, 128K context), nemotron-3-nano:4b (2.8 GB, 256K)
Default embedding model: nomic-embed-text
OpenAI compatible Cloud

Any service that speaks the OpenAI chat completions protocol — OpenAI, OpenRouter, Groq, or a local server such as LM Studio or vLLM. Requires an API key for cloud services. Supports embeddings where the service offers them.

Default URL: openrouter.ai/api
Local servers: point the entry's base URL at your local endpoint (e.g. localhost:1234/v1 for LM Studio)
Hefty platform Cloud

The managed gateway operated by Hefty — usable with a Hefty platform API key. The gateway address is fixed and cannot be changed; you configure the key and go. Ships with a default text model and a matching embedding model.

URL: gate1.llm.hefty.bot (fixed)
Default text model: qwen3.8-27b-nvfp4
Default embedding model: qwen3-vl-embedding-2b-awq
Upgrading from an older version?

Earlier Hefty releases offered additional provider choices (OpenAI, LM Studio, NVIDIA, Hugging Face, SGLang, vLLM, Gemini, and others). Existing entries are migrated automatically when your configuration loads: anything that isn't Ollama or the Hefty platform becomes OpenAI compatible, keeping its base URL, model, and API key. Note that Anthropic's native API is no longer supported — use a service or proxy that exposes an OpenAI-compatible endpoint instead.

Embedding Provider

Hefty uses a separate embedding model for its knowledge system (vector search). All three providers support embeddings, and the embedding provider can be the same as or different from your text model provider — for example, a cloud text model paired with a local Ollama instance running nomic-embed-text, or the Hefty platform gateway for both.