Advanced
Most configuration is done via the in-app AI Settings dialog and persisted to hefty.conf in your data directory. The environment variables below are deployment-level settings; the reference tables further down cover the AI settings keys the dialog writes.
Environment Variables Reference
Server
HEFTY_PORT- port the desktop app uses to reach the local server (GUI + API). Default:41890HEFTY_DATA_DIR- data storage location; yourhefty.confand everything Hefty remembers live here. Default:~/.heftyHEFTY_HEADLESS- whentrue, the server runs in API-only mode with no GUI routesHEFTY_VEC_EXTENSION_PATH- path to thesqlite-vecextension. Auto-detected in the desktop app.
Provider Setup Details
Ollama (Default)
Ollama runs on port 11434 by default, which matches Hefty's default configuration. No extra setup needed beyond pulling the models — Hefty can also download its recommended models for you during setup.
OpenAI Compatible
Point the entry's base URL at any OpenAI-compatible endpoint — OpenAI, OpenRouter, Groq, or a local server like LM Studio or vLLM all work. Add the service's API key, then pick a model from the ones the service offers or enter a custom model name.
Hefty Platform
Enter your Hefty platform API key. The gateway URL is fixed and cannot be changed; default models are provided by the platform.
hefty.conf AI Settings Reference
The AI Settings dialog writes these keys to hefty.conf on save. You can also hand-edit the file while Hefty is stopped — changes take effect at the next start. Per-entry API keys are stored by the dialog when you save; treat your hefty.conf as sensitive.
llm-providers[]
An ordered list — the first entry is the primary provider, the rest are fallbacks tried in order when one fails.
| Key | Default | Description |
|---|---|---|
provider | ollama | One of ollama, openai-compatible, or hefty. Unknown legacy ids are mapped automatically. |
base-url | http://localhost:11434 | The provider's endpoint. Managed gateway URL is applied for the Hefty platform. |
text-model | gemma4:e4b | The reasoning model to use. |
keep-alive-minutes | 15 | How long the model stays loaded in memory. |
stream-max-tokens | 16384 | Completion-token cap for streaming responses. Non-positive values fall back to the default. |
context-token-limit | 32000 | Maximum context window size. |
max-concurrent-requests | unset | Parallel request cap for this entry (1–64, out-of-range values are clamped with a warning). Applies across all users; when unset, the global priority-queue.max-concurrent and stream.max-concurrent-streams limits apply. |
think | unset | Tri-state extended thinking switch: true (on), false (off), unset (provider/model default). |
reasoning-effort | medium | Thinking budget for the model: low, medium, or high. Only sent while thinking is enabled. |
embedding-providers[]
| Key | Default | Description |
|---|---|---|
provider | ollama | Same three providers as the LLM list. |
base-url | http://localhost:11434 | The embedding provider's endpoint. |
embedding-model | nomic-embed-text | The embedding model to use. |
embedding-dimension | derived | Explicit vector dimension; derived from the model when unset. |
context-token-limit | 2048 | How much text can be embedded per request. |
Global AI settings keys
| Key | Default | Description |
|---|---|---|
max-reflection-loops | 10 | Maximum reasoning iterations per task. |
circuit-breaker.max-failures | 5 | Consecutive failures before the provider is tripped. |
circuit-breaker.reset-timeout | 30s | How long before a tripped provider is retried. |
priority-queue
Controls the shared queue that admits LLM calls. max-concurrent bounds all LLM calls dispatched in parallel — streaming and non-streaming alike (embeddings, titles, one-shot completions) — and serves as the global fallback for provider entries that don't set their own max-concurrent-requests. The aging settings promote requests that have waited a long time so low-priority work is never starved: a waiting request's effective priority grows by aging-step points every aging-interval-ms, so with the defaults a priority-1 request overtakes freshly enqueued priority-100 requests after roughly 25 minutes of waiting.
| Key | Default | Description |
|---|---|---|
enabled | true | Turn the priority queue on or off. |
max-concurrent | 3 | Global parallel cap for all LLM calls. Set to 1 on gateway tiers that count every completion against the same per-key in-flight cap; higher suits local (Ollama) and multi-provider setups. |
max-queue-depth | 100 | Maximum number of waiting requests. |
aging-interval-ms | 15000 | How often a waiting request's priority bonus accrues. |
aging-step | 1 | Priority points added per aging interval. |
stream
Streaming admission control. A stream holds an in-flight slot at the provider for its entire duration — minutes, not seconds — so concurrent streams from your instance serialize locally: up to max-concurrent-streams run at once, and the excess wait in a queue bounded by queue-timeout-ms before failing with a typed admission timeout (instead of contending at the gateway). Match max-concurrent-streams to your tier's in-flight cap; on local (Ollama) or other no-slot-cap deployments, raise it so parallel conversations aren't queued behind each other.
| Key | Default | Description |
|---|---|---|
max-concurrent-streams | 1 | How many streams may be in flight simultaneously. |
queue-timeout-ms | 300000 | How long a waiting stream may queue before failing with an admission timeout. |
Protocol & Integration
Hefty runs a unified server that serves both the web UI and the API on a single port (default 41890).
The API uses JSON-RPC over WebSocket, implementing the Model Context Protocol (MCP). Any MCP-compatible client can connect - not just the built-in web UI.
WebSocket endpoint: ws://localhost:41890/mcp/ws