Advanced

Most configuration is done via the in-app AI Settings dialog and persisted to hefty.conf in your data directory. The environment variables below are deployment-level settings; the reference tables further down cover the AI settings keys the dialog writes.

Environment Variables Reference

Server

  • HEFTY_PORT - port the desktop app uses to reach the local server (GUI + API). Default: 41890
  • HEFTY_DATA_DIR - data storage location; your hefty.conf and everything Hefty remembers live here. Default: ~/.hefty
  • HEFTY_HEADLESS - when true, the server runs in API-only mode with no GUI routes
  • HEFTY_VEC_EXTENSION_PATH - path to the sqlite-vec extension. Auto-detected in the desktop app.
Provider Setup Details

Ollama (Default)

Ollama runs on port 11434 by default, which matches Hefty's default configuration. No extra setup needed beyond pulling the models — Hefty can also download its recommended models for you during setup.

OpenAI Compatible

Point the entry's base URL at any OpenAI-compatible endpoint — OpenAI, OpenRouter, Groq, or a local server like LM Studio or vLLM all work. Add the service's API key, then pick a model from the ones the service offers or enter a custom model name.

Hefty Platform

Enter your Hefty platform API key. The gateway URL is fixed and cannot be changed; default models are provided by the platform.

hefty.conf AI Settings Reference

The AI Settings dialog writes these keys to hefty.conf on save. You can also hand-edit the file while Hefty is stopped — changes take effect at the next start. Per-entry API keys are stored by the dialog when you save; treat your hefty.conf as sensitive.

llm-providers[]

An ordered list — the first entry is the primary provider, the rest are fallbacks tried in order when one fails.

KeyDefaultDescription
providerollamaOne of ollama, openai-compatible, or hefty. Unknown legacy ids are mapped automatically.
base-urlhttp://localhost:11434The provider's endpoint. Managed gateway URL is applied for the Hefty platform.
text-modelgemma4:e4bThe reasoning model to use.
keep-alive-minutes15How long the model stays loaded in memory.
stream-max-tokens16384Completion-token cap for streaming responses. Non-positive values fall back to the default.
context-token-limit32000Maximum context window size.
max-concurrent-requestsunsetParallel request cap for this entry (1–64, out-of-range values are clamped with a warning). Applies across all users; when unset, the global priority-queue.max-concurrent and stream.max-concurrent-streams limits apply.
thinkunsetTri-state extended thinking switch: true (on), false (off), unset (provider/model default).
reasoning-effortmediumThinking budget for the model: low, medium, or high. Only sent while thinking is enabled.

embedding-providers[]

KeyDefaultDescription
providerollamaSame three providers as the LLM list.
base-urlhttp://localhost:11434The embedding provider's endpoint.
embedding-modelnomic-embed-textThe embedding model to use.
embedding-dimensionderivedExplicit vector dimension; derived from the model when unset.
context-token-limit2048How much text can be embedded per request.

Global AI settings keys

KeyDefaultDescription
max-reflection-loops10Maximum reasoning iterations per task.
circuit-breaker.max-failures5Consecutive failures before the provider is tripped.
circuit-breaker.reset-timeout30sHow long before a tripped provider is retried.

priority-queue

Controls the shared queue that admits LLM calls. max-concurrent bounds all LLM calls dispatched in parallel — streaming and non-streaming alike (embeddings, titles, one-shot completions) — and serves as the global fallback for provider entries that don't set their own max-concurrent-requests. The aging settings promote requests that have waited a long time so low-priority work is never starved: a waiting request's effective priority grows by aging-step points every aging-interval-ms, so with the defaults a priority-1 request overtakes freshly enqueued priority-100 requests after roughly 25 minutes of waiting.

KeyDefaultDescription
enabledtrueTurn the priority queue on or off.
max-concurrent3Global parallel cap for all LLM calls. Set to 1 on gateway tiers that count every completion against the same per-key in-flight cap; higher suits local (Ollama) and multi-provider setups.
max-queue-depth100Maximum number of waiting requests.
aging-interval-ms15000How often a waiting request's priority bonus accrues.
aging-step1Priority points added per aging interval.

stream

Streaming admission control. A stream holds an in-flight slot at the provider for its entire duration — minutes, not seconds — so concurrent streams from your instance serialize locally: up to max-concurrent-streams run at once, and the excess wait in a queue bounded by queue-timeout-ms before failing with a typed admission timeout (instead of contending at the gateway). Match max-concurrent-streams to your tier's in-flight cap; on local (Ollama) or other no-slot-cap deployments, raise it so parallel conversations aren't queued behind each other.

KeyDefaultDescription
max-concurrent-streams1How many streams may be in flight simultaneously.
queue-timeout-ms300000How long a waiting stream may queue before failing with an admission timeout.
Protocol & Integration

Hefty runs a unified server that serves both the web UI and the API on a single port (default 41890).

The API uses JSON-RPC over WebSocket, implementing the Model Context Protocol (MCP). Any MCP-compatible client can connect - not just the built-in web UI.

WebSocket endpoint: ws://localhost:41890/mcp/ws