AI Settings

The AI Settings dialog is where admins connect Hefty to LLM providers and configure how the agent processes information. Access it via User Menu → AI Settings (admin only).

Mandatory on First Setup

When cognition has not been configured yet, the dialog opens automatically and cannot be dismissed until the configuration is saved. After the initial setup, it becomes optional and can be opened and closed freely.

LLM Tab

The LLM tab is the primary tab for configuring text model providers. Providers are managed as an ordered list of entries: each collapsed row shows the provider and its model along with a status dot for the latest connection test, and can be moved up/down or removed. New entries start from the Ollama defaults and can be switched to any of the three supported providers (see AI Providers).

Expanding an entry reveals its settings:

  • Provider — Ollama, OpenAI compatible, or Hefty platform
  • Base URL — pre-filled with the provider's default and editable for custom endpoints; for the Hefty platform it is fixed to the managed gateway and cannot be changed
  • API key — required for cloud providers, with a show/hide toggle
  • Text model — pick from the models the connected provider offers and from Hefty's recommended models, or enter a custom model name
  • Connection Test — verifies connectivity, shows provider version info on success, and displays error messages on failure

Each entry also carries a collapsible per-entry settings group:

  • Temperature — optional per-entry override (slider); when disabled, the provider's default applies
  • Keep-alive minutes — how long the model stays loaded in memory
  • Stream max tokens — a safety cap on tokens per streaming response
  • Context token limit — maximum context window size
  • Max concurrent requests — how many requests Hefty may send to this provider in parallel (1–64), across all users; when unset, the global limits apply. Raise it for powerful providers so parallel users aren't serialized behind each other's requests
  • Thinking mode — Default, On, or Off — controls extended thinking for models that support it
  • Reasoning effort — a low, medium, or high thinking budget for the model (hidden when thinking is off)
Fallback Chain

Entries form an ordered fallback chain: the first entry is the primary provider, and the remaining entries are tried in order when one fails. Reorder entries with the up/down buttons to change which provider is used first.

Embeddings Tab

The Embeddings tab configures the embedding models used for Hefty's knowledge system (vector search). It uses the same entry-list UI as the LLM tab. Each entry provides:

  • Provider and Base URL — the same three providers; the Hefty platform URL is fixed
  • Embedding model — a model name, pre-filled with the provider's default on selection
  • API key — for cloud providers, with a show/hide toggle
  • Embedding dimension — optional explicit vector dimension; derived from the model when left unset
  • Context token limit — how much text can be embedded per request

The embedding provider can be the same as or separate from your text model provider — for example, a cloud text model with local Ollama embeddings.

Global Tab

Global settings apply across all providers and control the agent's reasoning behavior:

  • Max reflection loops — the maximum number of reasoning iterations the agent can perform for a single task (controls how deeply Hefty can think through complex problems)
  • Reasoning timeout — maximum time in seconds for a single reasoning pass before it is aborted
  • Circuit breaker: max failures — the number of consecutive failures allowed before the circuit breaker trips and stops sending requests to the provider
  • Circuit breaker: reset timeout — how long to wait before the circuit breaker resets and allows requests to resume
  • Max concurrent streams — how many streaming responses Hefty runs in parallel. A stream holds a slot at the provider for its entire duration (which can be minutes), so keep this at your provider's in-flight limit for tight cloud tiers; local providers like Ollama can go higher so parallel conversations aren't queued behind each other
  • Stream queue timeout — how long a stream that is waiting for a free slot may wait before it fails with an admission timeout

Model Downloads

Downloads apply to Ollama entries only — remote providers are probed for availability but never downloaded from. When Ollama models need downloading, the dialog switches to a download view:

  • A progress bar shows download percentage and bytes downloaded out of total
  • The dialog auto-refreshes every 2 seconds during active downloads
  • The dialog cannot be closed while critical models are still downloading
  • Failed downloads display error states with details and a link to the Ollama download page
  • The dialog closes itself automatically once downloads complete

Changes take effect immediately — Hefty hot-reloads the cognition subsystem when the configuration is saved, with no restart needed. The configuration is persisted to hefty.conf in the data directory (see Advanced for the full key reference).