Models and performance
LlamaBoss ships with llama.cpp for CPU and NVIDIA CUDA, finds GGUF files in your models folder, and can list models from AI providers you've connected.
Local GGUF models
Use Settings → Model → Download models for the curated catalog, or put .gguf files in your models folder. Manage models also accepts a direct HTTPS link to one GGUF file.
A Hugging Face file-page link is converted to its download link automatically. A repository homepage or one part of a split model is not supported by this downloader. Gated Hugging Face models may need to be downloaded in your browser after you accept their license.
The downloader checks for a GGUF file and saves through a temporary file. A failed replacement preserves the existing model. Manage models also lets you delete a model after it is unloaded. The default folder is:
%LOCALAPPDATA%\LlamaBoss\models
Use your own model folder
In Settings → Model → Location, click Change to use another folder, for example one on a bigger SSD. The downloader and the model picker both switch to it. LlamaBoss remembers the choice and never copies model files. Reset returns to the default.
Vision models
A model that can see images usually needs a matching projector file with mmproj in its name. Keep the model and its projector together in one folder. Curated vision downloads set this up for you.
Don't keep several unrelated projector files next to one model. If LlamaBoss can't tell which one belongs to the model, it won't guess.
Context length and 8-bit KV cache
Settings → Context length goes from 4k to 256k tokens. More context lets the model remember more of the conversation and tool output, but uses more memory and takes longer to process.
8-bit KV cache is recommended. Compared with a 16-bit KV cache, it roughly halves the KV memory requirement; model weights and other GPU allocations still use memory. Changing either setting reloads the model.
The model has to support the size you pick. Choosing 256k cannot give a model a reliable 256k context if it was not trained for one. The running llama.cpp server may also report a smaller window; check the context meter and server log for the effective size.
Multi-token prediction
Multi-token prediction (auto) speeds up generation on models built with MTP heads, such as some GLM and Qwen builds. LlamaBoss detects these from the GGUF file and does nothing on other models. Turn it off in Settings → Context length if a model misbehaves with it. Changing it reloads the model.
Thinking
Reasoning models can think before they answer. Choose how much from the thinking chip in the model pill, or with /think. The setting belongs to the conversation and applies from your next message.
| Setting | What LlamaBoss asks for |
|---|---|
| Auto | Nothing. The model or provider uses its own default. |
| Off | No thinking. |
| On | Thinking on. |
| Low, Medium, High | A thinking effort level. Local models treat any level as On. |
Support depends on the model and provider. Some reasoning models can't turn thinking off; LlamaBoss tells you and switches to the lowest level they accept.
Keeping it fast
- Pick a model that fits entirely in GPU VRAM when you can.
- If generation suddenly gets very slow, the model is probably spilling into system RAM.
- Close games, image generators and other GPU-heavy apps before loading a large model.
- A smaller quantization or a smaller model is often much faster for a modest loss in quality.
- A long context uses memory even before the answer gets long.
To compare response speed on your own machine, use /bench help and then /bench. It runs a fixed test separately from the chat and saves timing results. Benchmarks against a hosted provider use that provider's normal billing.
Remote models
Models from AI providers appear in the same picker, grouped by provider. They run on the provider's servers, so your messages and attachments leave your PC when you use them. See AI providers.