Documentation

Getting started with LlamaBoss

LlamaBoss is a native Windows app for chatting with AI models. Run GGUF models on your own PC, add a hosted provider when you want one, and let the agent work with your files under rules you control.

Reviewed against LlamaBoss 0.1.20 source. Updated October 3, 2026.

A chat with Agent Mode on: the model searches, reads part of a file, answers, then runs a test command after approval.
Documentation and installer versions

This guide describes the 0.1.20 source snapshot. The downloadable app release is 0.1.19, so some features may require a newer build. Check your installed version with the ⓘ button and use Check for Updates to see what is available.

Install

  1. Download the installer from the LlamaBoss home page.
  2. Run the .exe and finish the installer.
  3. Open LlamaBoss from the Start menu or the desktop shortcut.
Windows SmartScreen

The installer isn't code-signed yet, so SmartScreen may warn you. If you downloaded it from llamaboss.com, choose More info, then Run anyway. The full source is on GitHub.

Pick your first model

On first launch with no model installed, LlamaBoss opens the model downloader. You can also click the model pill in the top bar, or open Settings → Model, to download a curated model or paste a direct GGUF download link in Manage models, choose a GGUF already in your models folder, point LlamaBoss at a different folder, or pick a model from a provider you've connected.

Start small. A small model loads quickly and confirms your CPU or GPU runtime works before you try something larger.

Start a chat

  • Type a message and press Enter. Use Shift+Enter for a new line.
  • Add files with the paperclip, by dragging them onto the window, or by pasting.
  • Click the robot button to turn Agent Mode on or off for this chat.
  • Click · Auto next to the model name to change how much the model thinks before answering.
A good first test

Ask the model to explain a small text file or a screenshot. Then turn on Agent Mode and ask it to list its workspace. That checks plain chat, attachments, and a read-only tool in about a minute.

Explore the features

Local first, not local only

With a local GGUF model, your conversation is processed on your PC and works offline once the model is downloaded. LlamaBoss only goes online for features you use: model downloads, update checks and installer downloads, AI providers you connect, web pages the agent fetches, and Python packages you approve. Code run by the agent can also access the network.

Privacy & data lists every storage location and network feature.