FAQ and troubleshooting
Find your symptom below. If you need more detail, the logs are in the LlamaBoss data folder.
Windows says it protected your PC
If you downloaded the installer from llamaboss.com, choose More info, then Run anyway. An unsigned or low-reputation installer can trigger SmartScreen; this warning is different from an antivirus detection. Verify the download source before continuing.
A local model won't load
- Check that the file is a real GGUF and not a saved web page or partial download.
- Try a smaller model, or close other apps using the GPU.
- Set the context length back to 4k or 8k for the test.
- Turn on 8-bit KV cache.
- Make sure the models folder in Settings is the one containing the file.
- Turn off Multi-token prediction if the model has MTP heads and fails to start.
- Open
server.logfor the exact llama.cpp error. - If the log says the architecture or model is unsupported, the model may be newer than the bundled llama.cpp runtime. Check for a newer LlamaBoss package; changing context length will not add architecture support.
Replies are much slower than expected
The model may not fit in VRAM and be running partly from system RAM, another window may be using the shared model, or a large context may be using up memory. Try a smaller model or context and close other GPU work. With a thinking model, set thinking to Low or Off for quick questions.
The model ignores an image
Use a vision-capable model with its matching mmproj projector. Text-only models can't see images. If several unrelated projector files sit next to each other, move each vision model and its projector into its own folder.
The model talks about a tool but doesn't use it
- Check that the robot button is on.
- Try a model that's better at tool calling.
- For a provider model, check that Allow agent tools is on and run Test selected model for Chat Completions. For OpenAI Responses models, test directly in chat with Agent Mode enabled. If the tool test fails, try the other tool protocol under Advanced.
- Type
/cdto check the working directory, then ask the agent to list that folder. The old typed file-tool commands are no longer handled by the app. - See whether the reply stopped at the tool-step limit (
/agent_steps) or a loop guard.
An AI provider fails
- Open Edit AI provider and click Refresh models. That checks the key and the address at once.
- Under Advanced, check the base URL, chat path and authentication type.
- Use the exact model ID the provider expects.
- If the provider rejects the thinking setting, run
/think auto, or pick a different Reasoning format under Advanced. - A 401 usually means an authentication problem. A 403 can also mean missing permissions, account restrictions or model access; check the provider's error message.
The context meter turns amber or red
Amber means older tool output is being trimmed from what the model sees. Red means the context window is nearly full and answers may get worse. Older output may be saved under Workspace\Vars so the model can search and read selected ranges. Start a new chat with a concise handoff or a Project, or raise local context length if your model and GPU support it.
Logs
%LOCALAPPDATA%\LlamaBoss\logs\llamaboss.log
%LOCALAPPDATA%\LlamaBoss\logs\server.logRemove keys and personal paths before posting logs anywhere. Report bugs you can reproduce on GitHub Issues.