Install
System requirements
Axon runs its models on your own machine, so what decides whether it is usable is the GPU. It picks the size of model that fits the machine; more GPU memory buys a larger model and better answers.
At a glance
| Minimum | Recommended | |
|---|---|---|
| Windows / Linux | A graphics card with 4 GB of its own memory, 8 GB RAM | A card with 8 GB or more, 16 GB RAM |
| Mac | Apple silicon, 8 GB | Apple silicon, 16 GB or more |
| Disk | 10 GB free | More, for large document sets |
A machine without a GPU — or with only the graphics built into the processor, including every Intel Mac — will install and open Axon, but answers come too slowly to be practical.
Operating system
| System | Supported |
|---|---|
| macOS | 10.15 Catalina or later, Apple silicon or Intel |
| Windows | Windows 10 or 11, 64-bit (x64) |
| Linux | 64-bit (x86-64), Ubuntu 24.04 or a distribution of the same age or newer |
On Windows, Axon uses Microsoft Edge WebView2. Windows 11 includes it; on
Windows 10 the installer fetches it if it is missing. The Linux .deb
package needs WebKitGTK 4.1 and GTK 3, which apt installs with it.
Graphics
A GPU is required for Axon to be usable.
- Windows and Linux — a graphics card with at least 4 GB of its own memory (NVIDIA or AMD). Axon uses it through Vulkan, which current drivers provide; keep the driver up to date.
- Apple silicon — the GPU is built in and used automatically.
- Intel Macs, and PCs with only built-in graphics — Axon runs on the processor, which is too slow for everyday use.
The models Axon runs
Everything Axon does with your documents runs on these models, all on your machine and all fetched by the Setup assistant on first launch:
| Job | Model | Size | When it runs |
|---|---|---|---|
| Chat — answers questions, runs skills | Qwen3.5, sized for your machine (below) | 0.5–5.7 GB | Whenever you ask something |
| Summaries — a summary of each document, written while indexing | The chat model, unless you pick another in Settings → LLM → Models | — | While indexing |
| Embedding — turns passages into what search matches against | BGE-large-en-v1.5 | 0.7 GB | While indexing, and on every question |
| OCR / vision — reads scanned PDF pages and images | PaddleOCR-VL 1.6 | 1.8 GB | Only while indexing scans and images; unloaded after two minutes idle |
The chat and embedding models stay loaded side by side. The vision model joins them while you index scanned documents, so a large batch of scans is when the GPU is busiest.
Which chat model you get
The chat model is always Qwen3.5; Axon picks its size from the memory it will run in, after leaving room for the embedding model.
With a graphics card, that is the card's free memory, not the computer's. Part of it is always taken by the embedding model and working space, so in practice:
| Graphics card | Chat model |
|---|---|
| 4 GB | Qwen3.5 2B |
| 6 GB or 8 GB | Qwen3.5 4B |
| 10 GB or more | Qwen3.5 9B |
On Apple silicon, it is the Mac's memory, which the GPU shares:
| Memory | Chat model |
|---|---|
| 8 GB | Qwen3.5 2B |
| 16 GB | Qwen3.5 4B |
| 32 GB or more | Qwen3.5 9B |
Close other GPU-heavy programs before a long session — what they hold is not free for Axon. Settings → LLM → On this machine shows which model Axon chose and the GPUs it uses.
Disk space
| What | Size |
|---|---|
| The local models | 2.5 GB for embedding and OCR, plus the chat model — about 3.8 GB in all with the 2B, 8.2 GB with the 9B |
| Search indexes | Grows with the documents you add; plain text takes far less than its original files |
Allow 10 GB free to install Axon and start working. Your documents are not copied — a project indexes a folder where it already sits.
Network
Axon needs the internet three times: to download it, to download the local models on first launch, and to fetch updates when you ask for them. After that it answers questions with no connection at all.
Some features use a network on purpose, and only when you turn them on:
- Sharing a project (Pro) uses your local network between the two computers.
- Cloud-drive search (Pro) and web search for connectors reach the service you connected.
- Pro connects to exaflux.io now and then to renew the licence.
- Your own Ollama server, if you point Axon at one.
See Settings & providers for exactly what runs where.
Does something here not match what you see in Axon? Tell us.

