Install

System requirements


Axon runs its models on your own machine, so what decides whether it is usable is the GPU. It picks the size of model that fits the machine; more GPU memory buys a larger model and better answers.

At a glance

Minimum Recommended
Windows / Linux A graphics card with 4 GB of its own memory, 8 GB RAM A card with 8 GB or more, 16 GB RAM
Mac Apple silicon, 8 GB Apple silicon, 16 GB or more
Disk 10 GB free More, for large document sets

A machine without a GPU — or with only the graphics built into the processor, including every Intel Mac — will install and open Axon, but answers come too slowly to be practical.

Operating system

System Supported
macOS 10.15 Catalina or later, Apple silicon or Intel
Windows Windows 10 or 11, 64-bit (x64)
Linux 64-bit (x86-64), Ubuntu 24.04 or a distribution of the same age or newer

On Windows, Axon uses Microsoft Edge WebView2. Windows 11 includes it; on Windows 10 the installer fetches it if it is missing. The Linux .deb package needs WebKitGTK 4.1 and GTK 3, which apt installs with it.

Graphics

A GPU is required for Axon to be usable.

  • Windows and Linux — a graphics card with at least 4 GB of its own memory (NVIDIA or AMD). Axon uses it through Vulkan, which current drivers provide; keep the driver up to date.
  • Apple silicon — the GPU is built in and used automatically.
  • Intel Macs, and PCs with only built-in graphics — Axon runs on the processor, which is too slow for everyday use.

The models Axon runs

Everything Axon does with your documents runs on these models, all on your machine and all fetched by the Setup assistant on first launch:

Job Model Size When it runs
Chat — answers questions, runs skills Qwen3.5, sized for your machine (below) 0.5–5.7 GB Whenever you ask something
Summaries — a summary of each document, written while indexing The chat model, unless you pick another in Settings → LLM → Models — While indexing
Embedding — turns passages into what search matches against BGE-large-en-v1.5 0.7 GB While indexing, and on every question
OCR / vision — reads scanned PDF pages and images PaddleOCR-VL 1.6 1.8 GB Only while indexing scans and images; unloaded after two minutes idle

The chat and embedding models stay loaded side by side. The vision model joins them while you index scanned documents, so a large batch of scans is when the GPU is busiest.

Which chat model you get

The chat model is always Qwen3.5; Axon picks its size from the memory it will run in, after leaving room for the embedding model.

With a graphics card, that is the card's free memory, not the computer's. Part of it is always taken by the embedding model and working space, so in practice:

Graphics card Chat model
4 GB Qwen3.5 2B
6 GB or 8 GB Qwen3.5 4B
10 GB or more Qwen3.5 9B

On Apple silicon, it is the Mac's memory, which the GPU shares:

Memory Chat model
8 GB Qwen3.5 2B
16 GB Qwen3.5 4B
32 GB or more Qwen3.5 9B

Close other GPU-heavy programs before a long session — what they hold is not free for Axon. Settings → LLM → On this machine shows which model Axon chose and the GPUs it uses.

Disk space

What Size
The local models 2.5 GB for embedding and OCR, plus the chat model — about 3.8 GB in all with the 2B, 8.2 GB with the 9B
Search indexes Grows with the documents you add; plain text takes far less than its original files

Allow 10 GB free to install Axon and start working. Your documents are not copied — a project indexes a folder where it already sits.

Network

Axon needs the internet three times: to download it, to download the local models on first launch, and to fetch updates when you ask for them. After that it answers questions with no connection at all.

Some features use a network on purpose, and only when you turn them on:

  • Sharing a project (Pro) uses your local network between the two computers.
  • Cloud-drive search (Pro) and web search for connectors reach the service you connected.
  • Pro connects to exaflux.io now and then to renew the licence.
  • Your own Ollama server, if you point Axon at one.

See Settings & providers for exactly what runs where.

Does something here not match what you see in Axon? Tell us.