I am still working on this post, which is mostly complete.
Thanks for your visit!
TL;DR:
Install Ollama (native on Apple Silicon), add Open WebUI via Docker, pull a quantised model, and connect your own documents — without sending a single token to the cloud.
Quantisation is the equaliser: Ollama's default Llama 3.2 is a 3B-parameter Q4_K_M build of around 2 GB, comfortable even on an 8 GB M1, and larger models scale from there.
A summer of model access changes — the export-control withdrawal of Claude Fable 5 and Mythos 5, and GitHub Copilot's move to usage-based AI Credits — made the same point twice: a rented model can be revoked or repriced on someone else's schedule, an open-weight model on your own disk cannot.
Why I wanted a local AI lab
It has been a stranger few months than usual for anyone who builds on top of rented AI models, and by September the pattern was hard to ignore.
On 1 June, GitHub Copilot moved to usage-based billing. The flat monthly fee gained a meter: Copilot plans now bill against GitHub AI Credits, with extra usage charged once your monthly allowance runs out.
And while all that was happening, OpenCode — an open-source, model-agnostic coding agent that is perfectly happy talking to a local Ollama server — reached 7% developer adoption in JetBrains’ 2026 developer survey, with 42% of developers aware of it despite no big company name behind it.
The common thread is hard to miss. The cloud model you build on can be revoked, repriced, or rate-limited on someone else’s timetable; an open-weight model sitting on your own disk cannot. That has always been the quiet argument for a local AI lab, and this summer made it loud.
The good news is that the local AI ecosystem has hit a practical tipping point at exactly the right moment. Models that would have required a data centre two years ago now run on a mid-range laptop, and the tooling to serve, interact with, and augment those models has become genuinely approachable.
This post walks through a complete local AI lab: a model engine, a web interface that feels like ChatGPT, and a private RAG pipeline that lets your AI answer questions grounded in your own documents.
I run this stack on an Apple M1 MacBook, so where the steps differ from other platforms I have called that out explicitly.
A companion post covers the security picture — specifically, what changes when you add autonomous agents to this setup. Once the lab is actually answering, a follow-up on calling it from Python covers the two APIs and the jobs that only make sense after both doors are wired.
What we’re actually building
Component
Job
Ollama
Loads and runs the LLM
Open WebUI
Gives you the chat interface
Embedding model
Turns document chunks and queries into vectors
RAG
Finds relevant chunks before asking the LLM
Docker
Runs Open WebUI in an isolated container
Here is how those pieces fit together — worth pinning down before you install anything, because it is easy to assume Docker is somehow running your model. It is not. Open WebUI is the interface and orchestration layer; Ollama does the actual inference.
The whole stack runs offline once everything is downloaded. You can pull the Wi-Fi cable and it keeps working — more on what that actually means, and does not mean, later in this post.
By the end of this post, you will have:
Llama 3.2 running locally through Ollama as llama3.2:latest
a ChatGPT-style Open WebUI interface, reachable at http://localhost:3000
persistent chats and settings that survive a container restart
a private knowledge base built from your own documents
a test proving retrieval actually works, not just that it looks installed
a test proving the whole stack still works with no internet connection
Phase 1: The model engine — Ollama
Ollama is a local LLM runtime that handles model downloads, GPU routing, and serving behind a single CLI — the fastest way to get a model running on your own hardware.
Install Ollama
Go to ollama.com and download the installer for your OS. Once installed, verify it is running:
ollama --version
On Apple Silicon (M1/M2/M3/M4): the macOS .pkg installer works natively — no Rosetta needed. On Apple Silicon, Ollama runs natively and uses Apple’s GPU through Metal automatically. No CUDA setup or separate GPU configuration is required. (Metal is Apple’s equivalent of CUDA — if you have ever set up an NVIDIA card on a PC, this is the same idea, just handled for you.) If you prefer Homebrew:
brew install ollama
brew services start ollama
The brew services start command keeps Ollama running in the background across reboots, which is handy once you start using Open WebUI.
Pull and run a model
ollama run llama3.2
Running ollama run llama3.2 downloads the model and opens an interactive terminal chat. For a first test it works fine. For everyday use you will want the web interface.
Check the API, not just the CLI
Type /bye to exit the interactive chat, then check that Ollama’s HTTP server is actually listening:
curl http://localhost:11434/api/tags
You should get back JSON containing "name": "llama3.2:latest". ollama pull llama3.2 (or ollama run llama3.2) writes that tagged name to disk — there is no bare llama3.2 entry, and Open WebUI’s model selector and API use the tagged id. This is worth doing now rather than later, because Open WebUI never talks to the ollama CLI — it talks to this HTTP API on port 11434. If this curl fails, there is no point installing Open WebUI yet; fix Ollama first.
Two more commands worth knowing the difference between:
ollama list
ollama ps
ollama list shows every model you have downloaded to disk. ollama ps shows which of those are currently loaded into memory and actively serving. A model can appear in list and not in ps — that just means it is on disk but not running right now, which is normal and not a fault.
Worth knowing before you pull anything larger: Ollama’s default Llama 3.2 is a 3B-parameter, Q4_K_M build of around 2 GB — the Llama 3.2 text family only ships as 1B and 3B, there is no 7B version of it. That makes it a genuinely comfortable first model, even on an 8 GB M1. Larger 7B–8B models from other families (Llama 3.1, Mistral, Qwen) are entirely possible, but they leave much less headroom for macOS, the context cache, and whatever else you have open.
Choosing the right model size
Ollama models use a name:tag convention. A tag may specify the model’s size, variant, or quantisation level — for example llama3.2:1b, llama3.2:latest, or llama3.2:1b-instruct-q4_K_M. The tags containing names such as Q4_K_M, Q5_K_M, and Q8_0 are the ones that tell you how a model’s weights have been quantised — a compression technique that reduces the precision of a model’s weights (from, say, 16-bit to 4-bit) to shrink memory use.
Tag suffix
Quantisation
Typical memory (7B-class model)
Notes
:q4_K_M
4-bit
~4 GB
Best default for most users
:q5_K_M
5-bit
~5 GB
Slightly better quality
:q8_0
8-bit
~8 GB
Near full quality
(no tag)
Default
varies
Usually q4 or q5
Moving from 16-bit weights to a good 4-bit quantisation dramatically reduces storage and memory requirements, usually with a relatively small quality trade-off for ordinary chat. The exact difference depends on the model and the task — reasoning, coding, and precise factual work can show quantisation degradation even when everyday chat seems unaffected, so I would not treat “Q4 is basically free” as a universal law.
Phase 2: The web interface — Open WebUI
🔒 Subscribe to keep reading.
Phase 3: Private RAG — your AI on your documents
🔒 Subscribe to keep reading.
Phase 4: Prove the lab
🔒 Subscribe to keep reading.
Phase 5: Operate the lab
🔒 Subscribe to keep reading.
How much hardware do you actually need?
🔒 Subscribe to keep reading.
What “private” actually means
🔒 Subscribe to keep reading.
Where agents change the threat model
🔒 Subscribe to keep reading.
Quick-start checklist for a local AI lab
🔒 Subscribe to keep reading.
References
🔒 Subscribe to keep reading.
You've hit a Deep Dive tutorial.
I spend dozens of hours researching, coding, and breaking things to write these guides. This content is free, but reserved for my subscriber community. Drop your email below to unlock this guide (and all past/future deep dives):
Full content temporarily unavailable — refresh in a moment
You're in
I'll send weekly notes on AI tools, Python, and what I'm actually building. Check your inbox for Set a password to unlock articles — the form does not log you in.
New subscribers get an inbox mail: Set a password to unlock articles. The form does not log you in — use the same email afterwards.
References
About Elena
Elena, a PhD in Computer Science, simplifies AI concepts and helps you use machine learning.
Citation
Elena Daehnhardt. (2026) 'Build a Private AI Lab on Your Own Hardware: Why Local Models Matter More in 2026', daehnhardt.com, 08 September 2026. Available at: https://daehnhardt.com/blog/2026/09/08/private-ai-lab-ollama-open-webui-rag/
Want the Fantastic AI 2026 toolkit and the Git cheatsheet in your inbox?
Fantastic AI: The 2026 ToolkitGit Commands & WorkflowVibe Coding Checklist
You're in
I'll send weekly notes on AI tools, Python, and what I'm actually building. Check your inbox for Set a password to unlock articles — the form does not log you in.