How to Run Ollama with GPU on Ubuntu Using Docker: Install NVIDIA Driver, Container Toolkit & Open WebUI
Running your own AI chatbot used to mean sending every prompt, document and customer question to a third-party service, paying per message or per token, and hoping your data stays private. With a dedicated GPU server, you can skip all of that. In this guide, you will set up a private, ChatGPT-style assistant using Ollama (the engine that runs open-source language models on your NVIDIA GPU) and Open WebUI (a clean chat interface with user accounts, document upload and an OpenAI-compatible API for your own apps). Your prompts and files stay on your server, there are no usage limits, and your cost is simply the fixed price of the server.
We will build this on Ubuntu using Docker, so every component stays
isolated, easy to update and easy to back up. You will install the NVIDIA driver and the NVIDIA
Container Toolkit, run Ollama and Open WebUI with a single Docker Compose file, and put Caddy in
front of them for automatic HTTPS. Only Caddy is exposed to the internet, while Ollama and Open
WebUI stay on a private Docker network. The whole setup takes about 20 to 30 minutes, needs only
copy-and-paste commands, and includes fixes for the most common errors, such as
nvidia-smi failing or Ollama running on the CPU instead of the GPU.
Table of Contents
- Step 1: Update the server and check your GPU
- Step 2: Install the NVIDIA driver
- Step 3: Install Docker
- Step 4: Install the NVIDIA Container Toolkit
- Step 5: Create the project folder & configuration
- Step 6: Open the firewall
- Step 7: Start everything
- Step 8: Download your first model & verify GPU
- Step 9: Open chat interface & create admin account
- Point 1: Which model should I run? (VRAM guide)
- Point 2: Using the API from your own apps
- Troubleshooting
- Maintenance & Security Checklist
- Conclusion
Requirements
| Item | Description |
|---|---|
| Server | Dedicated server with an NVIDIA GPU (GTX 10-series or newer; RTX/Quadro/Tesla/A-series all work) |
| OS | Ubuntu 22.04 or 24.04 LTS (clean install recommended) |
| Access | Root or a sudo user, via SSH |
| RAM | 16 GB or more recommended |
| Disk | 50 GB+ free (each model takes 2 to 40+ GB) |
| Domain (optional) | A domain or subdomain pointed to your server IP, needed for HTTPS |
Step 1 Update the server and check your GPU
Log in over SSH and update the system:
sudo apt update && sudo apt upgrade -y
Confirm Ubuntu can see your GPU:
lspci | grep -i nvidia
You should see a line with your GPU model. If nothing appears, the GPU is not detected: check with your hosting provider before continuing.
Step 2 Install the NVIDIA driver
List the recommended driver for your GPU:
sudo ubuntu-drivers devices
Install it automatically (this picks the recommended version):
sudo ubuntu-drivers autoinstall
Reboot:
sudo reboot
After reconnecting, verify:
nvidia-smi
You should see a table showing your GPU name, driver version and VRAM. If you see it, the driver works.
Step 3 Install Docker
curl -fsSL https://get.docker.com | sudo sh
Allow your user to run Docker without sudo:
sudo usermod -aG docker $USER
Log out and log back in (or run newgrp docker), then test:
docker run --rm hello-world
Step 4 Install the NVIDIA Container Toolkit
This is the piece that lets Docker containers use your GPU.
Add NVIDIA's repository:
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | \
sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | \
sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
Install it:
sudo apt update
sudo apt install -y nvidia-container-toolkit
Tell Docker to use the NVIDIA runtime, then restart Docker:
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Test GPU access inside a container
docker run --rm --gpus all ubuntu nvidia-smi
If you see the same GPU table as in Step 2, GPU passthrough to Docker works. Do not continue until this test passes.
Step 5 Create the project folder
mkdir -p ~/ai-server && cd ~/ai-server
5.1 Create the docker-compose.yml
nano docker-compose.yml
Paste this:
services:
ollama:
image: ollama/ollama:latest
container_name: ollama
restart: unless-stopped
volumes:
- ollama_data:/root/.ollama
environment:
- OLLAMA_KEEP_ALIVE=30m # keep model loaded in VRAM for 30 min
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
# No "ports:" section on purpose. Ollama has no login,
# so it must NOT be reachable from the internet.
open-webui:
image: ghcr.io/open-webui/open-webui:main
container_name: open-webui
restart: unless-stopped
depends_on:
- ollama
environment:
- OLLAMA_BASE_URL=http://ollama:11434
- WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY}
- ENABLE_SIGNUP=${ENABLE_SIGNUP:-true}
volumes:
- open_webui_data:/app/backend/data
caddy:
image: caddy:2
container_name: caddy
restart: unless-stopped
depends_on:
- open-webui
ports:
- "80:80"
- "443:443"
volumes:
- ./Caddyfile:/etc/caddy/Caddyfile:ro
- caddy_data:/data
- caddy_config:/config
volumes:
ollama_data:
open_webui_data:
caddy_data:
caddy_config:
Save with Ctrl+O, Enter, then Ctrl+X.
5.2 Create the .env file
Generate a random secret key:
echo "WEBUI_SECRET_KEY=$(openssl rand -hex 32)" > .env
echo "ENABLE_SIGNUP=true" >> .env
5.3 Create the Caddyfile
Option A: with a domain (recommended, free automatic HTTPS). First point your domain's DNS A record to your server IP, then:
cat > Caddyfile << 'EOF'
ai.yourdomain.com {
reverse_proxy open-webui:8080
}
EOF
Replace ai.yourdomain.com with your real domain.
Option B: no domain (testing only, plain HTTP).
cat > Caddyfile << 'EOF'
:80 {
reverse_proxy open-webui:8080
}
EOF
Step 6 Open the firewall
Only SSH, HTTP and HTTPS are needed:
sudo ufw allow 22/tcp
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable
ports:.
That is why Ollama has no published port in the compose file above. Never publish
11434 to the internet.
Step 7 Start everything
docker compose up -d
Check that all three containers are running:
docker compose ps
Step 8 Download your first model
docker exec -it ollama ollama pull llama3.1:8b
Test it from the command line:
docker exec -it ollama ollama run llama3.1:8b "Explain what a dedicated server is in two sentences."
Verify the model is using the GPU (the most important check)
While a model is loaded, run:
docker exec -it ollama ollama ps
Look at the PROCESSOR column:
| What you see | Meaning |
|---|---|
| 100% GPU | Perfect: the whole model is in VRAM |
| 40%/60% | Model is too big for VRAM, partly on CPU (slow). Use a smaller model or lower quantization |
| CPU/GPU | Model is too big for VRAM, partly on CPU (slow). Use a smaller model or lower quantization |
| 100% CPU | GPU is not being used. See Troubleshooting, Problem 3 |
You can also watch GPU usage live in a second terminal:
watch -n 1 nvidia-smi
Step 9 Open the chat interface and create your admin account
Visit https://ai.yourdomain.com (or http://YOUR_SERVER_IP for Option
B).
- Click Get Started.
- Create your account. The first account becomes the administrator.
- Pick your model from the dropdown at the top and start chatting.
Lock down registration (do this right after creating your admin)
Otherwise anyone who finds your URL can sign up.
sed -i 's/ENABLE_SIGNUP=true/ENABLE_SIGNUP=false/' .env
docker compose up -d
Point 1Which model should I run? (VRAM guide)
Rule of thumb for 4-bit quantized models (the Ollama default): model size in GB is roughly VRAM needed, plus 1 to 2 GB for context.
| Your GPU VRAM | Comfortable model size | Example use |
|---|---|---|
| 6 to 8 GB | 3B to 8B | Chat, summaries, simple coding help |
| 12 to 16 GB | 8B to 14B | Better reasoning, coding |
| 24 GB | up to ~30B | Strong general assistant |
| 48 GB | up to ~70B | Near top-tier open models |
| 80 GB+ | 70B+ with long context | Heavy or multi-user workloads |
Browse available models at ollama.com/library and pull any with:
docker exec -it ollama ollama pull MODEL_NAME:TAG
List installed models, and delete ones you no longer need:
docker exec -it ollama ollama list
docker exec -it ollama ollama rm MODEL_NAME:TAG
Point 2 Using the API from your own apps
Ollama is OpenAI-compatible. Because we did not publish the port, call it from inside the Docker network, or temporarily test from the server itself:
docker exec -it ollama curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.1:8b",
"messages": [{"role": "user", "content": "Hello!"}]
}'
For remote API access, put it behind an authenticated gateway or use a VPN such as WireGuard. Never expose raw Ollama to the internet.
Troubleshooting
Problem 1 nvidia-smi says "NVIDIA-SMI has failed because it couldn't communicate with the NVIDIA driver"
Likely causes and fixes:
- You have not rebooted after installing the driver. Run
sudo reboot. - Secure Boot is on. Check with
mokutil --sb-state. If enabled, disable Secure Boot in BIOS/IPMI or enroll the MOK key. - Kernel and driver mismatch after a kernel update. Reinstall:
sudo apt install --reinstall linux-headers-$(uname -r)
sudo ubuntu-drivers autoinstall
sudo reboot
- The open-source
nouveaudriver is loaded. Check withlsmod | grep nouveau. If it shows, blacklist it:
echo -e "blacklist nouveau\noptions nouveau modeset=0" | sudo tee /etc/modprobe.d/blacklist-nouveau.conf
sudo update-initramfs -u && sudo reboot
Problem 2 docker: Error response from daemon: could not select device driver "" with capabilities: [[gpu]]
The NVIDIA Container Toolkit is missing or Docker was not configured. Re-run:
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
Then retest: docker run --rm --gpus all ubuntu nvidia-smi
Problem 3 Ollama is running on CPU, not GPU ("Ollama not using GPU")
Work through this checklist in order:
- Does
nvidia-smiwork on the host? If not, go to Problem 1. - Does
docker run --rm --gpus all ubuntu nvidia-smiwork? If not, go to Problem 2. - Does your compose file contain the
deploy.resources.reservations.devicesblock for theollamaservice? Without it, the container gets no GPU. - Check the Ollama logs for GPU detection messages:
docker logs ollama 2>&1 | grep -i -E "gpu|cuda|vram"
- Is the model simply too large for your VRAM? Run
ollama ps. A split such as30%/70% CPU/GPUmeans it is overflowing. Use a smaller model.
Problem 4 GPU works at first, then Ollama falls back to CPU after some time
This is a known issue on some Linux setups where the GPU driver state is lost. Try reloading the UVM module and restarting the container:
sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm
docker restart ollama
If it keeps happening, make sure you are on a recent driver version and keep the system updated.
Problem 5 Open WebUI shows no models, or "Ollama connection error"
- Confirm
OLLAMA_BASE_URL=http://ollama:11434is set and both containers are in the same compose project. - Check Ollama is up:
docker compose psanddocker logs ollama. - Confirm you actually pulled a model:
docker exec -it ollama ollama list. - Restart:
docker compose restart open-webui.
Problem 6 HTTPS does not work / Caddy shows certificate errors
- DNS A record must point to your server IP before starting Caddy. Check with
dig +short ai.yourdomain.com. - Ports 80 and 443 must be open in UFW and any provider-level firewall.
- Read the logs:
docker logs caddy. - If you hit certificate rate limits from repeated failed attempts, wait and retry, or test with Option B first.
Problem 7 "CUDA out of memory" or the model fails to load
- Use a smaller model or a smaller quantization.
- Reduce context length in Open WebUI (Settings → Advanced Params → Context Length).
- Make sure no other process is holding VRAM:
nvidia-smishows processes using the GPU.
Problem 8 Responses are very slow
- Run
ollama ps: if it is not100% GPU, that is your cause. - Large context windows slow things down. Lower the context length.
- First request after idle is slower because the model loads into VRAM.
OLLAMA_KEEP_ALIVE=30m(already set) reduces this.
Problem 9 Port 3000 or 11434 is reachable from the internet
You should not publish these ports. Remove any ports: entries for ollama and open-webui, then run docker compose up -d. Remember Docker bypasses UFW for published ports.
Maintenance
Update to the latest versions:
cd ~/ai-server
docker compose pull
docker compose up -d
docker image prune -f
Back up your data (chats, users, settings):
docker run --rm -v ai-server_open_webui_data:/data -v $(pwd):/backup ubuntu \
tar czf /backup/open-webui-backup.tar.gz /data
(The volume name is your folder name plus _open_webui_data. Check yours with
docker volume ls.)
Useful daily commands:
| Task | Command |
|---|---|
| See running containers | docker compose ps |
| View logs | docker compose logs -f |
| Loaded models | docker exec -it ollama ollama ps |
| Stop everything | docker compose down |
| Start everything | docker compose up -d |
| GPU usage | nvidia-smi |
Conclusion
That's it! You now have a fully functional, private AI chatbot running on your own hardware. Your data is secure, you have an API ready for your own applications, and you completely avoid those expensive per-token API costs.
Self-hosting AI is incredibly rewarding, but it does require the right hardware to run smoothly. If you are currently looking for a reliable machine to host your models, we at MIG servers provide high-performance dedicated GPU servers that are perfectly suited for running Ollama and other heavy AI workloads.