Clear, practical technology insights BSOD Code Lookup · Windows Error Code Lookup · Wi-Fi Troubleshooting · PC Troubleshooting Checklist

How to Run a Local AI Model on a Raspberry Pi With Ollama

Install a small language model on a 64-bit Raspberry Pi, test it with Ollama, add an optional web interface, and diagnose memory, storage, and thermal problems.

Table of Contents

A Raspberry Pi can run a small language model locally, but expectations matter. It will not match a desktop GPU or a cloud service, and larger models may be unusably slow or fail to load. The practical goal is an offline assistant for experimentation, simple automation, or learning how local inference works.

This guide uses Ollama for model management and TinyLlama as a lightweight first test. Start with the command line; add Open WebUI only after the model works reliably.

What you need

  • A Raspberry Pi 5 is preferable; a Pi 4 can work with a sufficiently small model.
  • A 64-bit Raspberry Pi OS installation.
  • At least 4GB RAM; 8GB gives more room for the operating system and interface.
  • Reliable power and active cooling for sustained workloads.
  • Ample free storage. A USB 3 SSD is preferable to a small or low-endurance microSD card.

Raspberry Pi prepared for local AI inference

Check the architecture and available resources:

uname -m
free -h
df -h /

The architecture should report aarch64. Leave several gigabytes free for the software, model, temporary files, and updates.

1. Update Raspberry Pi OS

sudo apt update
sudo apt full-upgrade -y
sudo reboot

Reconnect after the restart. Keeping the OS current reduces the chance that an old kernel, container runtime, or library causes an unrelated failure.

2. Install and start Ollama

Use the current Linux instructions on the official Ollama documentation. Its installer command is:

curl -fsSL https://ollama.com/install.sh | sh

Then enable and check the service:

sudo systemctl enable --now ollama
systemctl status ollama --no-pager

Read an installation script before running it if that is part of your security policy. Do not expose Ollama’s network port to the internet.

3. Download a small model

Begin with TinyLlama rather than a multi-billion-parameter model:

ollama pull tinyllama
ollama run tinyllama

At the prompt, ask a short factual or formatting question. Enter /bye when finished. Responses will be slower than on a conventional AI workstation, and a small model may produce weak or inaccurate answers. The test is successful if it loads consistently without exhausting memory.

List installed models and remove one you no longer need:

ollama list
ollama rm MODEL_NAME

4. Add Open WebUI only if needed

Open WebUI connected to Ollama on a Raspberry Pi

A browser interface uses additional memory. If the command-line test is stable and Docker is already installed, follow the current Open WebUI quick start. A typical container connecting to Ollama on the host looks like this:

docker run -d   -p 3000:8080   --add-host=host.docker.internal:host-gateway   -e OLLAMA_BASE_URL=http://host.docker.internal:11434   -v open-webui:/app/backend/data   --name open-webui   --restart unless-stopped   ghcr.io/open-webui/open-webui:main

Open http://PI_ADDRESS:3000 from a device on the same trusted network. Create the first administrator account and restrict access with your firewall or router. Do not forward port 3000 from the public internet.

Monitor temperature, memory, and errors

watch -n 2 free -h
vcgencmd measure_temp
journalctl -u ollama -n 100 --no-pager
dmesg -T | grep -i -E 'out of memory|killed process'

If a model loads and then the process disappears, look for an out-of-memory event. Choose a smaller or more heavily quantized model before adding swap. Heavy swap on a microSD card is slow and can increase writes; if swap is unavoidable, use conservative settings and preferably an SSD.

If speed falls during a long response, check temperature and throttling:

vcgencmd get_throttled

A result other than 0x0 can indicate current or previous undervoltage or throttling. Confirm that the power supply and cooling solution are appropriate for the Pi model.

Common problems

The model exits or the board freezes

Close the desktop and other services, stop Open WebUI, and try a smaller model. Confirm available RAM and check the kernel log for a killed process.

The download fails

Check free disk space, DNS, and the system clock. Partially downloaded model data can consume significant storage.

Open WebUI cannot reach Ollama

Confirm that the Ollama service is running and that the container has the correct OLLAMA_BASE_URL. Review its log with docker logs open-webui.

The Pi becomes very slow

Check temperature, power warnings, memory pressure, and storage activity. Running the graphical desktop, a container interface, and inference together may exceed a low-memory Pi’s practical limits.

What this setup is good for

A Pi-based model can support private experiments, a local text interface, or a small application that tolerates slow replies. It is not a dependable source of facts and should not control safety-critical devices without separate validation and safeguards.

Once the small model is stable, change one variable at a time—model size, context length, interface, or storage. That makes failures easier to diagnose and shows where the Raspberry Pi’s real limit lies.

Discussion

Reader Comments 0

Sign in with email or Google to join the discussion.