In short
Install Ollama, run a small model and check which processor is used. A successful response does not prove Radeon acceleration. Current documentation includes a Vulkan route for additional GPU support.
Requirements: Windows 11, current graphics drivers, storage and internet for initial downloads. This procedure follows upstream documentation; it is not a claimed execution test on a specific device.
Install Ollama
Download the official Windows installer from ollama.com. After installation, open a new PowerShell window and run:
ollama --version
ollama run qwen2.5:3bThe first run downloads model data. Complete the download before judging response speed.
Make Ollama use the Radeon iGPU
Install the appropriate AMD graphics driver, restart Ollama and run a prompt. Current documentation says Vulkan is enabled by default when the backend is installed. Inspect ollama ps while the model is loaded. CPU-only use calls for a driver/backend check; the NPU is a separate path.
Pick a model that fits your memory
Begin with the 3B example. Context and other applications increase memory use beyond downloaded weights. Close unnecessary workloads and reduce context before assuming a larger computer is the only answer to an allocation error.
Add a chat UI (Open WebUI)
With uv installed, use the documented Windows route:
$env:DATA_DIR="C:\open-webui\data"
uvx --python 3.11 open-webui@latest serveOpen the address printed by the process and connect to Ollama at http://localhost:11434. Keep access local unless you deliberately configure authenticated wider use.
Frequently asked questions
Does Ollama work on AMD GPUs?
Yes, through supported backends and drivers. Check current platform support and verify actual offload.
How much RAM do I need for Ollama?
Small quantised models use modest memory, but leave room for Windows and context. 32GB is a useful starting point for broader experimentation.