How to Install Ollama on a VPS (and Whether Yours Can Actually Run It)
The install is one command and it will work on almost any VPS. Whether the result is useful depends on a decision you should make before you run it.
On a CPU-only server, generation speed is bounded by memory bandwidth, not cores. A quantised 7–8B model produces output at roughly reading speed or slower. That is perfectly fine for summarising documents overnight, enriching records in a queue, or anything where nobody is watching. It is genuinely unpleasant for a chat interface, and no amount of vCPU will fix it — the bottleneck isn’t compute.
So decide which of these you’re building first:
- Batch or asynchronous work — a CPU VPS is fine, and this guide’s track one covers it.
- Anything interactive — you want a GPU, and track two covers that.
Everything else here applies to both.
Install it
From Ollama’s own documentation:
curl -fsSL https://ollama.com/install.sh | sh
That installs the binary and registers a systemd service. Confirm it’s up:
systemctl status ollama
ollama --version
Pull a model and talk to it:
ollama pull llama3.1:8b
ollama run llama3.1:8b
If you’re upgrading from an older install, Ollama notes you should remove the old libraries first with sudo rm -rf /usr/lib/ollama.
The security step most guides bury
Do this before you expose anything, because the default is the safe one and it’s easy to break it without noticing.
Ollama binds 127.0.0.1 port 11434 by default — that’s the documented behaviour, and it means a fresh install is not reachable from the internet. Good. The problem starts when a tutorial tells you to “make it accessible” and you set:
Environment="OLLAMA_HOST=0.0.0.0:11434"
That publishes an unauthenticated inference endpoint to every interface, including your public IP. Ollama has no built-in authentication. Anyone who finds port 11434 open can list your models and run inference on your hardware.
This is not theoretical. Research by the Intruder team, published in May 2026, queried 5,200+ internet-facing Ollama servers with a test prompt, and 31% answered. The same scan found 518 instances wrapping paid frontier-model APIs — strangers spending someone else’s token budget.
Leave it on localhost. When you need remote access, Ollama’s documentation points at three approaches that don’t involve publishing the port:
- A reverse proxy such as Nginx in front of it, where you can add authentication and TLS
- Cloudflare Tunnel
- ngrok
A private network — WireGuard or Tailscale — is the fourth and often the simplest: the service stays bound to an interface the internet can’t route to, and your laptop joins the network rather than the server joining the internet.
If you’ve already changed the bind address on a live server, verify what’s actually reachable from outside before assuming a firewall covered it. Our exposure audit walks that in five minutes — and note that if you run Ollama in Docker, ufw will not filter a published port.
How to change the bind address properly
If you do need to change it — for a reverse proxy on the same host, for instance, 127.0.0.1 is still correct. To edit the service environment, Ollama’s documented method is:
sudo systemctl edit ollama.service
Add under [Service]:
[Service]
Environment="OLLAMA_HOST=127.0.0.1:11434"
Then:
sudo systemctl daemon-reload
sudo systemctl restart ollama
Editing the unit file directly gets overwritten on upgrade; systemctl edit creates an override that survives.
Track one: CPU VPS
RAM is the hard constraint. Model weights stay resident in memory, so this is a floor rather than a target:
| Model | RAM floor | Realistic on CPU? |
|---|---|---|
| 7–8B quantised | 8GB | Yes, at reading speed or slower |
| 13B quantised | 16GB | Marginal |
| 30B | 32GB | No |
| 70B quantised | 40GB+ | No |
Add swap before you need it. Most VPS images ship with none, which turns a memory spike into a kill rather than a slowdown — and when the kernel kills Ollama, your application log shows nothing at all. If a model load has ever died silently, that’s an OOM kill, and 2–4GB of swap is the cheapest insurance available.
Measure your own throughput rather than trusting anyone’s benchmark, including ours:
ollama run llama3.1:8b --verbose
Ollama documents that flag as “show timings for response” — you get the evaluation rate in tokens per second after each reply. Run the prompt you actually care about, on the hardware you actually have, and you’ll know in two minutes whether the answer is usable.
Track two: GPU
For anything interactive, VRAM is the constraint and it works the same way — the model has to fit:
| Model | VRAM needed | Typical card |
|---|---|---|
| 7–8B quantised | 8–12GB | RTX 4000 Ada (20GB), RTX 4090 (24GB) |
| 13B quantised | 16–20GB | RTX 4090, L40S (48GB) |
| 70B quantised | 40GB+ | A100 / H100 (80GB) |
If the model doesn’t fit in VRAM it spills to system memory and the speed advantage largely evaporates, which is the single most common disappointment with cheap GPU instances. The second most common is renting a GPU that Ollama never uses — running Ollama on a GPU server covers the driver and container-runtime setup, and the one command that tells you which is happening.
The economics here are genuinely counterintuitive, and worth a moment before you rent anything: for an interactive endpoint that idles, per-hour GPU rental is usually cheaper than owning one continuously — we worked the crossover in RunPod vs DigitalOcean GPU Droplets, and it lands much later than most people assume. If the GPU will genuinely be busy, a dedicated monthly server wins instead: RunPod vs Hetzner for Ollama compares the same 96GB card on both and finds the dedicated box materially cheaper per month.
Where to run it
For a CPU VPS, the thing to optimise is RAM per pound, since RAM is what decides which models you can load at all. Hostinger’s KVM plans pair 8GB and 16GB tiers with the simplest setup path and a one-click route if you’d rather not touch a terminal — that’s the 7–8B and 13B brackets covered. Hetzner is the cheapest per gigabyte when its CX tier is in stock — worth checking before you commit, because that tier was showing as unavailable when we last looked, and the CPX line you’d fall back to is several times dearer. That assumes you’re comfortable with a plainer panel, which matters when the 16GB tier is the one you need.
Our best VPS for Ollama guide prices that comparison properly, and 16GB VPS covers the tier most 13B users end up on.
For GPU, RunPod is the sensible default — per-second billing, a serverless tier that costs nothing while idle, and cards from consumer through H100. Hetzner’s GEX line is the cheaper option if the GPU will run constantly.
Keeping it running
Two settings worth knowing once it works.
Model load time is the latency you’ll notice most, because loading several gigabytes of weights takes seconds. Ollama can keep a model resident — its documentation covers preloading and controlling how long a model stays in memory, which turns a slow first request into a fast one.
Watch memory rather than CPU. ollama ps shows what’s currently loaded; free -h shows whether you’re near the edge. CPU at 100% during generation is normal and expected. Memory near the ceiling is the thing that ends with a dead process.
How we checked this
The install command, the systemd override method, and the statement that “Ollama binds 127.0.0.1 port 11434 by default” are from Ollama’s own documentation, read in August 2026, as are the reverse proxy, Cloudflare Tunnel and ngrok options for remote access. The --verbose flag and its “show timings for response” description are from Ollama’s CLI definition. The exposure figures — 5,200+ servers probed, 31% answering — are from research published by the Intruder team in May 2026; we have not repeated that scan.
What we did not do: we did not benchmark tokens per second on any of the hardware above. The brief for this article asked for measured throughput on a CPU VPS, and we’d rather hand you the one-line command that measures your own than publish a number from hardware that isn’t yours — quantisation, model choice and context length move it more than the vendor does. The “reading speed or slower” characterisation for CPU inference is a qualitative claim we’re comfortable making; a specific tokens-per-second figure isn’t.
RAM and VRAM figures are published floors for the model alone. Running two models, or a model plus a browser, adds them together plus the operating system.
The hosting links above are affiliate links. The security advice — keep it on localhost — costs nothing and is the most important paragraph on this page.
FAQ
Can I run Ollama on a cheap VPS?
Yes, and it will work. A quantised 7–8B model needs about 8GB of RAM and runs at roughly reading speed or slower on CPU. That’s usable for batch and background work, and frustrating for anything a person waits on.
How much RAM does Ollama need?
Roughly 8GB for a quantised 7–8B model, 16GB for a 13B, and 40GB+ for a quantised 70B. Weights stay resident in memory, so treat these as floors. Add swap regardless — most VPS images have none, and a memory spike without it kills the process silently.
How do I access Ollama remotely?
Not by setting OLLAMA_HOST=0.0.0.0, which publishes an unauthenticated endpoint. Put a reverse proxy in front of it with authentication, use Cloudflare Tunnel or ngrok, or join the server to a private network like Tailscale and leave Ollama bound to localhost.
Is Ollama secure by default?
The default bind of 127.0.0.1:11434 is safe because it isn’t reachable from outside. Ollama itself has no authentication, so the moment you change that bind address you have an open inference endpoint — which is how 31% of 5,200 scanned servers ended up answering strangers’ prompts.
How do I know how fast my server actually is?
ollama run <model> --verbose prints timings for each response, including the evaluation rate in tokens per second. Run your real prompt on your real hardware; it’s more informative than any published benchmark.
Do I need a GPU for Ollama?
Only for interactive use or models above about 13B. CPU inference is bounded by memory bandwidth, so a GPU is the fix when someone is waiting on the output — and renting one by the hour is usually cheaper than owning one if the workload idles.
Why did my model load fail with no error?
Almost certainly the kernel killed it for memory. Check with sudo dmesg -T | grep -i -E 'killed process|out of memory' — Ollama can’t log its own SIGKILL, which is why the application log simply stops.