How to Install Ollama on a VPS (and Whether Yours Can Actually Run It)

Axel Grubba, September 26, 2026
Start selling digital products with Crevio
Crevio E-Commerce Platforms logo
Crevio
Sponsored
5.0
(1)
Free plan available
Crevio is an AI-powered platform that runs your business while you sleep. Describe what you want to se... Learn more about Crevio
Get an AI summary of this post on:

The install is one command and it will work on almost any VPS. Whether the result is useful depends on a decision you should make before you run it.

On a CPU-only server, generation speed is bounded by memory bandwidth, not cores. A quantised 7–8B model produces output at roughly reading speed or slower. That is perfectly fine for summarising documents overnight, enriching records in a queue, or anything where nobody is watching. It is genuinely unpleasant for a chat interface, and no amount of vCPU will fix it — the bottleneck isn’t compute.

So decide which of these you’re building first:

  • Batch or asynchronous work — a CPU VPS is fine, and this guide’s track one covers it.
  • Anything interactive — you want a GPU, and track two covers that.

Everything else here applies to both.

Decision diagram: if nobody is waiting on the output, a CPU VPS is fine for batch jobs, summarisation and queues, needing 8GB of RAM for a 7-8B model plus swap; if someone is waiting, you want a GPU for chat, agents and user-facing work, where the model must fit in VRAM or the speed gain is lost, and renting by the hour wins while it idles since a monthly box only wins once it is busy

Install it

From Ollama’s own documentation:

curl -fsSL https://ollama.com/install.sh | sh

That installs the binary and registers a systemd service. Confirm it’s up:

systemctl status ollama
ollama --version

Pull a model and talk to it:

ollama pull llama3.1:8b
ollama run llama3.1:8b

If you’re upgrading from an older install, Ollama notes you should remove the old libraries first with sudo rm -rf /usr/lib/ollama.

The security step most guides bury

Do this before you expose anything, because the default is the safe one and it’s easy to break it without noticing.

Ollama binds 127.0.0.1 port 11434 by default — that’s the documented behaviour, and it means a fresh install is not reachable from the internet. Good. The problem starts when a tutorial tells you to “make it accessible” and you set:

Environment="OLLAMA_HOST=0.0.0.0:11434"

That publishes an unauthenticated inference endpoint to every interface, including your public IP. Ollama has no built-in authentication. Anyone who finds port 11434 open can list your models and run inference on your hardware.

This is not theoretical. Research by the Intruder team, published in May 2026, queried 5,200+ internet-facing Ollama servers with a test prompt, and 31% answered. The same scan found 518 instances wrapping paid frontier-model APIs — strangers spending someone else’s token budget.

Leave it on localhost. When you need remote access, Ollama’s documentation points at three approaches that don’t involve publishing the port:

  • A reverse proxy such as Nginx in front of it, where you can add authentication and TLS
  • Cloudflare Tunnel
  • ngrok

A private network — WireGuard or Tailscale — is the fourth and often the simplest: the service stays bound to an interface the internet can’t route to, and your laptop joins the network rather than the server joining the internet.

If you’ve already changed the bind address on a live server, verify what’s actually reachable from outside before assuming a firewall covered it. Our exposure audit walks that in five minutes — and note that if you run Ollama in Docker, ufw will not filter a published port.

How to change the bind address properly

If you do need to change it — for a reverse proxy on the same host, for instance, 127.0.0.1 is still correct. To edit the service environment, Ollama’s documented method is:

sudo systemctl edit ollama.service

Add under [Service]:

[Service]
Environment="OLLAMA_HOST=127.0.0.1:11434"

Then:

sudo systemctl daemon-reload
sudo systemctl restart ollama

Editing the unit file directly gets overwritten on upgrade; systemctl edit creates an override that survives.

Track one: CPU VPS

RAM is the hard constraint. Model weights stay resident in memory, so this is a floor rather than a target:

Model RAM floor Realistic on CPU?
7–8B quantised 8GB Yes, at reading speed or slower
13B quantised 16GB Marginal
30B 32GB No
70B quantised 40GB+ No

Add swap before you need it. Most VPS images ship with none, which turns a memory spike into a kill rather than a slowdown — and when the kernel kills Ollama, your application log shows nothing at all. If a model load has ever died silently, that’s an OOM kill, and 2–4GB of swap is the cheapest insurance available.

Measure your own throughput rather than trusting anyone’s benchmark, including ours:

ollama run llama3.1:8b --verbose

Ollama documents that flag as “show timings for response” — you get the evaluation rate in tokens per second after each reply. Run the prompt you actually care about, on the hardware you actually have, and you’ll know in two minutes whether the answer is usable.

Track two: GPU

For anything interactive, VRAM is the constraint and it works the same way — the model has to fit:

Model VRAM needed Typical card
7–8B quantised 8–12GB RTX 4000 Ada (20GB), RTX 4090 (24GB)
13B quantised 16–20GB RTX 4090, L40S (48GB)
70B quantised 40GB+ A100 / H100 (80GB)

If the model doesn’t fit in VRAM it spills to system memory and the speed advantage largely evaporates, which is the single most common disappointment with cheap GPU instances. The second most common is renting a GPU that Ollama never uses — running Ollama on a GPU server covers the driver and container-runtime setup, and the one command that tells you which is happening.

The economics here are genuinely counterintuitive, and worth a moment before you rent anything: for an interactive endpoint that idles, per-hour GPU rental is usually cheaper than owning one continuously — we worked the crossover in RunPod vs DigitalOcean GPU Droplets, and it lands much later than most people assume. If the GPU will genuinely be busy, a dedicated monthly server wins instead: RunPod vs Hetzner for Ollama compares the same 96GB card on both and finds the dedicated box materially cheaper per month.

Where to run it

For a CPU VPS, the thing to optimise is RAM per pound, since RAM is what decides which models you can load at all. Hostinger’s KVM plans pair 8GB and 16GB tiers with the simplest setup path and a one-click route if you’d rather not touch a terminal — that’s the 7–8B and 13B brackets covered. Hetzner is the cheapest per gigabyte when its CX tier is in stock — worth checking before you commit, because that tier was showing as unavailable when we last looked, and the CPX line you’d fall back to is several times dearer. That assumes you’re comfortable with a plainer panel, which matters when the 16GB tier is the one you need.

Our best VPS for Ollama guide prices that comparison properly, and 16GB VPS covers the tier most 13B users end up on.

For GPU, RunPod is the sensible default — per-second billing, a serverless tier that costs nothing while idle, and cards from consumer through H100. Hetzner’s GEX line is the cheaper option if the GPU will run constantly.

Keeping it running

Two settings worth knowing once it works.

Model load time is the latency you’ll notice most, because loading several gigabytes of weights takes seconds. Ollama can keep a model resident — its documentation covers preloading and controlling how long a model stays in memory, which turns a slow first request into a fast one.

Watch memory rather than CPU. ollama ps shows what’s currently loaded; free -h shows whether you’re near the edge. CPU at 100% during generation is normal and expected. Memory near the ceiling is the thing that ends with a dead process.

How we checked this

The install command, the systemd override method, and the statement that “Ollama binds 127.0.0.1 port 11434 by default” are from Ollama’s own documentation, read in August 2026, as are the reverse proxy, Cloudflare Tunnel and ngrok options for remote access. The --verbose flag and its “show timings for response” description are from Ollama’s CLI definition. The exposure figures — 5,200+ servers probed, 31% answering — are from research published by the Intruder team in May 2026; we have not repeated that scan.

What we did not do: we did not benchmark tokens per second on any of the hardware above. The brief for this article asked for measured throughput on a CPU VPS, and we’d rather hand you the one-line command that measures your own than publish a number from hardware that isn’t yours — quantisation, model choice and context length move it more than the vendor does. The “reading speed or slower” characterisation for CPU inference is a qualitative claim we’re comfortable making; a specific tokens-per-second figure isn’t.

RAM and VRAM figures are published floors for the model alone. Running two models, or a model plus a browser, adds them together plus the operating system.

The hosting links above are affiliate links. The security advice — keep it on localhost — costs nothing and is the most important paragraph on this page.

FAQ

Can I run Ollama on a cheap VPS?

Yes, and it will work. A quantised 7–8B model needs about 8GB of RAM and runs at roughly reading speed or slower on CPU. That’s usable for batch and background work, and frustrating for anything a person waits on.

How much RAM does Ollama need?

Roughly 8GB for a quantised 7–8B model, 16GB for a 13B, and 40GB+ for a quantised 70B. Weights stay resident in memory, so treat these as floors. Add swap regardless — most VPS images have none, and a memory spike without it kills the process silently.

How do I access Ollama remotely?

Not by setting OLLAMA_HOST=0.0.0.0, which publishes an unauthenticated endpoint. Put a reverse proxy in front of it with authentication, use Cloudflare Tunnel or ngrok, or join the server to a private network like Tailscale and leave Ollama bound to localhost.

Is Ollama secure by default?

The default bind of 127.0.0.1:11434 is safe because it isn’t reachable from outside. Ollama itself has no authentication, so the moment you change that bind address you have an open inference endpoint — which is how 31% of 5,200 scanned servers ended up answering strangers’ prompts.

How do I know how fast my server actually is?

ollama run <model> --verbose prints timings for each response, including the evaluation rate in tokens per second. Run your real prompt on your real hardware; it’s more informative than any published benchmark.

Do I need a GPU for Ollama?

Only for interactive use or models above about 13B. CPU inference is bounded by memory bandwidth, so a GPU is the fix when someone is waiting on the output — and renting one by the hour is usually cheaper than owning one if the workload idles.

Why did my model load fail with no error?

Almost certainly the kernel killed it for memory. Check with sudo dmesg -T | grep -i -E 'killed process|out of memory' — Ollama can’t log its own SIGKILL, which is why the application log simply stops.

Founder & Software Review Editor
Axel Grubba is the founder of Findstack, a B2B software comparison platform, with his background spanning management consulting and venture capital where he invested in software. Recently, Axel has developed a passion for coding and enjoys traveling when he is not building and improving Findstack.
Business Software Reviews SaaS Product Evaluation CRM Software
Subscribe, get software deals straight to your inbox.
Join 8,200+ other entrepreneurs staying up-to-date on all the latest deals.
Zero spam. Unsubscribe at any time.