October 11, 2026
Tutorials

How Much RAM and CPU You Really Need to Run AI Models on a VPS

How Much RAM and CPU You Really Need to Run AI Models on a VPS

Most people shopping for a VPS to run a local AI model ask the wrong question first. They compare core counts and clock speeds, when the thing that actually decides whether a model runs well, runs badly, or refuses to run at all is memory. RAM sets the ceiling on which models you can load, and memory bandwidth sets how fast they respond. CPU cores matter, but far less than the pricing pages suggest.

This guide gives you real numbers for matching a VPS to a model, explains where the usual rules of thumb break down, and shows you how to check what your server can actually handle before you spend money on a bigger plan.

Table of Content

The One Formula of RAM and CPU That Explains Almost Everything

Running a language model comes down to two separate questions, and mixing them up is the most common mistake in this whole topic.

Will it fit? This depends on RAM. The model’s weights have to sit entirely in memory, plus extra room for the conversation context and the operating system.

Will it be fast enough? This depends mostly on memory bandwidth, meaning how quickly your server can read those weights out of RAM. Generating each word requires reading through essentially the whole model once. So a rough ceiling for speed is your memory bandwidth divided by the size of the model in memory.

A model that fits but sits on slow memory will work, just slowly. A model that does not fit will either crash or start using swap on disk, and swap is a performance disaster. One well-known report described a 13B model on a laptop falling to roughly one token per minute once it spilled into swap, down from about six tokens per second for a smaller model that fit comfortably. Keep that picture in mind whenever you are tempted to squeeze a model onto a server that is slightly too small.

How Much RAM and CPU You Really Need to Run AI Models on a VPS

It depends heavily on whether you are using external APIs (like OpenAI or OpenRouter) or hosting a local AI model directly on the server.

For a typical 7B to 14B AI model on a VPS, you need 16 GB to 32 GB of RAM to load the model weights and operating system.

You require a minimum of 4 to 8 vCPU cores to handle the processing load, though CPU inference will remain much slower than using a GPU.

How Much RAM Each Model Size Needs

Models are usually run quantized, meaning the numbers inside are stored at lower precision to save space. The most common format for local use is Q4_K_M, which holds a model at roughly 4 bits per parameter. A reliable rule of thumb at that level is about 0.6 GB of memory per billion parameters for the weights alone.

Here is what that looks like in practice, with headroom included for the operating system and context.

Model size

Weights at Q4

Recommended VPS RAM

Typical examples

1B to 3B

1 to 2 GB

4 GB

Llama 3.2 3B, Phi-4 Mini

7B to 8B

4 to 5 GB

8 GB, ideally 16 GB

Llama 3.3 8B, Qwen3 8B, Mistral 7B

13B to 14B

8 to 9 GB

16 GB

Qwen3 14B

30B to 35B

18 to 22 GB

32 GB

Qwen3 32B class models

70B

40 to 42 GB

64 GB or more

Llama 3.3 70B

That recommended column is deliberately higher than the raw weight size. Your VPS also has to hold the operating system, Docker, and a chat interface like Open WebUI, which together can take one to two gigabytes before the model even loads. Context adds more on top. Long conversations and large documents grow what is called the KV cache, and that memory scales with how much text you feed in.

Quantization choice moves these numbers significantly. Q8 roughly doubles the memory needed compared with Q4 and gives a small quality gain. Going the other direction, more aggressive 2- and 3 bit formats shrink the footprint further but start to cost real accuracy. For most VPS setups, Q4_K_M is the sensible default.

What About a GPU

Everything above assumes CPU-only inference, which is what most standard VPS plans give you. If you can rent a GPU instance, the math shifts, because video memory is much faster than system RAM.

The same rule applies, with the model needing to fit in VRAM, and a 7B model at Q4 fits comfortably in an 8 GB card while a 70B model needs around 40 GB or more of VRAM. If you want snappy responses from anything above 13B, a GPU is usually the better use of money than piling on more CPU cores.

How Many CPU Cores Do You Actually Need

This is where VPS marketing misleads people, because core count looks like the main spec when it is closer to a secondary one.

During text generation, the work is limited by how fast memory can feed the processor, not by how many cores are crunching numbers. That means doubling your cores does not double your speed. Past a handful of cores, extra threads mostly sit waiting on memory.

Cores matter more during prompt processing, the stage where the model reads your input before it starts replying. Long prompts, pasted documents, and retrieval workflows all lean on this stage, and more cores genuinely help there.

A practical guide for CPU-only inference:

  • 2 to 4 vCPUs is fine for light personal use with a small model, though long prompts will feel sluggish.
  • 4 to 8 vCPUs is the sweet spot for most self-hosted assistants running a 7B to 14B model for one or a few users.
  • 8 or more vCPUs pays off if several people use the server at once, or if you regularly process large documents.

Two quieter details are worth checking. Look for a CPU that supports AVX2, and ideally AVX-512, since inference software uses these instruction sets for much faster math. Also be aware that vCPUs are often hyperthreads, so a plan advertising eight vCPUs may only give you four physical cores worth of real compute, and a shared host can slow you further if neighbors are busy.

Realistic Speed Expectations on a VPS

Honest speed numbers are hard to give, because VPS memory bandwidth varies a lot between providers and hardware generations. What you can do is use the formula from earlier. A Q4 model of about 5 GB on a server delivering roughly 40 GB per second of effective bandwidth has a theoretical ceiling near 8 tokens per second. Real results land below that, often in the range of 3 to 8 tokens per second for an 8B model on a decent CPU plan.

For comparison, comfortable reading speed sits around 5 to 8 tokens per second, so an 8B model on a good CPU VPS is usable for chat, just not instant. A 70B dense model reading around 40 GB per token on the same hardware would crawl at a fraction of a token per second, which is why large models on CPU-only servers are rarely practical no matter how much RAM you add.

One exception worth knowing about is mixture-of-experts models. These have a large total parameter count but only activate a small portion for each token, so they need plenty of RAM to load while generating faster than a dense model of the same size. They are a smart match for servers with lots of RAM and modest bandwidth.

A Simple Sizing Checklist

Run through these steps before choosing a plan.

  1. Decide on the model you actually want, then find its Q4 file size. That is your minimum for weights.
  2. Add 2 to 4 GB for the operating system, Docker, and your chat interface.
  3. Add more if you plan on long context windows or document chat.
  4. Choose the next RAM tier up from the total, since running at the edge invites swapping.
  5. Pick 4 to 8 vCPUs for a single-user setup, more for shared use.

How to Check Your Own Server

If you already have a VPS and want to know what it can handle, a few commands answer most of it.

Check total and available memory:

free -h

check total and available memory

See your CPU model, core count, and supported instruction sets:

lscpu | grep -E ‘Model name|^CPUs|Thread|Flags’ | head -20

check cpu detail

Confirm AVX2 or AVX-512 support:

lscpu | grep -o -E ‘avx2|avx512[a-z]*’ | sort -u

check cpu detail

Check whether swap is active, which you generally want to avoid for model loading:

swapon –show

check swap

Ways to Make a Small VPS Go Further

If your server is tight on resources, a few adjustments help without spending anything.

Use a smaller context window in Ollama by setting num_ctx lower, since a huge context reserves memory you may never use. Pick a smaller or more aggressively quantized model rather than forcing a large one to fit. Keep only the models you use loaded, and remove the rest.

Disable swap or keep it minimal so a model that is too large fails loudly instead of slowing to a crawl. And close anything else running on the server that you do not need, since every gigabyte counts on a small plan.

Conclusion

Sizing a VPS for AI models is much simpler once you stop treating CPU cores as the headline number. Match RAM to the model’s quantized size plus a couple of gigabytes of breathing room, remember that memory bandwidth controls how fast responses come back, and use four to eight vCPUs for most single-user setups.

For a first project, a 16 GB server running an 8B model at Q4 is the most reliable starting point, and it gives you room to learn what your own workload actually demands before you commit to something larger. If you are ready to put a model on that server, our walkthrough on running a private ChatGPT style assistant with Ollama and Open WebUI covers the full setup from there.

Frequently Asked Questions

1. How much RAM do I need to run a local AI model on a VPS?

For a 7B to 8B model at Q4 quantization, plan on 8 GB at the very least and 16 GB to be comfortable. Larger models scale from there, with 13B to 14B models wanting 16 GB, 30B class models wanting 32 GB, and 70B models needing 64 GB or more.

2. Can I run an AI model on a VPS without a GPU?

Yes. Smaller models in the 1B to 14B range run on CPU only, with enough RAM to hold them. Responses are slower than on a GPU, often a handful of words per second for an 8B model, but it is a perfectly usable setup for personal assistants and light internal tools.

3. Does more CPU cores mean faster AI responses?

Only up to a point. Text generation is mostly limited by memory bandwidth rather than core count, so doubling your cores does not double your speed. Extra cores help most with processing long prompts and with serving several users at the same time.

4. What happens if my VPS does not have enough RAM for the model?

The model either fails to load or starts spilling into swap space on disk, which makes responses painfully slow, sometimes dropping to a fraction of a token per second. It is better to choose a smaller model that fits fully in RAM than to force a larger one onto a server that cannot hold it.

5. Is 8 GB of RAM enough to run Ollama?

It is enough for small models, such as 3B parameter models and some 7B models at Q4, but it leaves little room once the operating system, Docker, and a chat interface are running. For a smooth experience with an 8B model, 16 GB is a much safer choice.

Leave feedback about this

  • Quality
  • Price
  • Service

PROS

+
Add Field

CONS

+
Add Field
Choose Image
Choose Video