Overview
This guide shows a practical, repeatable approach to run open-source LLMs (for example Llama 2 variants) on low-cost cloud virtual machines using CPU-optimized runtimes and quantized GGUF models. You'll get concrete steps to provision a VM, prepare a quantized model, run inference with llama.cpp style runtimes, and expose a minimal HTTP API for development. Tradeoffs, performance caveats, and security notes are included so you can pick the right setup for your needs.
What this article covers
- Which model sizes and runtimes are reasonable for cheap VMs
- How to build a CPU-optimized runtime and use pre-quantized models
- A minimal Flask wrapper that exposes a simple HTTP API
- Operational tradeoffs and security hardening pointers
High-level tradeoffs
- Model size: smaller models (7B-ish) can often run on CPU with quantization; larger (13B+) typically need more RAM or a GPU for acceptable latency.
- Runtime: llama.cpp-style runtimes are best for low-cost CPU inference using GGUF quantized models. Web UI stacks (text-generation-webui) add convenience but increase resource use.
- Latency vs cost: restarting a binary per request is easy but slow. Keeping a persistent process or using a dedicated server/runtime reduces latency but uses more memory.
Choose model and runtime
Pick a model and runtime that match your constraints. Practical options:
- llama.cpp + GGUF quantized model — good for CPU inference on cheaper VMs.
- text-generation-webui — adds a web UI and API but needs more memory.
- vLLM / Triton / Rayserve — high-performance options for production with GPU clusters.
Important: always verify the model license before deployment and, when possible, download pre-quantized GGUF artifacts to avoid lengthy conversions on a tiny VM.
Step-by-step: from VM to API
1) Provision an instance
Choose a cloud VM with enough RAM for your model. As a rule of thumb, a quantized 7B GGUF often needs single-digit GBs of RAM plus swap; a 13B model will need significantly more. If you expect higher throughput or lower latency, plan for persistent processes or GPU-enabled instances.
2) Install system packages and build a CPU runtime
Example commands to install essentials and build a C++-based CPU runtime (replace package manager commands for your distro).
sudo apt update && sudo apt install -y build-essential git cmake python3 python3-venv curl unzipgit clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
make -j"$(nproc)"The repository build produces a binary (commonly named main) you can use to run GGUF models.
3) Get a quantized GGUF model
For a small VM, use a pre-quantized GGUF file when possible. Download it to models/. Example (replace URL with a valid model file host):
mkdir -p ~/models
cd ~/models
curl -L -o model.gguf "https://example.com/path/to/pre-quantized-model.gguf"If you must convert, perform conversion on a larger machine and transfer the GGUF to the VM to avoid CPU-heavy work on a tiny instance.
4) Quick test from the VM (CLI)
A simple CLI invocation to run a prompt with the built binary. This starts a one-shot inference process.
cd ~/llama.cpp
./main -m ~/models/model.gguf -p "Write a short greeting in plain text." -n 128That command launches the binary, supplies a prompt, and prints generated tokens to stdout. One-shot invocation is simple but has high per-request latency because it loads the model each time.
5) Minimal HTTP API wrapper (development-ready)
For rapid integration, use a tiny Python Flask app that calls the binary. This example restarts the binary per request (simple, but slow). Use a persistent server or a socket-based runtime for production.
from flask import Flask, request, jsonify
import subprocess
app = Flask(__name__)
@app.route('/generate', methods=['POST'])
def generate():
data = request.get_json(force=True)
prompt = data.get('prompt', '')
if not prompt:
return jsonify({'error': 'prompt required'}), 400
# Run the llama.cpp binary as a simple one-shot invocation
cmd = ['./main', '-m', '/home/ubuntu/models/model.gguf', '-p', prompt, '-n', '128']
try:
res = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
output = res.stdout.strip() or res.stderr.strip()
return jsonify({'output': output})
except subprocess.TimeoutExpired:
return jsonify({'error': 'inference timeout'}), 504
if __name__ == '__main__':
app.run(host='0.0.0.0', port=8080)Query it locally with curl:
curl -X POST localhost:8080/generate -H 'Content-Type: application/json' -d '{"prompt":"Say hello"}'Notes on improving this wrapper:
- Keep the model resident in memory using a persistent process or a long-running runtime (avoid reloading model on every request).
- Use a supervisor (systemd, tmux, or a container) to restart on crashes and control resources.
- For concurrency, consider a queue and worker pool or integrate with a runtime that supports multiple streams.
Operational recommendations
- Swap and cgroups: Configure swap cautiously: it can prevent OOM but greatly increases latency. Use cgroups/limits to prevent noisy-neighbor issues on shared hosts.
- Security: Put the API behind a reverse proxy (Nginx), restrict access with authentication, and enable a firewall to limit exposure. Avoid exposing inference APIs publicly without rate limits and auth.
- Monitoring: Capture memory, CPU, request latencies, and error rates. Start small and scale the instance size if latency or memory errors appear.
- Persistence: For production, run a persistent inference server (vLLM or Triton for GPUs, or a long-running llama.cpp-based server) to avoid model reload overhead.
When to choose GPU or managed services
If your workload requires low latency, high throughput, or larger model sizes (13B+), a GPU-backed instance or a managed inference platform is usually more cost-effective than forcing large models onto tiny CPU VMs. Use low-cost CPU VMs for development, prototyping, or low-traffic edge use cases.
Licensing and compliance
Confirm model licensing and attribution requirements for the model variant you use. Self-hosting carries responsibilities for user data handling, logging, and privacy; ensure your deployment aligns with legal and organizational policies.
Conclusion
Running open LLMs like Llama 2 on low-cost cloud VMs is practical for prototyping and light workloads when you:
- Choose the right model size and a quantized GGUF artifact
- Use CPU-optimized runtimes (llama.cpp) or persistent servers for production
- Harden and monitor the service (firewall, auth, resource limits)
For higher throughput, lower latency, or larger models, migrate to GPU instances or managed inference solutions. Start with a small VM for development and iterate toward a production architecture that matches your SLA.
Further reading: a practical step-through that inspired this workflow is available on DEV Community (see source notes below).
Tip: If you want to avoid cloud entirely, the same approach works on a local workstation with enough RAM and a compatible CPU; replace cloud VM provisioning with local virtualization or containers.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment