Why self-host an open LLM?
Self-hosting an open model (for example, Llama 2 or a community GGUF variant) gives you predictable costs, full data control, and the ability to tune runtime behavior. This guide focuses on practical, repeatable patterns that work on low-cost VPS instances (CPU-first) and small GPU machines.
High-level pattern
- Choose model and runtime: CPU-friendly (llama.cpp bindings) or GPU-focused (TGI, vLLM).
- Quantize to fit memory and improve throughput where possible.
- Expose a small authenticated API that proxies to the local runtime.
- Containerize or run as a systemd service; protect with a reverse proxy and rate limits.
Pick a runtime and understand tradeoffs
- llama.cpp / llama-cpp-python: excellent for CPU-only or low-RAM setups when using GGUF quantized files. Lower throughput but predictable costs.
- text-generation-inference (TGI) / vLLM: targets GPU inference for higher throughput and lower latency, but needs a GPU and more complex setup.
- Tradeoffs: CPU setups: cheap and portable, higher latency. GPU setups: faster and more concurrent, more expensive and needs driver/tooling upkeep.
Minimal secure API: Flask + llama-cpp-python
The example below shows a small authenticated API that calls a local llama-cpp binding. It’s intentionally minimal so you can adapt it to FastAPI, Django, or a Go service later.
from flask import Flask, request, jsonify
from llama_cpp import Llama
app = Flask(__name__)
API_KEY = "REPLACE_WITH_STRONG_KEY"
# point this at your local quantized model file (GGUF / ggml), downloaded and licensed appropriately
llm = Llama(model_path="/srv/models/llama2-13b.gguf")
@app.route('/v1/generate', methods=['POST'])
def generate():
auth = request.headers.get('Authorization', '')
if auth != f"Bearer {API_KEY}":
return jsonify({'error': 'unauthorized'}), 401
payload = request.json or {}
prompt = payload.get('prompt', '')
if not prompt:
return jsonify({'error': 'prompt required'}), 400
# tune max_tokens / temperature for your use case
resp = llm.create(prompt=prompt, max_tokens=256, temperature=0.7)
# response structure depends on binding version; check your installed package
text = resp.get('choices', [{}])[0].get('text', '')
return jsonify({'text': text})
if __name__ == '__main__':
app.run(host='0.0.0.0', port=8080)Notes
- llama-cpp-python may require native build tools during installation. Use a Docker image that contains build deps for repeatable builds.
- Store API keys in environment variables or a secrets manager in production; never hardcode them.
Dockerize the API
A simple Dockerfile pattern that installs system deps, Python packages, and runs a WSGI server.
FROM python:3.11-slim
RUN apt-get update \
&& apt-get install -y build-essential cmake libopenblas-dev git \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
CMD ["gunicorn", "-b", "0.0.0.0:8080", "app:app", "--workers", "1"]requirements.txt example: flask, llama-cpp-python, gunicorn.
Reverse proxy and header injection (nginx)
Use a reverse proxy to terminate TLS, enforce rate limits, and inject the Authorization header so backends never see public API keys.
server {
listen 443 ssl;
server_name llm.example.com;
location /v1/ {
proxy_pass http://127.0.0.1:8080/;
proxy_set_header Authorization "Bearer REPLACE_WITH_STRONG_KEY";
proxy_set_header Host $host;
}
}Deploy patterns
- Single small VPS (CPU-only): good for experimentation and tiny workloads; use quantized GGUF models and limit concurrency.
- GPU instance: use vLLM or TGI in containers; prefer 24GB+ GPU memory for mid-sized models unless heavily quantized.
- Autoscaling: for production load, front the API with aingress/load-balancer and spin GPU instances behind an autoscaler; keep a CPU fallback for low-cost burst handling.
Quantization and model size tradeoffs
- Quantization reduces memory at the cost of some model quality. 8-bit and 4-bit quant formats are common. Test output on representative prompts before committing.
- Smaller parameters/models reduce latency and RAM needs. If you must support many concurrent users, prefer smaller or quantized models over a single large unquantized model.
- Beware compatibility: runtime bindings expect specific file formats (GGUF, GGML) — convert with community tools where available.
Operational checklist
- Verify model license and redistribution rules before downloading/hosting (deployment write-up for reference).
- Run local load tests to determine concurrency and CPU/RAM demands.
- Limit request sizes and tokens per request to avoid resource exhaustion.
- Monitor process memory, CPU, and disk; rotate logs and clear caches periodically.
- Use HTTPS and short-lived API keys or a proper auth layer for production.
Example systemd unit for a single-instance deployment
[Unit]
Description=LLM API
After=network.target
[Service]
User=www-data
WorkingDirectory=/srv/llm-api
ExecStart=/usr/bin/gunicorn -b 0.0.0.0:8080 app:app --workers 1
Restart=on-failure
[Install]
WantedBy=multi-user.targetTesting and local development
- Run the Flask app locally with smaller quantized models first.
- Use simple curl tests and a short prompt to verify the pipeline:
curl -s -X POST https://llm.example.com/v1/generate \
-H 'Content-Type: application/json' \
-d '{"prompt": "Write a one-line commit message: Fix issue with login"}'Common pitfalls
- Underestimating memory: a single request may inflate memory usage; set concurrency limits.
- Ignoring licensing rules for model weights; check source/host policies.
- Assuming CPU inference will match cloud latency guarantees; always load-test against your expected traffic.
Conclusion
Self-hosting open LLMs on a budget VPS is practical for experimentation and low-volume production use. The core pattern is: pick a runtime that matches your hardware, quantize when needed, expose a minimal and authenticated API, and run behind a reverse proxy for TLS and rate limits. Start small, measure, and scale horizontally or move to GPU instances as demand grows.
Further reading
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment