Why self-host Llama 3.2 with Ollama?
Self-hosting a modern open-weight LLM with a lightweight inference layer such as Ollama gives you control over latency, data privacy, model versions, and cost predictability. This guide focuses on practical steps to get a working inference endpoint on a small VPS or a compact Kubernetes cluster, plus actionable tips for productionizing inference safely.
Prerequisites
- Access to a Linux VPS or Kubernetes cluster with at least 4–8GB RAM per worker for small models (adjust for model size).
- Docker and kubectl access (or a managed Kubernetes offering).
- Basic familiarity with kubectl, YAML manifests, and networking (ingress/load balancer).
- Ollama knowledge: Ollama provides a simple server that runs models and exposes a local API. See the official docs for latest commands: Ollama docs.
Architecture overview
Keep the architecture simple at first: one Ollama container (the inference process) per node, model files on local disk or PV, and a fronting service (Ingress / LoadBalancer) for client requests. Add autoscaling and batching later.
Key components
- Ollama inference server (container)
- Model storage (local disk, hostPath, or PersistentVolume)
- Service/Ingress for client traffic
- Horizontal Pod Autoscaler (HPA) or external scaler for traffic spikes
Quick start: single-node (VPS) using Docker
On a single VPS you can test quickly with Docker and Ollama's CLI. Replace example model names with the one you want.
sudo apt update && sudo apt install -y docker.io curl
sudo usermod -aG docker $USER
# (log out and back in for group to apply)
# Pull and run Ollama (example image - check official repo for current image)
docker run --rm -it \
-p 11434:11434 \
-v /opt/ollama:/root/.ollama \
ollama/ollama:latest
# In another shell, pull a model via Ollama CLI (inside container or using installed CLI):
# ollama pull llama-3.2
# Simple curl test (assuming model 'llama-3.2' is available and the server exposes a REST endpoint):
curl -X POST "http://localhost:11434/v1/completions" \
-H "Content-Type: application/json" \
-d '{"model":"llama-3.2","prompt":"Write a short haiku about code."}'Deploying on Kubernetes: minimal manifests
Below is a minimal Deployment + Service. This example mounts a model directory and runs Ollama. Adapt resource requests/limits for your model and cluster.
apiVersion: v1
kind: Namespace
metadata:
name: ollama
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: ollama
namespace: ollama
spec:
replicas: 1
selector:
matchLabels:
app: ollama
template:
metadata:
labels:
app: ollama
spec:
containers:
- name: ollama
image: ollama/ollama:latest
ports:
- containerPort: 11434
volumeMounts:
- name: models
mountPath: /root/.ollama/models
resources:
requests:
memory: "4Gi"
cpu: "1000m"
limits:
memory: "8Gi"
cpu: "2000m"
volumes:
- name: models
emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
name: ollama
namespace: ollama
spec:
type: ClusterIP
selector:
app: ollama
ports:
- port: 11434
targetPort: 11434Apply with kubectl apply -f. For external access, add an Ingress or change the Service type to LoadBalancer.
Load, scaling, and autoscaling
Start with conservative resources and measure. Add a HorizontalPodAutoscaler that scales on CPU or custom metrics (requests in queue). Example HPA (CPU):
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: ollama-hpa
namespace: ollama
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: ollama
minReplicas: 1
maxReplicas: 3
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60Notes:
- If you run large models, GPU nodes are often mandatory — use nodeSelectors/taints and the NVIDIA device plugin.
- Ollama may support model-specific optimizations (quantized GGML weights, int8) — prefer smaller quantized variants for CPU-only deployments.
Calling the model from your app
Most Ollama setups expose a JSON HTTP completions endpoint. Example curl (adjust for your URL):
curl -sS -X POST "http://ollama.example.com/v1/completions" \
-H "Content-Type: application/json" \
-d '{"model":"llama-3.2","prompt":"Summarize the following text...","max_tokens":150}'Wrap calls with a retry/backoff and timeouts. For throughput, use a request queue or an API gateway to apply rate limits and circuit breakers.
Operational tips and tradeoffs
- Model size vs latency: Larger models (bigger parameter counts) produce better quality at higher compute cost and latency. Prefer quantized or distilled variants for low-cost inference.
- Batching: If latency can be relaxed, batch requests on the server side to increase throughput per GPU/CPU cycle.
- GPU vs CPU: GPUs greatly reduce latency for large models. On small clusters without GPUs, choose smaller or quantized models.
- Storage: Use fast local NVMe for model storage when possible. Network-mounted models can add startup latency.
- Security: Run the inference service in a restricted network namespace. Do not expose it publicly without authentication or a gateway layer.
- Updates & reproducibility: Pin Ollama image and model versions. Have a rollout strategy (canary) for model or runtime updates.
Debugging checklist
- Check Pod logs for model load errors and OOM kills (
kubectl logs). - Inspect events for scheduling issues (insufficient resources).
- Verify model files exist at the mounted path inside the container.
- Measure CPU / memory and end-to-end latency; adjust requests/limits and HPA thresholds accordingly.
When to use managed inference instead
Self-hosting is great for control and privacy. Consider managed LLM services when you need:
- Predictable SLA and automatic scaling without cluster ops
- Less operational overhead for GPU resource management
- Compliance and audit tooling bundled with the service
Conclusion
Deploying Llama 3.2 with Ollama on small nodes is an achievable way to get private, controllable inference. Start small: choose an appropriately sized model, pin images and versions, expose the service behind a gateway with authentication, and add autoscaling and batching as usage grows. Measure continuously and prefer quantized models or GPUs depending on latency and cost constraints.
Further reading
Original walkthrough used as a starting point for practical steps: How to deploy Llama 3.2 with Ollama + Kubernetes. Always consult the latest Ollama docs for CLI flags and deployment changes: Ollama docs.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment