Sechno
Devops

Practical Guide: Deploying Llama 3.2 with Ollama on Small Kubernetes Nodes

Step-by-step, pragmatic blueprint for running Llama 3.2 via Ollama on low-cost VPS or small Kubernetes clusters: install, manifest examples, client calls, sizing tips, and tradeoffs for reliable inference.

SSechno Team 5 min read 49 views
Practical Guide: Deploying Llama 3.2 with Ollama on Small Kubernetes Nodes

Why self-host Llama 3.2 with Ollama?

Self-hosting a modern open-weight LLM with a lightweight inference layer such as Ollama gives you control over latency, data privacy, model versions, and cost predictability. This guide focuses on practical steps to get a working inference endpoint on a small VPS or a compact Kubernetes cluster, plus actionable tips for productionizing inference safely.

Prerequisites

  • Access to a Linux VPS or Kubernetes cluster with at least 4–8GB RAM per worker for small models (adjust for model size).
  • Docker and kubectl access (or a managed Kubernetes offering).
  • Basic familiarity with kubectl, YAML manifests, and networking (ingress/load balancer).
  • Ollama knowledge: Ollama provides a simple server that runs models and exposes a local API. See the official docs for latest commands: Ollama docs.

Architecture overview

Keep the architecture simple at first: one Ollama container (the inference process) per node, model files on local disk or PV, and a fronting service (Ingress / LoadBalancer) for client requests. Add autoscaling and batching later.

Key components

  • Ollama inference server (container)
  • Model storage (local disk, hostPath, or PersistentVolume)
  • Service/Ingress for client traffic
  • Horizontal Pod Autoscaler (HPA) or external scaler for traffic spikes

Quick start: single-node (VPS) using Docker

On a single VPS you can test quickly with Docker and Ollama's CLI. Replace example model names with the one you want.

sudo apt update && sudo apt install -y docker.io curl
sudo usermod -aG docker $USER
# (log out and back in for group to apply)
 
# Pull and run Ollama (example image - check official repo for current image)
docker run --rm -it \
  -p 11434:11434 \
  -v /opt/ollama:/root/.ollama \
  ollama/ollama:latest
 
# In another shell, pull a model via Ollama CLI (inside container or using installed CLI):
# ollama pull llama-3.2
 
# Simple curl test (assuming model 'llama-3.2' is available and the server exposes a REST endpoint):
curl -X POST "http://localhost:11434/v1/completions" \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.2","prompt":"Write a short haiku about code."}'

Deploying on Kubernetes: minimal manifests

Below is a minimal Deployment + Service. This example mounts a model directory and runs Ollama. Adapt resource requests/limits for your model and cluster.

apiVersion: v1
kind: Namespace
metadata:
  name: ollama
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: ollama
  namespace: ollama
spec:
  replicas: 1
  selector:
    matchLabels:
      app: ollama
  template:
    metadata:
      labels:
        app: ollama
    spec:
      containers:
      - name: ollama
        image: ollama/ollama:latest
        ports:
        - containerPort: 11434
        volumeMounts:
        - name: models
          mountPath: /root/.ollama/models
        resources:
          requests:
            memory: "4Gi"
            cpu: "1000m"
          limits:
            memory: "8Gi"
            cpu: "2000m"
      volumes:
      - name: models
        emptyDir: {}
---
apiVersion: v1
kind: Service
metadata:
  name: ollama
  namespace: ollama
spec:
  type: ClusterIP
  selector:
    app: ollama
  ports:
  - port: 11434
    targetPort: 11434

Apply with kubectl apply -f. For external access, add an Ingress or change the Service type to LoadBalancer.

Load, scaling, and autoscaling

Start with conservative resources and measure. Add a HorizontalPodAutoscaler that scales on CPU or custom metrics (requests in queue). Example HPA (CPU):

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: ollama-hpa
  namespace: ollama
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: ollama
  minReplicas: 1
  maxReplicas: 3
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 60

Notes:

  • If you run large models, GPU nodes are often mandatory — use nodeSelectors/taints and the NVIDIA device plugin.
  • Ollama may support model-specific optimizations (quantized GGML weights, int8) — prefer smaller quantized variants for CPU-only deployments.

Calling the model from your app

Most Ollama setups expose a JSON HTTP completions endpoint. Example curl (adjust for your URL):

curl -sS -X POST "http://ollama.example.com/v1/completions" \
  -H "Content-Type: application/json" \
  -d '{"model":"llama-3.2","prompt":"Summarize the following text...","max_tokens":150}'

Wrap calls with a retry/backoff and timeouts. For throughput, use a request queue or an API gateway to apply rate limits and circuit breakers.

Operational tips and tradeoffs

  • Model size vs latency: Larger models (bigger parameter counts) produce better quality at higher compute cost and latency. Prefer quantized or distilled variants for low-cost inference.
  • Batching: If latency can be relaxed, batch requests on the server side to increase throughput per GPU/CPU cycle.
  • GPU vs CPU: GPUs greatly reduce latency for large models. On small clusters without GPUs, choose smaller or quantized models.
  • Storage: Use fast local NVMe for model storage when possible. Network-mounted models can add startup latency.
  • Security: Run the inference service in a restricted network namespace. Do not expose it publicly without authentication or a gateway layer.
  • Updates & reproducibility: Pin Ollama image and model versions. Have a rollout strategy (canary) for model or runtime updates.

Debugging checklist

  1. Check Pod logs for model load errors and OOM kills (kubectl logs).
  2. Inspect events for scheduling issues (insufficient resources).
  3. Verify model files exist at the mounted path inside the container.
  4. Measure CPU / memory and end-to-end latency; adjust requests/limits and HPA thresholds accordingly.

When to use managed inference instead

Self-hosting is great for control and privacy. Consider managed LLM services when you need:

  • Predictable SLA and automatic scaling without cluster ops
  • Less operational overhead for GPU resource management
  • Compliance and audit tooling bundled with the service

Conclusion

Deploying Llama 3.2 with Ollama on small nodes is an achievable way to get private, controllable inference. Start small: choose an appropriately sized model, pin images and versions, expose the service behind a gateway with authentication, and add autoscaling and batching as usage grows. Measure continuously and prefer quantized models or GPUs depending on latency and cost constraints.

Further reading

Original walkthrough used as a starting point for practical steps: How to deploy Llama 3.2 with Ollama + Kubernetes. Always consult the latest Ollama docs for CLI flags and deployment changes: Ollama docs.

Was this helpful?

Share this post

Comments (0)

Want to join the conversation?

Log in or sign up to leave a comment and share your thoughts.

Log in to Comment