Why self-host LLMs?
Running open-source models on your own GPU droplet gives you predictable per-request latency, reduced long-term API spend for heavy workloads, data control, and the freedom to use quantized model variants. This guide focuses on a practical, production-ready stack: vLLM for efficient GPU batching and scheduling, and KServe to expose a Kubernetes-native inference API.
High-level architecture
- Model store: Persistent volume (PVC) with model files (float or quantized format).
- vLLM predictor: Container running vLLM (or another optimized runtime) to load the model and serve inference requests.
- KServe: Manages InferenceService CRs, autoscaling, routing, and gives a standard REST/gRPC endpoint.
- Autoscaling & monitoring: HPA/KNative (optional), Prometheus metrics from the predictor, and logging to a central store.
Before you start — practical checklist
- Pick a GPU instance with enough VRAM for your target model and quantization (e.g., 8–24GB depending on model+quant).
- Decide quantization strategy (AWQ / QLoRA / GPTQ / ggml) — tradeoffs below.
- Ensure Kubernetes cluster has GPU nodes and the NVIDIA device plugin installed.
- Provision a PVC or object store accessible from the pod for model files.
- Prepare a small canary workload and latency budgets to validate performance.
Example: KServe InferenceService for a vLLM-based predictor
The YAML below shows a minimal InferenceService that mounts a model from a PVC and requests a GPU. Adapt the image and args to your vendor/runtime image.
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: mistral-vllm
namespace: ml-prod
spec:
predictor:
custom:
container:
image: your-registry/vllm-kserve:latest
args:
- "--model-path=/models/mistral-7b"
- "--port=8080"
resources:
limits:
nvidia.com/gpu: "1"
requests:
cpu: "2"
memory: "8Gi"
volumeMounts:
- name: model-volume
mountPath: /models
components:
- name: model-storage
persistentVolumeClaim:
claimName: mistral-model-pvcNotes:
- Use a production-ready container image that runs vLLM (or your chosen runtime) and exposes the KServe-compatible REST/gRPC endpoint.
- Use resource requests & limits to ensure scheduler places the pod on a GPU node and the node has adequate CPU/memory.
- For high availability, run at least two instances and front them with KServe routing or an external load balancer.
Python: calling the KServe endpoint
Once KServe exposes the model, your app can call the predictor via REST. Example client:
import requests
# Replace with your cluster address or an ingress endpoint
url = "http://mistral-vllm-predictor.ml-prod.svc.cluster.local/v1/models/mistral/predict"
payload = {"instances": ["Summarize the following text concisely: ..."]}
resp = requests.post(url, json=payload, timeout=60)
resp.raise_for_status()
print(resp.json())Model preparation & quantization (practical notes)
Quantization reduces VRAM and inference cost but may affect output quality. Common approaches:
- GPTQ / AWQ / AWQ-like: Good VRAM savings with modest accuracy impact when done properly.
- 8-bit (bitsandbytes): Widely used with PyTorch-based runtimes; lower accuracy loss but needs compatible runtime.
- CPU-optimized formats (ggml): Useful for CPU inference but increases latency compared to GPU runs.
Typical conversion steps (conceptual):
- Download the HF model weights to a staging VM or container with sufficient disk.
- Run the chosen quantization/conversion tool to produce artifact files in the format your runtime supports.
- Store the converted model in the PVC or an object store and update your InferenceService to point at it.
Because conversion tools vary, keep a small staging workflow that validates model outputs (sanity prompts + unit tests) before pushing to production.
Operational considerations & tradeoffs
- Latency vs. Cost: Single-GPU inference is lowest-latency; batching in vLLM can increase throughput but adds tail latency. If strict p95 latency is required, tune batch sizes and max wait times.
- Model size vs. quantization: More aggressive quantization reduces VRAM and cost but can degrade quality. Measure against a small evaluation set representative of your app.
- Autoscaling: KServe integrates with Kubernetes autoscaling but GPU autoscaling is coarse. Consider request queuing, pre-warming, or a hybrid strategy (small fleet of GPUs + CPU fallback for lightweight tasks).
- Reliability: Keep model artifacts immutable and use versioned paths. Use readiness probes that validate the runtime has loaded the model before routing traffic.
- Security & data privacy: Ensure network policies and ingress rules prevent exposure to unauthorized networks. If sending user data to the model, plan encryption and retention policies.
Monitoring & SLOs
- Expose metrics: request latency (p50/p95/p99), GPU utilization, batch sizes, queue length, OOM events.
- Use Prometheus + Grafana and set alerts for high queue length or GPU memory pressure.
- Run periodic output-quality checks if you retrain or switch quantized artifacts.
Small reproducible workflow (quick dev loop)
- Develop & test locally using a small model or CPU-only build.
- Convert and validate the model in a staging namespace on the cluster.
- Deploy a single KServe canary instance behind an ingress and run load tests to collect p95 and throughput.
- Iterate on batch settings and quantization until you meet SLOs, then promote to production.
Conclusion
Self-hosting open LLMs with vLLM + KServe is a practical path to predictable latency and cost control for production inference. The key engineering tasks are: choose the right quantization/format for your accuracy/VRAM tradeoff, provision GPU nodes with correct resource requests, validate models with representative tests, and instrument robust monitoring. Start small (single canary GPU), measure p50/p95 latency and output quality, and iterate on batching and scaling policies.
Useful further reading
Original walkthroughs and community posts on deploying vLLM + KServe and inexpensive GPU droplets informed this guide. See the referenced deployment post for concrete container images and hands-on examples: Deploy Mistral 7B with vLLM + KServe on a $10/month GPU droplet.
Was this helpful?
Share this post
Comments (0)
Want to join the conversation?
Log in or sign up to leave a comment and share your thoughts.
Log in to Comment