Announcing KServe v0.20 - Confidential Model Serving, Traffic Splitting, Canary Rollouts, and KV Cache Offloading
Published on August 6, 2026
We are excited to announce the release of KServe v0.20. This release delivers major new capabilities across the platform with key highlights including:
- Confidential model serving with TEE-based encrypted model decryption
- Traffic splitting API for progressive LLMInferenceService deployments
- Canary rollout support for InferenceService in RawDeployment mode
- KV cache offloading with CPU memory tiering for vLLM workloads
- Distributed tracing API for LLMInferenceService components
- Managed DRA (Dynamic Resource Allocation) for GPU provisioning
- Anthropic Messages API routing support
- Native OCI ImageVolume mounting (
oci+native://) for model storage - vLLM as a standalone runtime for InferenceService
- AutoGluon server with time series inference support











