In-Depth

vMotion for Containers? How DevZero Uses CRIU to Cure Kubernetes and AI Sprawl

When KubeCon + CloudNativeCon North America lands in Salt Lake City this November (9--12), its inaugural AI Inference + Agentic track will take center stage. This is a clear indicator that Kubernetes is solidifying its role as the de facto control plane for production AI.

DevZero Kubernetes and AI infrastructure overview
[Click on image for larger view.]

In preparation for the event, I caught up with Debo Ray, co-founder and CEO of DevZero, to talk about where container infrastructure is heading. We both live in the Pacific Northwest, and after comparing notes on the twisting curves and scenic mountain peaks on the North Cascade Highway, as well as roadside oyster stands along the coast, we dove straight into the core infrastructure challenge keeping platform engineers and finance teams awake at night: the spiraling economics of Kubernetes and AI.

By way of background, Debosmit "Debo" Ray is the co-founder and CEO of DevZero, a platform engineering and developer productivity startup founded alongside fellow ex-Uber engineer Rob Fletcher. Drawing from his years as a staff infrastructure and security engineer scaling distributed systems at Uber, Ray created DevZero to eliminate developer toil. It initially replaced brittle local environments with cloud-based development environments (CDEs) and then, at customers' urging, expanded into autonomous Kubernetes and cloud resource optimization. The company emerged from stealth in early 2023, backed by $26 million from Anthos Capital and Foundation Capital. These experiences have made Ray a prominent voice in the platform engineering space, where he regularly writes and speaks on engineering efficiency, cloud architecture and modernizing developer workflows.

I have spent decades watching enterprise IT swing between centralization and sprawl. We saw it when physical servers gave way to virtual machines; we saw it again during the early days of EC2 instance sprawl; and now we are watching history repeat itself with Kubernetes clusters and AI token usage. Under the hood, solving this challenge takes far more than generic dashboards or stern, sometimes downright nasty emails from finance. It demands rethinking how container runtimes manage resources dynamically at the kernel level.

The FinOps Trap and the Tragedy of the Commons
To understand why this is happening, remember that Kubernetes was engineered from day one to deliver workload reliability and fault tolerance, not economic efficiency. When developers configure a deployment, they declare CPU and memory requests. Because nobody wants a late-night wake-up call triggered by an Out of Memory (OOM) or severe CPU throttling, engineers routinely err on the side of caution and over-provision their pods.

In infrastructure architecture, economists call this the tragedy of the commons. A shared platform cluster looks like a free, bottomless well to individual product teams. Every team optimizes its local maximum to keep its service running smoothly, leaving cluster-wide utilization hovering in the single digits, often between 5% and 6%. For every dollar committed to cloud compute, massive portions are routinely wasted on idle reservations.

Now, layer generative AI and agentic workflows directly on top of that baseline. With large language models, consumption shifted from predictable CPU cycles to variable token streams. Developers interacting with models across an organization rarely have the time, desire or motivation to calculate the downstream computing cost of long context windows or multistep agent iterations while prototyping. As Debo pointed out during our discussion, multiple enterprise engineering teams now spend significantly more on AI inference and GPU infrastructure than on their entire core cloud infrastructure.

Kubernetes resource utilization and AI infrastructure costs
[Click on image for larger view.]

Assigning post-incident blame through retrospective tagging only treats the symptoms. By the time the monthly cloud invoice arrives, that capital has already been spent. True resource efficiency requires active, programmatic intervention before and during runtime, rather than reactive accounting drills.

True Live Migration for Containers: Why CRIU Changes the Equation
The missing piece in container orchestration has always been non-destructive runtime mobility. In the virtualization world, we take techniques like VMware vSphere vMotion and KVM live migration for granted. When a host runs hot or requires maintenance, the hypervisor transfers the guest memory and CPU state to another host with zero dropped connections or application downtime.

Kubernetes was not designed for this capability and has historically lacked it. If you want to right-size a pod to smaller resource boundaries or drain a worker node, the control plane sends a SIGTERM, waits for a timeout, issues a SIGKILL and schedules a fresh pod somewhere else. For stateful applications, database caches or large inference workloads, a cold start means dropped network connections, rebuilt local caches and costly latency penalties.

DevZero tackled this architectural limitation by leveraging CRIU (Checkpoint/Restore in Userspace). By freezing the container process hierarchy, checkpointing memory pages and active socket descriptors to disk, and restoring them on another target machine, DevZero migrates running workloads without service interruption, closing the long-standing gap between hypervisors and container engines.

Preserving transport state across hosts is where container migration usually breaks down. When a pod migrates to a different worker node, its network identity must follow it smoothly without breaking active TCP sessions or confusing upstream callers. DevZero handles this behind the scenes using low-level netfilter and iptables routing tricks to fool existing connections into believing they are communicating with the exact same uninterrupted socket.

DevZero container migration and network continuity
[Click on image for larger view.]

Work is also moving forward to bring these native capabilities deeper into the upstream Kubernetes project, targeting checkpoint-restore improvements across the container runtime interface (CRI-O and containerd). Bringing core CRIU committers onto the internal team has let DevZero push non-disruptive rightsizing beyond theoretical benchmarks and into active production clusters.

Real-World Impact: Network Topology and AZ Outages
When resource optimization factors in actual application network topology, the operational and financial payoffs compound immediately.

A compelling case study Debo shared involves DataBahn, a fast-growing analytics and logging company processing massive data volumes. As DataBahn scaled, cross-availability zone (AZ) network transfer costs on AWS began eroding gross margins. Public cloud providers levy substantial egress charges when microservices exchange heavy traffic across availability zone boundaries.

By deploying custom telemetry operators, DevZero mapped the inter-service communication patterns of DataBahn's workloads. The platform dynamically scheduled tightly coupled dependencies into the same physical availability zone and, when feasible, co-located them on the same node. The result was lower latency and an overall spend reduction exceeding 70% on AWS and 60% on Azure.

DevZero workload scheduling across availability zones
[Click on image for larger view.]

Consolidating into a single AZ, of course, introduces obvious availability risks if that zone fails. In one case, during a subsequent localized AWS infrastructure event, real-time telemetry alerted DevZero to rising failure rates. Within 2 minutes, the platform leveraged its checkpoint-restore pipeline to live-migrate critical pods over to healthy availability zones. Network traffic costs briefly ticked upward while running across zones, but the applications stayed online throughout the outage without dropping active client connections.

The GPU Dilemma and the Next Step in Agentic AI
Taking checkpoint-restore into the GPU landscape introduces a whole new set of technical hurdles. Moving standard system memory across physical nodes, as we do with VMs, is a solved problem. Moving gigabytes of CUDA device context, memory pointers and tensor states from an NVIDIA H100 or H200 card is a completely different, and ugly beast.

Standard CRIU dumps have no native visibility into GPU hardware state. To address this, DevZero has integrated emerging capabilities such as CRIUgpu and NVIDIA's cuda-checkpoint APIs. These tools coordinate process locks, flush remaining kernels, capture unified CPU-GPU device memory to the host file system and reconstruct the GPU context on a target node.

This technology directly targets the cold-start barrier that currently prevents AI inference workloads from running on dynamic spot instances. Loading an open-weights model like DeepSeek or LLaMA from cold storage into high-bandwidth GPU memory takes substantial time. By restoring from pre-warmed, serialized checkpoints, teams can spin up inference instances 6 to 10 times faster while slashing idle reservation waste.

DevZero GPU checkpointing for AI inference workloads
[Click on image for larger view.]

Looking ahead to KubeCon + CloudNativeCon, the cloud-native ecosystem is expanding rapidly around inference routing and governance. DevZero is rolling out an inference aggregation and model routing platform that dynamically evaluates cost, latency and performance across hundreds of external serverless providers, automatically routing incoming requests to the optimal endpoint.

Final Thoughts
As organizations deploy complex autonomous agents that spin up infrastructure, generate code and execute multitier workflows on the fly, robust runtime observability and non-disruptive migration are no longer optional luxuries. Without autonomous, kernel-level efficiency built directly into the container fabric, cloud and AI budgets will simply continue to outpace the value they deliver.

DevZero will be at KubeCon 2026, so stop by their booth to talk to Debo and the rest of the folks from DevZero.

About the Author

Tom Fenton has a wealth of hands-on IT experience gained over the past 30 years in a variety of technologies, with the past 20 years focusing on virtualization and storage. He previously worked as a Technical Marketing Manager for ControlUp. He also previously worked at VMware in Staff and Senior level positions. He has also worked as a Senior Validation Engineer with The Taneja Group, where he headed the Validation Service Lab and was instrumental in starting up its vSphere Virtual Volumes practice. He's on X @vDoppler.

Featured

Subscribe on YouTube