DevZero’s cover photo
DevZero

DevZero

Software Development

Seattle, WA 2,341 followers

Continuous Kubernetes optimization. No restarts. No guesswork. No toil.

About us

DevZero is an autonomous infrastructure operations platform that helps engineering teams eliminate the traditional tradeoffs between performance, reliability, and cost. By continuously optimizing compute and inference environments in real time, DevZero helps teams improve utilization, reduce cloud costs, and maintain reliable performance for stateful and AI workloads at scale. The platform continuously profiles workloads, intelligently allocates CPU, memory, and GPU resources, and places applications on the most efficient infrastructure without disrupting running services. Built by engineers who have managed large-scale Kubernetes environments, DevZero was designed to solve uptime anxiety: the tendency to overprovision infrastructure because nobody wants a 3 am incident.

Website
https://devzero.io
Industry
Software Development
Company size
11-50 employees
Headquarters
Seattle, WA
Type
Privately Held
Specialties
kubernetes, cloud, aws, eks, openshift, oci, autoscaling, live rightsizing, cost optimization, cost monitoring, and live migration

Locations

Employees at DevZero

Updates

  • DevZero reposted this

    What if Kubernetes workloads could resume instead of restart? Pod-Level Checkpoint/Restore is one of the most interesting alpha changes targeted for Kubernetes 1.37. Today, checkpointing in Kubernetes primarily operates at the individual-container level. KEP-5823 takes an important step toward treating the Pod—the unit Kubernetes actually schedules—as a whole. The first implementation introduces CheckpointPod and RestorePod methods in the Container Runtime Interface, establishing a standard way for the kubelet and container runtimes to coordinate Pod-level checkpoint and restore. This is foundational work rather than a complete live-migration workflow, but it creates a path toward faster recovery, warm starts, less disruptive node maintenance, and workload migration without rebuilding all runtime state from scratch. The broader feature is a Kubernetes community effort. One of its key implementation contributions, the new CRI methods, was authored by Radostin Stoyanov, a CRIU maintainer and member of the DevZero team. That connection is particularly relevant because DevZero already applies CRIU-based checkpoint/restore in its own migration and rightsizing workflows, preserving in-memory state and open connections without restarting the application. The CPU/GPU distinction matters too. CRIU can capture Linux process state for CPU workloads. GPU workloads require additional GPU-aware support because CUDA and device state sit outside what CRIU alone manages. Kubernetes is moving in the right direction: standardizing primitives that the wider ecosystem has already been developing and applying in practice. Still alpha, but a meaningful step forward. Read more: https://lnkd.in/dTZTPK86 CRIU: https://lnkd.in/daRhvKCY #Kubernetes #CloudNative #CRIU #CheckpointRestore #DevZero

    • No alternative text description for this image
  • DevZero reposted this

    Next Week Cloud Native Seattle is hosting deep dive into #MCP Events and Live #K8s Right Sizing. Join us for 🍕pizza, 🗣️ networking 🤝 and engineering discussions in Seattle 👉🏼 lu.ma/h69eyb40 🎤 I will be previewing a new Extension currently in design that will finally enable real-time event subscriptions in the Model Context Protocol. 🎤 Debosmit Ray (debo) (the CEO of DevZero) will be talking about Live Right Sizing in Kubernetes - how to tune resources without restarts. Space is limited. Grab your RSVP here 👉🏼 lu.ma/h69eyb40 Colleen Harig Chris Crow Kaslin Fields John Craft Natalie Fisher Josh Berkus Alex Lawrence Stephen Linker

  • DevZero reposted this

    Support for event notifications in the #MCP protocol is limited. That's about to change, and our August meetup digs into how. 👇 I will be talking about a new protocol extension currently in the design stage, that will allow AI Agents to subscribe to events happening in real world (a Slack message, a GitHub push, a 3am page) and wake up when they occur. Debosmit Ray (debo) (CEO & Co-founder of DevZero, ex-Uber infra) will also tell us about: "Live rightsizing in Kubernetes" - how resource tuning happens in k8s without restarts. Join Us 👉 lu.ma/h69eyb40 📅 Thursday, August 13 🕠 5:30 PM — Networking, pizza & drinks 📍 Docker, Inc, Seattle Maritime Building (Alaskan Way) 🍕 Food & drinks courtesy of our sponsor, DevZero

    • August Meetup - Cloud Native Seattle - A First Look at MCP Events & Kubernetes RightSizing
  • DevZero reposted this

    Cloud Native Seattle is back — Thursday, August 13 — with talks about real-time AI agents and smarter Kubernetes resource management. 🎉 🎙️ Aman Singh (Principal Engineer, Microsoft Azure CTO's office) A first look at #MCP Events — the draft protocol extension that lets #AI #agents subscribe to events happening in the real world (a Slack message, a GitHub push, a 3am page) and wake up when they occur. Straight from the MCP working group. 🎙️ Debosmit Ray (debo) (CEO & Co-founder of DevZero, ex-Uber infra) Live rightsizing in Kubernetes: tuning resource requests and limits at runtime, without pod restarts. Where VPA falls short, what "live" actually means in production, and what to watch for when you optimize continuously. 👉 RSVP - tinyurl.com/cncf-sea-04 📅 Thursday, August 13 🕠 5:30 PM — doors, networking, pizza & drinks 📍 Docker, Inc, Seattle Maritime Building (Alaskan Way) 🍕 Food & drinks courtesy of our sponsor, DevZero Run clusters or build with agents — there's something here for you. Come for the talks, stay for the networking. #MCP #AIAgents #CloudNative #Kubernetes #Seattle

    • Flyer for Cloud Native Seattle Meetup on August 13th in Seattle Maritime Building.
  • DevZero reposted this

    An Airbnb owner pays the mortgage for all 30 days in a month. In 2026, Inference is using the same setup, except the house is sitting empty most of the time. As an Airbnb guest, I'll show up for 3 nights, pay for just the 3, and then leave. Then someone else comes and stays for the other 27 days. The same exact thing is happening in inference right now, and it's not going to work long term. I send one query and use your GPU for 2 minutes, maybe 3. You can’t spin the model down when I'm done, because if you do, the next person waits 30 minutes for an answer and then leaves. So the compute sits there warm, idle, and burning all your money. On the Kubernetes side, average cluster utilization runs at 12%. Inference is running on the same substrate, except now you can't even spin a model down cleanly when traffic dips, and this creates structural waste with no off switch. The cushion right now saving everyone's butts is venture capital. As long as someone keeps the money flowing, nobody has to measure utilization at all. But that can't (and won't) last. When the VC money dries up, every inference provider gets reduced to one number: how much of their compute is actually doing paying work versus sitting hot and idle. And everyone who actually looks hates the answer.

    • No alternative text description for this image
    • No alternative text description for this image
  • DevZero reposted this

    Kubernetes Notes #4 — What actually lives inside a Pod spec The Pod spec is the only thing kubelet cares about. Everything you want to run, where it runs, how much resource it gets — it all lives here. Most people have seen this: apiVersion: v1 kind: Pod metadata:  name: nginx spec:  containers:  - name: nginx   image: nginx:1.14.2   ports:   - containerPort: 80 This works. But it is missing almost everything that matters in real usage. Here is the same Pod done properly: apiVersion: v1 kind: Pod metadata:  name: nginx  annotations:   my-tool/config: "some-value" spec:  nodeName: worker-node-1  containers:  - name: nginx   image: nginx:1.14.2   ports:   - containerPort: 80   env:   - name: ENV_NAME    value: "hello"   resources:    requests:     memory: "64Mi"     cpu: "250m"    limits:     memory: "128Mi"     cpu: "500m"   volumeMounts:   - name: shared-data    mountPath: /data  volumes:  - name: shared-data   emptyDir: {} Now let me explain every field that was added. — nodeName Pins the Pod to a specific node. Bypasses the scheduler completely. Use this when something on that node is required — like a file on disk. — annotations Key-value metadata. Kubelet ignores them. But other tools can read and act on them. This is how you pass information without touching the core Kubernetes API. — resources.requests The minimum the container needs. Kubernetes uses this to pick a node. No node has enough — Pod stays Pending forever. — resources.limits The maximum the container can use. Cross the memory limit — OOMKilled. No limit set — your container can eat the entire node. Always set this. — env Environment variables injected at start. Can come from hardcoded values, a ConfigMap, or a Secret. — volumes + volumeMounts volumes defines storage at the Pod level. volumeMounts connects it into the container at a path. Multiple containers in the same Pod can mount the same volume. The first spec is a car with no seatbelt. The second one is production ready. Both run. One is safe. --- Note #5 coming — Pod lifecycle phases. What Pending, Running, Succeeded actually mean under the hood. #GSoC #Kubernetes #CRIU #OpenSource

  • DevZero reposted this

    In our latest conversation,Between Shantanu Das ↗️ Debo from DevZero explains how checkpoint restore and live migration are changing the way platform teams think about Kubernetes, cloud costs, and AI infrastructure. We also discuss: >The reality behind "zero restart" migration >Whether autonomous infrastructure can replace human decision making >Managing unpredictable LLM traffic at scale >Building developer trust without being open source >What platform engineers and SREs should expect over the next few years If you're building or managing cloud infrastructure, this episode is packed with practical insights and forward looking perspectives. Watch the full conversation and let us know which prediction stood out to you. #PlatformEngineering #DevOps #Kubernetes #CloudNative #SRE #AIInfrastructure #CloudComputing #DeveloperTools #Infrastructure #Infrasity

  • DevZero reposted this

    Back to reality. Laptop open. Inbox full. KubeCon badge probably still somewhere in the backpack. Before normal life fully takes over, here are a few snapshots from another packed few days around Kubernetes and cloud native tech. As always, some of the best moments were not on stage. They happened in hallway conversations, booth chats, crowded coffee lines, and random detours that started with Kubernetes but quickly turned into real production problems, team challenges, architecture decisions, security, cost, and developer experience. A lot to process. A lot of new faces. A few ideas we’ll be thinking about for a while. Thanks to everyone who stopped to talk, compare notes, challenge ideas, or just say hello. See you at the next one, after we answer a few emails first. #KubeCon #Kubernetes #CloudNative

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • DevZero reposted this

    Turns out Booth #14 has a scaling problem. A good one. KubeCon Mumbai has been busy in the best possible way. Every time the booth looks like it might catch its breath, another conversation starts, another question turns into a demo, and another group stops by with a very real infrastructure problem. Not a bad problem to have. The energy has been wild, the questions have been sharp, and the conversations have gone exactly where we hoped they would. Real-world platform problems, messy infrastructure tradeoffs, AI workloads, GPUs, noisy neighbors, capacity planning, and all the things that make Kubernetes both powerful and occasionally… character-building. We’re here again today with DevZero at Booth #14. Come by, say hi, see what we’re building, grab some swag, and enter the iPhone and Mac mini raffles. Or just come tell us what’s currently haunting your cluster. We’ll understand. 📍 KubeCon + CloudNativeCon India 2026 🎪 Booth #14 📅 Here today #KubeCon #CloudNativeCon #Kubernetes #PlatformEngineering #CloudNative #AIInfrastructure

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Day 1 at #KubeCon India was a success. We had great conversations with infra teams tackling some of the toughest challenges in Kubernetes today. A few themes kept coming up: • Balancing performance, reliability, and cost • Managing AI workloads efficiently • Reducing infrastructure complexity • Getting more out of existing resources without introducing risk Most importantly, we got to meet an incredible cloud native community that’s pushing the ecosystem forward. One thing we didn’t expect? You all cleaned us out of swag. 😅 If you missed us yesterday, don’t worry. We’ll be back at the booth today and we still have our biggest giveaways left: 📱 iPhone 17 💻 Mac mini Stop by, say hello, and enter for your chance to win. Thanks to everyone who stopped by on Day 1. Looking forward to another great day in Mumbai.

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image

Similar pages

Browse jobs