Skip to main content

Make cold starts faster

When you trigger a prod or staging run in the Trigger.dev cloud it takes 3 seconds on average for the machine to start up and your code to start executing. This same slowness happens when a run resumes, like after using wait.for if the delay is above a threshold when we shut the machine down.

They take this long because we're using Kubernetes for our cluster and a new pod takes a while to come up.

We're switching to using MicroVMs for the cloud machines. Our target is to get the p95 for starts and resumes to under 500ms.

Status: In Progress18 comments

Comments18

  • Anonymous User

    •

    Jul 30, 2024

    Will the open source version also use Firecracker?

    • mattaitken

      Team•

      Aug 1, 2024

      It won’t be open source for the foreseeable future, for a couple of reasons:

      1. Deploying it is a lot harder than just running some Docker compose files, or using Kubernetes. Not necessarily because hosting a single Firecracker instance is hard but because of all the other requirements with networking, cached storage of snapshots, fast scaling up and down, etc.

      2. Very fast run starts will be a differentiator for our cloud product. In the future we will offer on-prem under a separate commercial license agreement and that would include Firecracker too.

      • Vitrion B.V.

        •

        Aug 23, 2024

        Aah thats a bummer 😔, right now the thing that “bothers” me most is a simple console.log hello taking above 400 ms every call

        • mattaitken

          Team•

          Aug 28, 2024

          @Vitrion B.V. currently it takes about 4 seconds on average for a run to start in v3.

          I’m confused about your console.log example. When the run is executing the code is as fast as any Node.js environment and there are no delays. Where are you seeing this 400ms delay on every call?

  • Mitchell Sigley

    •

    Jan 9, 2025

    Is there an updated ETA on this? Will there be much migration required?

  • An Anonymous User

    •

    Apr 8, 2025

  • An Anonymous User

    •

    Jul 28, 2025

    Hello!

  • An Anonymous User

    •

    Aug 20, 2025

    Test

  • An Anonymous User

    •

    Dec 2, 2025

    When do you anticipate this feature will be done. The blog post I read from September suggested it would be complete soon?

  • An Anonymous User

    •

    Dec 18, 2025

    MicroVMs + Bun runtime should destroy slow cold starts! 🚀

  • danielwyb

    •

    Dec 29, 2025

    What’s the ETA on this?

  • An Anonymous User

    •

    Mar 9

    Here’s a concise follow-up you could post that pushes on priority + ETA and references Firecracker-based alternatives in a constructive way:


    Thanks for the update, this improvement would have a very large impact for us.

    At Omnia we run time-sensitive agentic workflows, and each run is triggered via Trigger.dev. If the cold start is ~3 seconds, the perceived responsiveness of the product drops dramatically, because the user is waiting for the workflow to even begin.

    A sub-500ms p95 would make a huge difference.

    Do you have any sense of:

    • Priority of this migration internally?

    • Rough ETA for when MicroVM-based machines might reach production?

    We’re asking because we’re evaluating architectures for workloads that need very fast startup times, and many platforms are moving to Firecracker-based microVMs to solve this exact problem.

    For example:

    • AWS Lambda moved to Firecracker and can achieve sub-second cold starts in many cases

    • Fly.io Machines use Firecracker and can boot in ~300–500ms

    • Modal also uses Firecracker to get very fast container startup for AI workloads

    Those platforms make it possible to keep the isolation benefits of VMs while achieving near-container startup speeds, which seems very aligned with what you’re targeting here.

    Trigger.dev’s programming model is excellent for workflows, so if the cold start drops into that <500ms range, it would make it much easier for us to keep more user-facing workflows inside Trigger.

    Would love to know how this is progressing and whether there’s anything users like us can do to test or provide feedback as the new infrastructure rolls out. 🚀

    • Miguel Fernández

      •

      Mar 9

      In spite of the AI trace from the message re-write, this is a genuine human asking for the ETA.

      • mattaitken

        Team•

        Mar 15

        We will ship an opt-in beta region this month

        • An Anonymous User

          •

          Apr 5

          Is that up to date?

        • An Anonymous User

          •

          Apr 17

          Hey Matt, any news on this? It’s been a month since your message.

        • Abdul kader Mousa Basha

          •

          May 15

          Hello matt,

          Any updates on this ?

          I hope this get implemented, this would honestly be very valuable.