Arena started in 2024 as a UC Berkeley Sky Computing Lab project for comparing models head to head. Today, it's grown into a platform for evaluating model performance across agents, text, code, images, and videos, running up to 600,000 E2B sandboxes a day. Each sandbox is an isolated cloud computer where an agent can write code, install dependencies, and work for hours on coding, research, reports, and presentations. That isolation keeps results trustworthy and secure: no session's code or files can reach another's and skew the comparison. That security at scale gets tested every time a frontier model drops. Before GPT-5 went public, people rushed to Code Arena to try it first. "When GPT-5 was about to come out, a lot of people came to Code Arena to experience the model firsthand because it wasn't out to the public yet. E2B was the backbone behind all of that. So when we had this massive surge, we didn't have to worry about whether we could handle it." - Aryan Vichare, Founding Engineer, Arena Read the full case study linked in the comments.
About us
Machines for AI agents
- Website
-
https://e2b.dev
External link for E2B
- Industry
- Technology, Information and Internet
- Company size
- 51-200 employees
- Headquarters
- San Francisco, California
- Type
- Privately Held
- Founded
- 2023
Employees at E2B
Locations
-
Primary
Get directions
Market St
San Francisco, California, US
-
Get directions
Na Perštýně 342/1
6th floor
Prague, 11000, CZ
Updates
-
Introducing E2B Embed: it brings AI agent sandboxes directly into your product and run them on infrastructure you control. E2B Embed packages the full E2B stack on a single node, giving you the flexibility to deploy in your own environment or a customer’s. Open source under Apache 2.0, it supports Docker Compose, Terraform on GCP, and Kubernetes. All with the same E2B SDK, CLI, and API. https://lnkd.in/gBvCAyaD
E2B Embed: a new way to run E2B
https://www.youtube.com/
-
Behind every great agent is a great stack. We’re bringing E2B, Fireworks AI, and Braintrust together in SF for a night on: → sandboxes → inference → observability → what it actually takes to run agents in production Lightning talks, 🍣 sushi, drinks, and builders. Come get STACKED (registration link in the comments)↓
-
-
Why make an agent choose one way of solving a problem when it can explore them all? Catch our very own Mish Ushakov and Ondrej Drapalik at AI Engineer Paris on Thursday, September 24, as they show how agents can fork live machines to explore multiple paths in parallel- each with its own copy of the machine’s state, filesystem, and running processes.
-
-
E2B reposted this
Testing AI-written code is one thing, testing it against real customer data raises the stakes. How do you give testing agents the access it needs while keeping that data isolated? Lark built an AI test platform that maps a customer's product, writes end-to-end tests, and keeps running them as the code changes. That means testing environments with real user data and credentials in play. Lark also needed to run Docker inside the sandbox itself to spin up a customer's own dev environment for testing. E2B's MicroVM isolation gave Lark strict security boundaries plus Docker-in-sandbox support, without having to build that isolation layer themselves. "They're testing their live environments with real customer data or real API keys, and so everything has to be securely isolated. We don't want an agent in one sandbox to be peeping into another sandbox. It's crucial to the security of our product." - Jack Brown, CEO, Lark Read the full case study linked in the comments.
-
We’re hiring a community builder to make E2B the best place on the internet for people building agents. Talk with builders on X. Grow our Discord + GitHub. Bring everyone together for workshops, dinners, and hackathons. Come build the agent community with us ↓ https://lnkd.in/g2Ni9s-w
-
E2B reposted this
Go from “we should build that” to a working agent faster with the Agents API. Now you can build on the same harness that powers Codex. We handle orchestration, long-running sessions, and context management, so your team can spend more time building something customers want. You define your agent’s tools and workflows, and choose where it runs code and works with files. Start with an OpenAI-hosted sandbox, bring your own environment, or connect a sandbox provider. Plus, we keep improving the harness as new model capabilities arrive. Less infrastructure upkeep on a roadmap that’s already full. Available in public beta: https://lnkd.in/e79vVXGM
-