computer provides isolated desktop boxes that agents can use to complete and drive real tasks. Each computer runs a Linux desktop inside a container or microVM, where an agent can open web pages, drive native apps, take screenshots, move the pointer, type text, run commands, and transfer files. You can also watch the agent work or take control through a browser when it needs help.
- Quick start
- Choose an interface
- Choose how to control the desktop
- Desktop operations
- Accessibility operations
- Browser operations
- Display and window operations
- Human control
- Configure a desktop
- Box lifecycle and history
- Server, REST, and MCP
- Runtimes
- Custom desktops and runtimes
- What is inside the box
- Security
- Examples
- Workspace crates
- Development and testing
Install Rust 1.85 or newer and one supported runtime:
- Docker
- Podman
- nerdctl
- microsandbox, for a microVM instead of a container
There's no need to fetch or manage a separate desktop image. The image is built automatically from source the first time you launch a desktop. The initial build takes a few minutes, but later desktops with the same configuration start in seconds.
curl -fsSL https://raw.githubusercontent.com/CITGuru/computer/main/scripts/install.sh | shBuild and install computer and computerd from the current source:
git clone https://github.com/CITGuru/computer.git
cd computer
cargo install --path . --lockedcomputer controls desktops from a shell and serves MCP over stdio. computerd keeps the REST and HTTP MCP service running.
Start and control a desktop:
BOX=$(computer new)
computer open "$BOX" https://example.com
computer screenshot "$BOX" screen.png
computer rm "$BOX"See the CLI guide for server and remote-fleet use.
Add the API without the command dependencies:
[dependencies]
computer = { git = "https://github.com/CITGuru/computer", default-features = false }
tokio = { version = "1", features = ["macros", "rt-multi-thread"] }Launch a desktop, open a page, send input, and take a screenshot:
use computer::{Button, Computer, Point};
#[tokio::main]
async fn main() -> computer::Result<()> {
let computer = Computer::launch().await?;
computer.open_url("https://example.com").await?;
computer.click(Point::new(640, 81), Button::Left).await?;
computer.type_text("driven from rust").await?;
let _png = computer.screenshot().await?;
computer.shutdown().await
}CLI:
BOX=$(computer new --url https://example.com)
computer mouse "$BOX" click 640 81 left
computer keyboard "$BOX" type "driven from the CLI"
computer screenshot "$BOX" screen.png
computer rm "$BOX"Computer::launch() starts one 1280x800 screen. viewer_url() returns a browser URL where you can watch it. shutdown() stops and removes the container.
From this repository, run the complete example:
cargo run --example quickstartThe computer command starts an ephemeral local server when no daemon is available. Start computerd when you need persistent traces, forks, remote access, or a fleet that survives command exits.
computerd
computer new --url https://example.com
computer lsSet COMPUTER_SERVER_URL and, when required, COMPUTER_SERVER_TOKEN to use a remote server. Add --local to bypass the server for the smaller set of direct local commands.
See the CLI guide.
The root computer package exports the desktop API. Disable its default cli feature when an application needs only the library:
computer = { git = "https://github.com/CITGuru/computer", default-features = false }Optional features add the E2B, Vercel Sandbox, Daytona, and Modal clients, microsandbox library binding, and daemon storage backends: e2b, vercel, daytona, modal, microsandbox, sqlite, postgres, and s3.
computerd exposes the complete remote API under /v1. The API supports box creation and lifecycle, action batches, frames, pages, windows, files, commands, takeovers, CDP access, traces, and forks.
Use computer-client from Rust or use HTTP directly. The wire types are in computer-api.
Use either interface:
computer mcp --stdiofor a local stdio MCP serverhttp://<server>/mcpfor Streamable HTTP served bycomputerd
Both interfaces expose the same boxes and tools. Hosts that support MCP Apps can show the live desktop beside tool results.
The crate provides three control methods. You can use more than one method in the same task.
Screen coordinates work with any visible application. Take a screenshot, select a point, and send mouse or keyboard input. Use this method when the application exposes no structured interface.
Accessibility nodes expose native widgets by role and name. Use them for file dialogs, settings panels, installers, and other native interfaces where coordinates can become stale.
Chrome DevTools controls Chromium pages directly. Use it to navigate, inspect the DOM, find elements, fill forms, evaluate JavaScript, and manage isolated browser sessions.
Coordinates use device pixels. The top-left corner is (0, 0).
let png = computer.screenshot().await?;
computer.move_to((640, 400)).await?;
computer.click((640, 400), Button::Left).await?;
computer.double_click((640, 400), Button::Left).await?;
computer.drag((100, 100), (400, 300), Button::Left).await?;
computer.type_text("hello").await?;
computer.press("ctrl+shift+p").await?;
computer.scroll((640, 400), Delta::down(3)).await?;
computer.scroll((640, 400), Delta::right(3)).await?;
let pointer = computer.cursor().await?;CLI:
computer screenshot "$BOX" frame.png
computer mouse "$BOX" move 640 400
computer mouse "$BOX" click 640 400 left
computer mouse "$BOX" click 640 400 left --double
computer mouse "$BOX" drag 100 100 400 300 left
computer keyboard "$BOX" type "hello"
computer keyboard "$BOX" press "ctrl+shift+p"
computer mouse "$BOX" scroll 640 400 down 3
computer mouse "$BOX" scroll 640 400 right 3
computer mouse "$BOX" atComplete Keyboard and Mouse reference:
computer keyboard <box> type <text> [--delay MS]
computer keyboard <box> press <key>... [--held shift,ctrl,alt,super]
computer keyboard <box> down <key> [--hold SECONDS]
computer keyboard <box> up <key>
computer mouse <box> move <x> <y> [--smooth|--human] [--seed N]
computer mouse <box> click <x> <y> [left|right|middle] [--double]
[--held shift,ctrl,alt,super]
[--smooth|--human] [--seed N]
computer mouse <box> drag <x1> <y1> <x2> <y2> [left|right|middle]
[--held shift,ctrl,alt,super]
[--smooth|--human] [--seed N]
computer mouse <box> path <x1> <y1> <x2> <y2> ... [left|right|middle]
[--held shift,ctrl,alt,super]
[--smooth|--human] [--seed N]
computer mouse <box> down [<x> <y>] [left|right|middle] [--hold SECONDS]
[--smooth|--human] [--seed N]
computer mouse <box> up [<x> <y>] [left|right|middle]
computer mouse <box> scroll [<x> <y>] up|down|left|right [NOTCHES]
computer mouse <box> scroll <x> <y> <DY> [DX]
computer mouse <box> at
The left button is the default. --held applies to a single click or a drag, not to --double. A scroll without coordinates uses the centre of the screen. In the signed form, positive DY moves down and positive DX moves right. down and up need the server and are not available with --local.
scroll moves the content below the given point. Positive dy moves down and positive dx moves right. A delta with both values sends one diagonal gesture.
Common key names work as expected. The crate converts names such as enter, cmd, and pageup to the display server's key names.
Keep these coordinate rules:
- Screenshots do not show the pointer. Use
cursor()to read its position. - Do not calculate a click from a scaled screenshot.
- Take a new screenshot after a person or another process changes the desktop.
- Take a new screenshot after an operation opens or raises a browser tab.
Hold modifiers as part of the pointer operation:
let screen = computer.primary();
let at = Point::new(640, 400);
let from = Point::new(100, 100);
let to = Point::new(400, 300);
screen.click_with(at, Button::Left, &[Held::Shift]).await?;
screen.drag_with(from, to, Button::Left, &[Held::Ctrl]).await?;CLI:
computer mouse "$BOX" click 640 400 left --held shift
computer mouse "$BOX" drag 100 100 400 300 left --held ctrlModifier pointer operations are available on X11 and on Wayland.
Move the pointer with a smooth line or a repeatable human-like curve:
computer mouse "$BOX" move 640 400 --smooth
computer mouse "$BOX" click 640 400 left --human --seed 42
computer mouse "$BOX" drag 100 100 400 300 left --human --seed 42The Rust API exposes the generated steps through motion::path and sends a full path in one operation:
use computer::{Desktop as _, Motion, Point, motion};
let from = Point::new(100, 100);
let to = Point::new(400, 300);
let steps = motion::path(from, to, Motion::Human, 42);
computer.move_along(&steps).await?;
computer
.drag_along(from, &steps, Button::Left, &[Held::Shift])
.await?;Instant motion stays the default.
Use separate button operations only when a normal drag cannot express the interaction:
computer mouse "$BOX" down 400 400 left --hold 10
computer mouse "$BOX" move 900 600 --smooth
computer mouse "$BOX" up 900 600 leftA key can stay down between commands in the same way. Everything typed or clicked while it is down carries it:
computer keyboard "$BOX" down shift --hold 20
computer mouse "$BOX" click 300 200
computer mouse "$BOX" click 300 400
computer keyboard "$BOX" up shiftdown takes one key, not a combination. The server lets a button or a key go after --hold (10 seconds unless told, 60 at most), when a person takes the screen over, and before the box is paused.
The Rust API exposes the same low-level operations:
use computer::Desktop as _;
computer
.button_down(Some(Point::new(400, 400)), Button::Left)
.await?;
computer.move_to(Point::new(900, 600)).await?;
computer
.button_up(Some(Point::new(900, 600)), Button::Left)
.await?;
computer.let_go(Button::Left).await?;The server holds a CLI down for 10 seconds by default and 60 seconds at most. It releases the button when that limit ends, when a person takes control, or at the end of an action batch.
Pace text for applications that drop characters:
computer
.type_text("typed at a visible pace")
.every(Duration::from_millis(20))
.await?;
computer
.press(["tab", "tab", "tab"])
.holding([Held::Alt])
.await?;computer keyboard "$BOX" type "typed at a visible pace" --delay 20
computer keyboard "$BOX" press tab tab tab --held altcursor() reads the pointer without moving it. find_cursor() may move the pointer on a display server where probing is the only way to find it:
let last_known = computer.cursor().await?;
let measured = computer.find_cursor().await?;Wait for drawing to stop instead of using a fixed sleep:
computer
.wait_until_still(
Duration::from_millis(400),
Duration::from_secs(10),
)
.await?;CLI:
computer wait "$BOX" --settle 400 --within 10000An animated screen can reach the deadline without becoming still.
screenshot() returns a full-size PNG. capture() can select a region or window and can reduce its size:
let screen = computer.primary();
let full = screen.screenshot().await?;
let region = screen
.capture(&Shot::region(Rect::new(
Point::new(100, 80),
400,
300,
)))
.await?;
let window_id = screen
.windows()
.await?
.into_iter()
.next()
.ok_or_else(|| computer::Error::denied("no window is open"))?
.id;
let window = screen.capture(&Shot::window(&window_id)).await?;
let small = screen
.capture(&Shot::window(&window_id).scaled(50))
.await?;CLI:
computer screenshot "$BOX" full.png
computer screenshot "$BOX" region.png --at 100,80 --size 400x300
computer screenshot "$BOX" window.png --window 42
computer screenshot "$BOX" small.png --window 42 --scale 50
computer screenshot "$BOX" pointer.png --pointerA window capture looks up the current window position when it runs. Scaling reduces the bytes sent to an agent, but scaled captures must not be used to calculate input coordinates.
Run commands and move files between the host and the desktop:
use computer::Search;
let output = computer.exec(["ls", "-la", "/tmp"]).await?;
let input = std::fs::read("input.png")?;
computer.write_file("/tmp/input.png", &input).await?;
let bytes = computer.read_file("/tmp/input.png").await?;
computer.upload("input.pdf", "/tmp/input.pdf").await?;
computer.download("/tmp/output.gif", "output.gif").await?;
let entries = computer.list_dir("/tmp").await?;
let matches = computer
.grep(&Search {
pattern: "needle".to_string(),
path: "/workspace".to_string(),
include: Some("*.rs".to_string()),
ignore_case: false,
limit: None,
})
.await?;
let paths = computer.glob("*.png", "/tmp", None).await?;
let logs = computer.logs().await?;CLI:
computer exec "$BOX" -- ls -la /tmp
computer file "$BOX" put ./input.png /tmp/input.png
computer file "$BOX" get /tmp/output.gif ./output.gif
computer file "$BOX" ls /tmp
computer file "$BOX" grep "needle" /workspace --include "*.rs"
computer file "$BOX" glob "*.png" /tmpSearch results are capped in the box before they cross the wire. Commands in the direct Rust API have a two-minute default limit. Use exec_within(argv, duration) when a command needs a different limit.
The CLI file commands use the server API and are not available with --local.
Each screen has its own clipboard and primary selection:
computer.set_clipboard("ready to paste").await?;
let text = computer.clipboard().await?;
computer
.set_selection(Selection::Primary, "middle-click paste")
.await?;
let selected = computer.selection(Selection::Primary).await?;CLI:
computer clip "$BOX" "ready to paste"
computer clip "$BOX"
computer clip "$BOX" "middle-click paste" --primary
computer clip "$BOX" --primaryCLIPBOARD is used by copy and paste. PRIMARY is filled when text is selected and is pasted by a middle click.
Selections can also contain binary data:
let targets = computer
.clipboard_targets(Selection::Clipboard)
.await?;
let png = computer
.clipboard_bytes(Selection::Clipboard, "image/png")
.await?;
computer
.set_clipboard_bytes(Selection::Clipboard, "image/png", &png)
.await?;While a person controls the screen, the program can read selections but cannot change them.
Accessibility operations use the semantic tree that native applications publish through AT-SPI. The built-in image has it by default:
let computer = Computer::builder().launch().await?;
let screen = computer.primary();CLI:
BOX=$(computer new)--minimal or Computer::builder().minimal() leaves it out.
Read the tree or search it by widget name and role:
let tree = screen.nodes(None, Some(4)).await?;
let street = NodeQuery {
query: "Street".to_string(),
role: Some("text".to_string()),
..NodeQuery::default()
};
let fields = screen.find_nodes(&street, Some(10)).await?;CLI:
computer widget "$BOX" tree --depth 4
computer widget "$BOX" find "Street" --role text --limit 10find_nodes returns the best match first. A query can match a widget's own name or the label beside it. The node.labelled field shows when a nearby label produced the match.
screen.focus_node(&street).await?;
screen.set_node(&street, "12 Bishop Street").await?;
let ok = NodeQuery {
query: "OK".to_string(),
..NodeQuery::default()
};
screen.invoke_node(&ok, None).await?;CLI:
computer widget "$BOX" focus "Street" --role text
computer widget "$BOX" fill "Street" "12 Bishop Street" --role text
computer widget "$BOX" press "OK"invoke_node runs a widget action without moving the pointer. The toolkit controls the action names. GTK can call an action click while Qt calls it Press.
Each node can also contain its centre point. Use that point when the application must receive a real pointer click:
if let Some(at) = fields.first().and_then(|node| node.at) {
screen.click(at, Button::Left).await?;
}CLI:
computer mouse "$BOX" click 640 400 leftApplications and custom widgets can publish incomplete trees. Use screenshots and normal input when a useful node is not available.
Browser operations use Chrome DevTools Protocol (CDP) to control Chromium by page structure instead of desktop pixels. They keep working when the Chromium window moves or another desktop window covers it.
Use three levels of browser control:
readreturns the rendered document as Markdown, text, or HTML.snapshot,find, and element actions inspect and operate on controls by query or reference.evalandPage::callare escape hatches for behavior that the higher-level API does not wrap.
Prefer page operations for websites. Use screenshots and desktop input for browser chrome, permission prompts, native file dialogs, and anything outside the page. Browser operations require a reachable CDP endpoint.
Open a page, inspect its controls, act by reference or name, wait for the result, and read only what changed:
BOX=$(computer new --url https://www.selenium.dev/selenium/web/web-form.html)
computer browser "$BOX" snapshot --urls --quiet 400
# @e2 textbox "Text input"
# @e8 combobox "Dropdown (select)"
# @e12 checkbox "Default checkbox"
# @e15 button "Submit"
computer browser "$BOX" fill @e2 "agent"
computer browser "$BOX" select @e8 "Two"
computer browser "$BOX" check @e12
computer browser "$BOX" click @e15
computer browser "$BOX" wait "Received!" --within 10000
computer browser "$BOX" snapshot --deltaA query can be visible text, an accessible name, an element ID, a CSS selector, or an @eN reference from snapshot. Use a reference after you inspect a page because it identifies one specific element. Use text or a selector when no snapshot exists. Add --exact when a partial text match could select the wrong control.
Use the operation that matches the control:
filltypes into text fields and assigns values to controls such as dates, colours, and sliders.checkanduncheckset checkbox state without toggling an already-correct value.options,select, anddeselectoperate on dropdowns.uploadassigns files directly to a file input; it does not open a native chooser.focus,hover,click, anddragsend the corresponding page interaction.
fill, focus, dropdown, and upload operations follow an explicit HTML label to its control, whether the label uses for or wraps the control. They do not guess that arbitrary nearby text names a field. If a query matches non-control content, the refusal names fields inside or beside it by reference or selector; when too many fields are nearby, it asks for a snapshot.
Do not use fill as a substitute for those specialized operations. It refuses a dropdown, checkbox, file input, or button and reports the correct operation.
Chrome DevTools controls a page without screen coordinates:
let browser = computer
.browser()
.expect("a published DevTools port");
let mut page = browser
.open_page("https://example.com", Duration::from_secs(20))
.await?;
let title = page.title().await?;
let links = page
.evaluate("Array.from(document.links).map(a => a.href)")
.await?;
let png = page.screenshot().await?;
page.navigate("https://example.org").await?;CLI:
TAB=$(computer open "$BOX" https://example.com)
computer browser "$BOX" eval "document.title" --tab "$TAB"
computer browser "$BOX" eval "Array.from(document.links).map(a => a.href)" --tab "$TAB"
computer browser "$BOX" screenshot page.png --tab "$TAB"
computer open "$BOX" https://example.org --target currentPage::call sends any Chrome DevTools Protocol method that the crate does not wrap.
Use open_page rather than open followed by a load wait. A new tab first shows about:blank, which is already loaded.
Pages provide higher-level operations for common browser tasks:
let fields = page.find("input", Some(10), None, None).await?;
page.fill("Email", "agent@example.com").await?;
page.choose(
"Country",
&["United Kingdom".to_string()],
false,
)
.await?;
page.upload("Attachment", &["/tmp/report.pdf".to_string()])
.await?;
page.click_on("Submit", Button::Left).await?;CLI:
computer browser "$BOX" find input --limit 10
computer browser "$BOX" fill "Email" "agent@example.com"
computer browser "$BOX" select "Country" "United Kingdom"
computer browser "$BOX" upload "Attachment" /tmp/report.pdf
computer browser "$BOX" click "Submit"These operations use page structure. They do not depend on the position of the Chromium window.
The browser API also has operations for focus, checkboxes, radio buttons, dropdown options, hover, drag, history, scrolling, and waits. A query can be visible text, an accessible name, an ID, a selector, or an element reference from a snapshot.
Read rendered page content as Markdown, text, or raw HTML:
let read = page
.read(Reading::Markdown, Some(4_000), Some(20))
.await?;
println!("{}\n{}", read.title, read.text);computer browser "$BOX" read --limit 4000A screenshot says where pixels are. A page read returns text beyond the viewport and the destination behind each link.
Snapshot the interactive controls in document order and address them by stable references:
let taken = page.snapshot(None, None).await?;
for element in &taken.elements {
println!("{}", element.text);
}
page.click_on("@e4", Button::Left).await?;
let changed = page.snapshot_delta(None, None).await?;computer browser "$BOX" snapshot --urls
computer browser "$BOX" click @e4
computer browser "$BOX" snapshot --deltaReferences stay with an element while the document stays loaded. A navigation clears them. A removed element leaves a hole instead of letting an old reference select a different control.
After a page has been snapshotted, server and CLI element actions also report the controls they made appear, change, or leave. This often removes the need for a second full snapshot.
Wait for content, its removal, an enabled control, a quiet page, a loaded document, or a JavaScript condition:
page.wait_for("Done", false, Duration::from_secs(10))
.await?;
page.quiet(
Duration::from_millis(400),
Duration::from_secs(10),
)
.await?;computer browser "$BOX" wait "Done" --within 10000
computer browser "$BOX" wait ".spinner" --gone
computer browser "$BOX" wait "Submit" --enabled
computer browser "$BOX" wait --load
computer browser "$BOX" wait --quiet 400
computer browser "$BOX" wait "Success" --or "Payment failed,Try again"
computer browser "$BOX" wait --fn "location.pathname === '/done'"Use browser waits after an action starts a fetch. --or waits for the first of several outcomes. --fn waits until a JavaScript expression is truthy. Use screen stillness for desktop drawing that browser state cannot describe.
A page screenshot excludes the desktop window, address bar, and pointer. It can capture the complete scrollable document:
computer browser "$BOX" screenshot page.png
computer browser "$BOX" screenshot page.jpg --full --format jpeg --quality 70open_url() opens and raises a new tab. Coordinates from an earlier screenshot then address the new frontmost page.
Use page visibility methods before you mix DevTools operations with screen coordinates:
let mut page = browser
.open_page("https://example.com", Duration::from_secs(20))
.await?;
computer.open_url("https://example.net").await?;
assert!(!page.visible().await?);
page.bring_to_front().await?;
assert!(page.visible().await?);CLI:
TAB=$(computer open "$BOX" https://example.com)
computer open "$BOX" https://example.net
computer browser "$BOX" tabs
computer browser "$BOX" switch "$TAB"
computer screenshot "$BOX" current-page.png --tab "$TAB"browser.visible_page().await? returns the page that the screen currently shows.
Mint a short-lived CDP address through computerd:
export CDP=$(computer cdp "$BOX")
agent-browser --cdp "$(computer cdp "$BOX" --ws)" snapshot -iPlaywright can pass $CDP to connectOverCDP, and browser-use can use it as cdp_url. The address contains a token in its path. It lasts one hour by default, can be changed with --ttl, and expires with the box.
computer cdp "$BOX" --direct prints the unguarded loopback port instead. Use it only on the box host. The proxied address is the correct form for remote servers.
computer cdp uses the server API and is not available with --local.
A browser group is a Chromium browser context. Groups share the Chromium process but keep cookies, local storage, IndexedDB, and service workers separate:
let group = browser.create_group().await?;
let mut page = group
.open_page("https://example.com", Duration::from_secs(20))
.await?;
page.evaluate("localStorage.setItem('agent', 'one')")
.await?;
group.close().await?;Groups do not create screens. Only one page can be frontmost. Call bring_to_front() before you use screen coordinates.
Export the parts of a browser session that belong to named origins:
let origins = ["https://example.com".to_string()];
let session = browser
.export_session(&origins, Carry::default())
.await?;
let tabs = browser.import_session(&session).await?;Carry::default() includes cookies and local storage. Enable IndexedDB for sites that store login state there. Session data grants account access and must be handled as a credential.
From the CLI, save to a file or leave it with the server under a name:
computer browser "$BOX" state save login.json --origin https://mail.example.com
computer browser "$OTHER" state load login.json
computer browser "$BOX" state save --name work
computer browser "$OTHER" state load --name work
computer browser state list
computer browser state rm workWith no --origin, the origins of the tabs open now are saved. The file is written readable by its owner only. A named state stays with the server until it restarts; the MCP tools save_state and load_state use names only, so the login never passes through the model.
Read, set and clear single cookies:
computer browser "$BOX" cookies --url https://example.com
computer browser "$BOX" cookies set theme=dark --url https://example.com
computer browser "$BOX" cookies set --curl "$(pbpaste)" # a browser's "copy as cURL"
computer browser "$BOX" cookies clear --url https://example.comThese are all commands under computer browser:
computer browser <box> read [--format markdown|text|raw] [--limit N] [--tab ID]
computer browser <box> snapshot [--scope QUERY] [--limit N] [--urls] [--delta]
[--quiet MS] [--tab ID]
computer browser <box> find [QUERY] [--role ROLE] [--exact] [--scroll]
[--limit N] [--tab ID]
computer browser <box> click <query> [--double] [--button left|right|middle]
[--new-tab] [--smooth|--human] [--seed N] [--tab ID]
computer browser <box> drag <from-query> <to-query> [--button left|right|middle]
[--smooth|--human] [--seed N] [--tab ID]
computer browser <box> hover <query> [--smooth|--human] [--seed N] [--tab ID]
computer browser <box> highlight <query> [--for SECONDS] [--tab ID]
computer browser <box> focus <query> [--tab ID]
computer browser <box> fill <query> <value> [--tab ID]
computer browser <box> check <query> [--tab ID]
computer browser <box> uncheck <query> [--tab ID]
computer browser <box> options <query> [--tab ID]
computer browser <box> select <query> <option>... [--tab ID]
computer browser <box> deselect <query> [<option>...] [--tab ID]
computer browser <box> upload <query> <file>... [--in-box] [--tab ID]
computer browser <box> wait [QUERY] [--gone] [--or TEXT,TEXT] [--within MS]
[--quiet MS] [--enabled] [--load] [--fn JS]
[--exact] [--tab ID]
computer browser <box> back [--tab ID]
computer browser <box> forward [--tab ID]
computer browser <box> reload [--tab ID]
computer browser <box> eval <expression> [--timeout MS] [--limit N] [--tab ID]
computer browser <box> console [--errors] [--clear] [--limit N] [--tab ID]
computer browser <box> errors [--clear] [--limit N] [--tab ID]
computer browser <box> screenshot [FILE] [--full] [--format png|jpeg]
[--quality N] [--annotate] [--tab ID]
computer browser <box> pdf [FILE] [--landscape] [--no-background] [--tab ID]
computer browser <box> dialog accept [TEXT] | dismiss
computer browser <box> tabs
computer browser <box> switch <tab>
computer browser <box> close <tab>
computer browser <box> state save <file> | --name NAME [--origin URL]...
[--session-storage] [--indexed-db]
[--no-local-storage]
computer browser <box> state load <file> | --name NAME
computer browser state list
computer browser state rm <name>
computer browser <box> cookies [--url URL]
computer browser <box> cookies set NAME=VALUE... --url URL [--domain D]
[--path P] [--secure] [--http-only]
[--expires SECONDS]
computer browser <box> cookies set --curl '<curl command>'
computer browser <box> cookies clear --url URL | --all
--tab takes a tab id or the label open --label gave it. read, snapshot, find, eval and console take --content-boundaries, which puts what the page wrote between two markers that hold a nonce the page cannot know; COMPUTER_CONTENT_BOUNDARIES=1 does the same for every call and for the MCP tools.
A query reaches into a frame of the same origin as the page, and a snapshot numbers the controls in it. A frame from another origin is left out.
A page dialog stops every page tool. An alert is accepted by itself, and the action that opened it says what it said. A confirm, a prompt or a leave-page dialog makes that action fail at once with its words, and any other page tool say the page is not answering; dialog accept or dialog dismiss answers it with a key on the screen.
Opening a page and exporting a CDP endpoint are top-level commands:
computer open <box> <url> [--target blank|current] [--label NAME]
computer cdp <box> [--ws] [--ttl MINUTES] [--direct]
window list reports each window's ID, geometry, class, and title. Use the ID for later commands. Prefer the class when waiting for a window because a title often changes with the open document.
let screen = computer.primary();
let windows = screen.windows().await?;
let window = screen
.active_window()
.await?
.or_else(|| windows.into_iter().next())
.ok_or_else(|| computer::Error::denied("no window is open"))?;
screen
.arrange(
&window.id,
Arrange::Size {
width: 800,
height: 600,
},
)
.await?;
screen
.arrange(&window.id, Arrange::At(Point::new(120, 90)))
.await?;The Rust API also exposes focus, window state, and close operations:
screen.focus(&window.id).await?;
screen.arrange(&window.id, Arrange::Maximise).await?;
screen.arrange(&window.id, Arrange::Minimise).await?;
screen.arrange(&window.id, Arrange::Restore).await?;
screen.close_window(&window.id).await?;CLI:
computer window "$BOX" list
computer window "$BOX" active
computer window "$BOX" 42 size 800 600
computer window "$BOX" 42 move 120 90
computer window "$BOX" 42 focus
computer window "$BOX" 42 max
computer window "$BOX" 42 min
computer window "$BOX" 42 restore
computer window "$BOX" 42 closeMatch a window by class when possible. A title can change with the open document.
Wait for a new window to appear and stop moving:
let dialog = screen
.wait_for_window("Mousepad", Duration::from_secs(10))
.await?;CLI:
computer window "$BOX" wait Mousepad --within 10The result reports the final geometry. A window manager can clamp or reject the requested size or position.
window wait --within uses seconds. computer wait --within and browser wait --within use milliseconds.
Complete CLI window reference:
computer window <box> list
computer window <box> active
computer window <box> wait <class> [--within SECONDS]
computer window <box> <id> focus
computer window <box> <id> close
computer window <box> <id> move <x> <y>
computer window <box> <id> size <width> <height>
computer window <box> <id> max
computer window <box> <id> min
computer window <box> <id> restore
focus raises a window and then reports the window that actually has keyboard focus. close asks the application to close the window. Arrange commands return the final geometry because the window manager can clamp a move, enforce size hints, or refuse a state change.
Screen IDs start at zero. Screen 0 starts with the desktop. Other screens start when they are first requested:
let second = computer.screen(ScreenId(1)).await?;
second.open_url("https://example.org").await?;The bundled image supports up to eight screens. Each screen has separate browser data, clipboard selections, and view and control servers.
screen() leases the screen to the current process. A second caller is refused instead of receiving the same screen.
let bytes = std::fs::read("background.png")?;
computer.set_wallpaper(&bytes).await?;
let second = computer.screen(ScreenId(1)).await?;
second.set_wallpaper(&bytes).await?;PNG and JPEG data are supported. A profile without wallpaper support returns Unsupported.
Add video packages to record MP4 inside the box:
let computer = Computer::builder()
.packages(Extras::video().packages)
.launch()
.await?;
computer
.record(Duration::from_secs(10), "/tmp/screen.mp4")
.await?;
computer
.download("/tmp/screen.mp4", "screen.mp4")
.await?;The duration-based Rust helper uses X11 capture. Use the profile-aware CLI recording commands for X11 or Wayland.
CLI:
BOX=$(computer new)
computer record "$BOX" start --fps 20
computer record "$BOX" status
computer record "$BOX" stop screen.mp4Add Extras::audio() when the recording also needs sound.
The recording example shows how to build an animated GIF from selected screenshots.
hand_over() gives a person exclusive input through a browser:
let takeover = computer.hand_over().await?;
let control_url = takeover.url();
let frame = computer.screenshot().await?;
assert!(
computer
.click(Point::new(640, 400), Button::Left)
.await
.is_err()
);
takeover.end().await?;CLI:
computer takeover "$BOX"
computer release "$BOX"The program can continue to read the screen while the person controls it. Its input operations return an error.
Use share() only when the person and the program must send input at the same time:
let shared = computer.share().await?;Shared input can race. Use exclusive handover when both sides do not need simultaneous control.
let takeover = computer.hand_over().await?;
computer
.wait_until_free(Duration::from_secs(600))
.await?;
takeover.end().await?;
let new_frame = computer.screenshot().await?;Always take a new screenshot after a handover.
An attached process can end a takeover that outlived its owner:
let computer = Computer::attach("my-box").await?;
if computer.person_driving().await {
computer.reclaim().await?;
}CLI:
computer release "$BOX"The image also refuses raw synthetic input from inside the box while a person has exclusive control.
computer new converts command flags into a portable Spec and Placement, starts the box, and prints its ID to standard output. Status text and the viewer URL go to standard error, so command substitution receives only the ID:
BOX=$(computer new --url https://example.com)
computer box "$BOX"Choose the display and capacity:
BOX=$(computer new \
--size 1920x1080 \
--screens 2 \
--wayland \
--memory 4g \
--cpus 2 \
--runtime podman)Build applications and features into the image:
BOX=$(computer new \
--app gimp,vscode \
--package jq \
--package ripgrep \
--audio)
MINIMAL_BOX=$(computer new --minimal)A box is one of two desktops:
--base, the default: CJK and emoji fonts, recording support (ffmpeg), the desktop launcher, native widget operations (accessibility), and Xwayland on a Wayland box.--minimal: none of them.
--audio adds the sound server to either.
Applications, packages, and features are image inputs. They cannot be added to a running box, and the first box with a new combination must build an image.
Set network and lifetime policy:
BOX=$(computer new --no-network --ttl 60 --idle 10)--ttl removes the box after a fixed number of minutes. --idle removes it after that many minutes without server activity. --no-network blocks outbound network access from the desktop.
Complete command reference:
computer new [--size WIDTHxHEIGHT] [--screens N] [--wayland] [--url URL]
[--app NAME]... [--package PACKAGE]... [--audio]
[--base | --minimal]
[--no-network] [--memory SIZE] [--cpus N]
[--runtime NAME] [--ttl MINUTES] [--idle MINUTES]
[--spec FILE|-]
Repeat --app and --package, or give comma-separated names. --spec - reads a create request from standard input. Flags override values from the file. computer --local new --name NAME can choose a local runtime name; a server always assigns its own box ID and refuses --name.
Use the builder to change the defaults:
let computer = Computer::builder()
.size(1920, 1080)
.network(false)
.memory("2g")
.runtime("podman")
.name("my-box")
.keep_on_drop(true)
.launch()
.await?;CLI:
BOX=$(computer new --size 1920x1080 --no-network --memory 2g --runtime podman --ttl 60)computer new keeps the desktop after the command exits. --app gimp and --package jq install into the image, and --spec box.json takes a file shaped like the body of POST /v1/boxes for what has no flag; a flag goes over the file.
network(false)blocks outbound network access from the desktop.runtime()also acceptsnerdctl.keep_on_drop(true)leaves the desktop running when the handle is dropped.expires_after(duration)removes the desktop after a fixed time.expires_when_idle(duration)removes it after a period without activity through that handle. Calltouch()when work reaches the box by another path.
Computer::builder().config()? returns the resolved image, ports, environment, and boot command without starting a desktop.
A Spec describes the desktop, installed applications, and access policy. A Placement describes where it runs and its resource and lifetime limits:
{
"spec": {
"desktop": {
"server": "x11",
"width": 1280,
"height": 800,
"packages": ["jq"]
},
"policy": {
"network": true
}
},
"placement": {
"runtime": "docker",
"memory": "2g",
"expires_after_secs": 3600
}
}BOX=$(computer new --spec box.json)The same specification can be placed in a container, microVM, or supported cloud sandbox. Unknown keys are refused instead of ignored. See examples/box.json and examples/from_spec.rs.
List the built-in application catalog, install applications into a new image, and open one by name:
computer apps
BOX=$(computer new --app gimp --app vscode)
computer app "$BOX" gimpapp waits until the application's window has drawn. Applications and packages are image inputs. They cannot be added to a running box.
The default image has CJK and emoji fonts, ffmpeg, the dock, and accessibility. Add Debian packages or sound when required, or leave the defaults out:
let desktop_with_packages = Computer::builder()
.packages(["vim", "curl"])
.audio()
.launch()
.await?;
let minimal_desktop = Computer::builder().minimal().launch().await?;CLI:
BOX=$(computer new --package vim,curl --audio)Extras::audio(), Extras::video(), Extras::accessibility(), and Extras::everything() provide common package sets.
Each builder package helper sets the complete extra package list. Use one combined list when you need custom packages and a preset together. The package list is part of the image tag. The first launch with a new list builds a new image.
Keep the complete Chromium profile in a named container volume:
let computer = Computer::builder()
.profiles("agent-work")
.launch()
.await?;The profile includes logins, history, extensions, and browser storage. It stays on the host and can be used by only one desktop at a time.
From the CLI or MCP, name the profile when the box is made:
computer new --profile workThe volume is computer-profile-work. A second box asking for a profile that a box holds, running or stopped, is refused and names that box. Before a box with a profile is removed or stopped, its browser is closed cleanly, so a cookie set a moment earlier is written. A fork does not take the profile. Session cookies end with the browser, as they do on any computer; a site's "remember me" cookie is the one that carries a login over. E2B, Vercel, Daytona, Modal, and microsandbox boxes refuse a profile, since they have no Docker volume.
Use session export when the data must move between hosts. Use a named profile when all browser state must stay on one host.
let computer = Computer::attach("my-box").await?;CLI:
computer ls
computer screenshot "$BOX" screen.pngCLI commands attach to the desktop named by their <box> argument.
The attached desktop keeps its windows, browser profile, and files. Dropping an attached handle does not remove a desktop that the handle did not create.
List all boxes known to the selected server:
computer lsEach row contains the box ID, screen size, screen count, and a non-ready state when applicable. Inspect one box in detail:
computer box "$BOX"computer box reports:
- ID and state:
Ready,Paused, orStopped - configured screen count and size
- the digest of the portable specification
- creation and expiry times
- viewer and direct DevTools URLs when available
The box ID is the value accepted by every command that takes <box>. Treat viewer and DevTools URLs as credentials when they contain access tokens.
computer box uses the server API and is not available with --local. Local computer ls asks the default Docker runtime what is still running.
computer pause "$BOX"
computer resume "$BOX"
computer stop "$BOX"
computer resume "$BOX"
computer rm "$BOX"A paused box keeps its memory and ports but uses no processor. A stopped box keeps its writable filesystem without keeping its memory. Resuming a stopped box starts a fresh desktop on new viewer ports. Removing a box deletes its files.
The Rust API exposes the same lifecycle:
computer.pause().await?;
computer.resume().await?;
computer.stop().await?;
let computer = computer.start(Duration::from_secs(90)).await?;Use --ttl MINUTES for a fixed lifetime and --idle MINUTES for an inactivity limit. computerd also sweeps expired boxes that outlive the process that created them.
A batch holds one screen across several operations, stops at the first refusal by default, and returns one final frame:
[
{ "type": "open_url", "url": "https://example.com/order" },
{
"type": "on_page",
"what": { "op": "fill", "query": "Name", "text": "Ada" }
},
{
"type": "on_page",
"what": { "op": "click", "query": "Continue" }
},
{
"type": "on_page",
"what": { "op": "wait_for", "query": "Details" }
}
]computer batch "$BOX" actions.json --settle 400Use --keep-going only when later steps do not depend on earlier steps. The REST equivalent is POST /v1/boxes/{id}/screens/{screen}/actions.
CLI batches, traces, forks, pause, stop, and resume use the server API. The ephemeral server can run a batch, but a long-lived computerd is required to retain useful history across commands.
computerd records actions, frames, commands, lifecycle changes, file transfers, and custody changes:
computer trace "$BOX"
NEW_BOX=$(computer fork "$BOX")A fork launches the same specification and replays the trace. It reconstructs the work; it does not copy a running machine. Page changes, timing, and unrecorded file or clipboard bytes can make the result differ.
The daemon listens on 127.0.0.1:8080 by default:
computerd
curl http://127.0.0.1:8080/v1/healthBind outside loopback only with a server token:
COMPUTER_SERVER_ADDR=0.0.0.0:8080 \
COMPUTER_SERVER_TOKEN="$(openssl rand -hex 32)" \
computerdClients send the token as Authorization: Bearer .... The gate protects REST and MCP. Viewer links use separate per-box credentials.
The REST API uses shared request and response types from computer-api. Box creation, action batches, and forks accept idempotency keys so a transport retry does not repeat a click or create a second box. See the server guide for routes and semantics.
By default every box the server creates lives an hour unless you change it from the confiuration. computerd offers one runtime per host runtime it finds — docker, podman and nerdctl for containers, smolvm for a microVM on libkrun — and GET /v1/runtimes says what each one runs a box in and what it can do. A placement names one; a box that names none lands on the default. COMPUTER_SERVER_CONFIG points at a file that tunes them, offers an engine twice, or turns one off, and COMPUTER_SERVER_SANDBOXES=e2b (or vercel) adds a remote vendor when the daemon was built with that vendor's support. A vendor can also be added while the server runs — computer runtime add cloud --provider e2b --api-key takes the key on stdin — and is then sealed in the store and offered again after a restart.
On restart, computerd takes back every box it recorded, through the runtime its record names, and then scans each runtime for boxes left labelled by an earlier server.
Create a box on a loopback server:
BASE=http://127.0.0.1:8080
BOX=$(
curl -fsS "$BASE/v1/boxes" \
-H 'content-type: application/json' \
-H 'idempotency-key: create-demo-1' \
-d '{
"spec": {
"desktop": { "width": 1280, "height": 800 },
"policy": { "network": true }
},
"placement": { "expires_after_secs": 3600 }
}' |
python3 -c 'import json,sys; print(json.load(sys.stdin)["id"])'
)Drive it with one action batch and one final frame:
curl -fsS "$BASE/v1/boxes/$BOX/screens/0/actions" \
-H 'content-type: application/json' \
-H 'idempotency-key: open-demo-1' \
-d '{
"actions": [
{ "type": "open_url", "url": "https://example.com" },
{ "type": "wait_still", "settle_ms": 400, "within_ms": 10000 }
],
"want": ["frame", "cursor"]
}'Inspect and remove it:
curl -fsS "$BASE/v1/boxes/$BOX"
curl -fsS -X DELETE "$BASE/v1/boxes/$BOX" \
-H 'x-computer-confirm-delete: true'Add Authorization: Bearer <token> to every request when COMPUTER_SERVER_TOKEN is set. API errors have one shape: code, message, and retryable.
The default in-memory store keeps traces only for the life of computerd. Select a durable backend when traces and frames must survive a restart:
COMPUTER_STORAGE_BACKEND=local \
COMPUTER_STATE_DIR=/var/lib/computer \
computerdSQLite, PostgreSQL, and S3 are optional build features:
cargo install --path . --locked --features sqlite,postgres,s3SQL backends use COMPUTER_STATE_URL. S3 uses COMPUTER_S3_ENDPOINT, COMPUTER_S3_BUCKET, credentials from the environment, and optional region and prefix settings.
By default, old frames are retained for two hours and trace entries for seven days. COMPUTER_KEEP_FRAMES_SECS, COMPUTER_KEEP_ENTRIES_SECS, and COMPUTER_PRUNE_SECS change those windows.
For an MCP host that launches a local process:
{
"mcpServers": {
"computer": {
"command": "computer",
"args": ["mcp", "--stdio"],
"env": {
"COMPUTER_SERVER_URL": "http://127.0.0.1:8080"
}
}
}
}Add COMPUTER_SERVER_TOKEN to the stdio environment when the daemon is gated.
For a remote host:
URL: https://boxes.example.com/mcp
Transport: Streamable HTTP
Authorization: Bearer <COMPUTER_SERVER_TOKEN>
Put TLS in front of computerd and forward WebSocket upgrades when an MCP Apps host must render the live screen. COMPUTER_PUBLIC_URL sets the public origin when a reverse proxy hides it.
An MCP Apps host can render ui://computer/screen.html beside launch_box, open_screen, and hand_over. The page receives a short-lived screen ticket under _meta; the model does not receive it.
See the MCP guide.
The desktop operations stay the same across runtimes. Startup, isolation, image storage, networking, and cleanup can differ.
Docker is the default. Podman and nerdctl use the same image and desktop API:
let computer = Computer::builder()
.runtime("podman")
.launch()
.await?;A container shares the host kernel. The runtime selects free host ports, and the desktop is removed when its owning handle is dropped unless configured otherwise.
A microVM boots its own kernel and gives a stronger isolation boundary. It starts more slowly and uses a separate image store.
The included integration targets microsandbox 0.6. Import the image into the hypervisor before the first launch. See examples/microvm.rs for the complete flow.
Other hypervisors can implement MicroVmApi.
The included E2B integration runs the desktop away from the local host. Build with the E2B HTTP client and set its API key:
computer = { git = "https://github.com/CITGuru/computer", default-features = false, features = ["e2b"] }export E2B_API_KEY=...E2B uses templates instead of local container images, and this crate builds one when it needs one: the first launch translates the bundled image into E2B's build steps, uploads the files it copies, waits for the build, and starts the box on it. A template it already has is used as it is, and one template is built per spec, so a box asking for GIMP gets its own.
To build one by hand instead — for a template you want to keep, or to see what the builder is given — write the context out and use their CLI:
python3 crates/computer-core/images/context.py \
crates/computer-core/images/desktop \
/tmp/e2b-ctx \
--for e2b
e2b template create computer-desktop \
-p /tmp/e2b-ctx \
-d Dockerfile \
-c "/usr/local/bin/computer-desktop" \
--ready-cmd "true" \
--cpu-count 2 \
--memory-mb 2048The transform removes instructions that E2B rejects or overrides and makes the desktop home writable by uid 1000. Pass the template ID printed by E2B to the example:
cargo run --features e2b --example e2b -- <template-id>Add --keep to leave the sandbox running until its deadline. The takeover example gives the public control viewer to a person:
cargo run --features e2b --example e2b_takeover -- <template-id> "search text"The viewer URL is withheld by default. A machine configured with public_viewer(true) is internet-reachable and must use Auth::Password or Auth::Token; launch is refused without that gate.
E2B sandboxes are created with public traffic off by default, so every port refuses a request without the traffic token that E2B issues for the sandbox. A browser cannot send that token, so such a box gives no direct viewer or takeover URL: watch it and take it over through computerd, which carries the token. public_traffic(true) on the machine, or --field public_traffic=true on computer runtime add, opens the ports, and the viewer and takeover URLs come back. Every port is then reachable by anyone with the sandbox ID: the viewer keeps its own token and the DevTools bridge its secret, but Chromium's own port 9222 is protected only by its refusal of any host name other than an IP address or localhost.
Page tools reach Chrome DevTools at the vendor's address for port 9223. A bridge in the box answers only requests that carry a secret made for that box. A template built from an older image does not have the bridge, so build it again.
The included Vercel Sandbox integration works the same way from the library. Build with the vercel feature and set the token, team, and project:
computer = { git = "https://github.com/CITGuru/computer", default-features = false, features = ["vercel"] }export VERCEL_TOKEN=... VERCEL_TEAM_ID=... VERCEL_PROJECT_ID=...
cargo run --features vercel --example vercelVercel starts sandboxes only from its own registry, for linux/amd64. With no image, the first launch builds the bundled image in a builder sandbox at Vercel, pushes it to vcr.vercel.com/<team-slug>/<project>/computer-desktop:<fingerprint>-x86_64, and removes the builder; later launches find the tag and start in seconds. An image directory is built the same way. A named image goes to Vercel as it is and must already be in the registry. Every published port has a public URL, so viewers need Auth::Token or Auth::Password, and page tools go through the same DevTools bridge with its secret. Vercel publishes at most 14 ports, so a box has at most 6 screens. See docs/guides/vercel-sandbox.md.
The included Daytona integration needs only DAYTONA_API_KEY and the daytona feature. Daytona builds from a Dockerfile but has no build context, so each COPY of a file in the bundled image or an image directory is sent as a RUN that writes the file from base64; Daytona keeps each build by the Dockerfile's content, so a later box starts in seconds. Each published port gets a signed preview URL that a browser opens with no header. See docs/guides/daytona.md.
The included Modal integration needs MODAL_TOKEN_ID, MODAL_TOKEN_SECRET, and the modal feature. Modal has no REST API, so the client speaks Modal's gRPC API with hand-written messages; Modal says that API is not public and can change. The image goes as Dockerfile lines with copied files inline, commands and files go through Modal's command router, and each published port is a public HTTPS tunnel. See docs/guides/modal.md.
X11 is the default display profile. Select the Wayland profile explicitly:
let computer = Computer::builder()
.profile(Arc::new(WaylandProfile))
.launch()
.await?;CLI:
BOX=$(computer new --wayland)The public desktop API is the same for both profiles.
The X11 image uses Xvfb, fluxbox, x11vnc, ImageMagick, and xdotool. The Wayland image uses headless sway, wayvnc, grim, and one virtual pointer and keyboard per screen that stay for the life of the screen, so a button or a modifier can be held between two steps as on X11.
Wayland cannot read the global pointer after a person moves it. cursor() returns Unsupported after a handover until the driver moves the pointer again.
Build a local Docker context:
let computer = Computer::builder()
.image_dir("images/my-desktop")
.launch()
.await?;Or pull an existing image:
let computer = Computer::builder()
.image("registry.example.com/desktop:1")
.launch()
.await?;A local context must contain a Dockerfile that implements the selected profile. A registry image cannot be combined with extra packages because the crate does not build that image.
An image can declare its contract:
LABEL computer.profile="computer-desktop"The crate refuses a declared profile mismatch before startup.
See examples/custom_image.rs and examples/images/acme/ for a small derived image.
Use ProfileBuilder when an image keeps most of a shipped contract:
let profile = ProfileBuilder::new(X11Profile)
.name("my-desktop")
.image_dir("images/mine")
.screen_commands(CommandScreen::new("my-screen"))
.wallpaper_runtime(CommandWallpaperRuntime::new("my-wallpaper"))
.build();
let computer = Computer::builder()
.profile(Arc::new(profile))
.launch()
.await?;A profile defines the image contract, ports, geometry, environment, capabilities, and display driver. A new display server implements Desktop and DesktopFactory.
RemoteApi is the common interface for E2B, Vercel, Daytona, Modal, and similar services. An adapter creates, finds, and removes a sandbox; runs commands; and reads and writes files.
RemoteMachine supplies shared lifetime, naming, and keep-alive behavior. RemoteProfile maps vendor endpoints and removes capabilities that the remote service cannot expose.
Run the local reference adapter before you connect an external API:
cargo run --example custom_sandboxSee examples/custom_sandbox.rs for a complete adapter and computer::testing::ScriptedRemote for tests without an account or network.
DesktopSupport states what a profile provides. audit tests those claims against a running desktop:
let audit = computer::audit(&computer).await;
assert!(audit.ok());The audit checks the screen, pointer, DevTools connection, clipboard, viewer, and handover where the profile claims support.
let removed = computer::sweep_expired(
&EngineMachine::default(),
SystemTime::now(),
)
.await?;CLI:
computer sweepcomputer sweep checks the local default Docker runtime, even when a remote server is configured. computerd reaps fleet boxes on its own cadence. Use deadlines for services that can stop before normal shutdown.
The default image is based on debian:bookworm-slim. It includes:
- Xvfb and fluxbox for the default virtual desktop
- Chromium with a separate profile for each screen
- x11vnc, websockify, and noVNC for browser viewing
xdotoolfor pointer and keyboard input- ImageMagick for PNG capture
xclipfor clipboard and primary selectionssocatfor the Chrome DevTools bridge- an input guard that blocks synthetic input during human control
The image tag contains a hash of its source files. A source change creates a new tag instead of reusing stale image contents.
The local viewer is open by default and is published only on loopback. Anyone who can reach a control port can drive the desktop.
computerd also refuses a non-loopback bind without a server token of at least 16 characters. /v1/health remains open for health checks.
Publishing outside loopback requires authentication:
let computer = Computer::builder()
.auth(Auth::Token)
.publish_on(Bind::Any)
.advertise("boxes.example.com")
.launch()
.await?;Auth::Password prompts in the browser and keeps the password out of the URL. Auth::Token puts a ticket in the URL so one link carries access.
Watch and control viewers use separate credentials. Changing the port on a watch URL does not create control access.
The box's raw Chrome DevTools port stays on loopback because CDP has no authentication. Use computer cdp to mint a short-lived address through the authenticated server. Treat that address as a credential.
network(false) blocks network access from the desktop. It does not protect the viewer.
A control port exists only during human handover. While it is open, the box rejects other synthetic input.
Browser sessions and persistent profiles contain authenticated account state. Treat them as credentials.
Start with these:
- quickstart launches, controls, and captures a desktop.
- browser controls Chromium through DevTools.
- elements finds and acts on page elements.
- capture captures regions and windows at different sizes.
- takeover gives control to a person and takes it back.
- tour controls two screens.
The repository also includes:
- serve leaves a named desktop running.
- attach connects to a running desktop.
- recording creates an animated GIF.
- waiting waits for browser state.
- research searches and reads web pages.
- from_spec launches from a portable specification.
- demo records a multi-step desktop animation.
- live_desktop audits the container image.
- custom_image builds and uses another image.
- microvm runs the desktop with microsandbox.
- custom_sandbox implements a sandbox vendor.
- e2b runs in an E2B sandbox.
- e2b_takeover gives an E2B desktop to a person.
- vercel runs in a Vercel sandbox.
- daytona runs in a Daytona sandbox.
- modal runs in a Modal sandbox.
- client drive exercises the REST client end to end.
Run an example with Cargo:
cargo run --example quickstart
cargo run -p computer-client --example drivecomputeris the root package. It re-exports the Rust desktop API and, by default, buildscomputerandcomputerd.computer-coreimplements boxes, desktops, profiles, display drivers, images, and runtime adapters.computer-typesholds portable specifications and shared values.computer-apidefines REST wire types.computer-clientis the Rust REST client.computer-serverimplementscomputerd.computer-cliimplements the commands.computer-mcpmaps MCP tools and the live screen app to REST.computer-storageprovides memory, local, SQLite, PostgreSQL, and S3 storage.
Run the same static checks as CI:
cargo build --workspace
cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace --no-fail-fast
RUSTDOCFLAGS="-D warnings" cargo doc --workspace --no-deps --all-features
cargo deny checkAlso check optional backends and code that is not compiled by the normal workspace commands:
cargo clippy -p computer --no-default-features -- -D warnings
cargo test -p computer-storage --features sqlite,postgres,s3
cargo test -p computer-server --features sqlite,s3
cargo clippy -p computer-core --features microsandbox --all-targets -- -D warnings
python3 scripts/check-page-scripts.py
crates/computer-mcp/ui/build.sh
git diff --exit-code -- crates/computer-mcp/ui/screen.html
for script in crates/computer-core/images/*/*.sh; do bash -n "$script"; doneThe normal test suite does not need a container runtime. Live tests are ignored by default:
cargo test --test live -- --ignored --nocapture
cargo test --test live_apps -- --ignored --nocapture
cargo test --test live_auth -- --ignored --nocapture
cargo test --test live_microvm -- --ignored --nocapture
cargo test --test live_extras -- --ignored --nocapture
cargo test --test live_session -- --ignored --nocapture
cargo test --test live_wayland -- --ignored --nocapture
cargo test --features e2b --test live_e2b -- --ignored --nocapturecomputer::testing provides test doubles for screens, container runtimes, hypervisors, and remote sandboxes.
MIT. See LICENSE.
