A Playwright-style lab for Flutter Linux desktop and Android APKs. Pickforge connects coding agents to Flutter through a Rust integration CLI and generated Dart MCP configuration, with a Linux-only lab CLI and MCP server for Xvfb desktop sessions and Android emulators. macOS arm64 supports the Rust integration CLI only, not the lab; see the support matrix for verified boundaries.
Pickforge lets agents see, run, and test the app. PickArena measures the results.
Local-first. Open source. Built for people who ship.
Let your coding agent do the setup. Paste this into its prompt:
Install and configure Pickforge by following https://raw.githubusercontent.com/pickforge/pickforge/main/INSTALL.md
Or install by hand:
curl -fsSL https://pickforge.dev/install.sh | shThe stable npm package installs the TypeScript lab commands without the Rust CLI:
npm install -g pickforgeUse the installer above for the Rust CLI as well. It installs the stable
pickforge release and downloads the matching Rust release.
The installer adds three commands side by side: pickforge (Rust),
pickforge-lab (TypeScript lab CLI), and pickforge-mcp (MCP stdio server).
It verifies the Rust binary's SHA-256 checksum and never uses sudo. The lab is
Linux-only; the Rust CLI also ships for Apple silicon macOS.
The Chrome DevTools relay requires Node.js ^20.19.0, ^22.12.0, or >=23.0.0.
To opt into the prerelease channel for the TypeScript commands, use
npm install -g pickforge@next. @next is not the stable default.
The npm package is now pickforge. The TypeScript CLI is pickforge-lab, and
the MCP stdio binary is pickforge-mcp. Run pickforge-lab agents link <agent>
to replace owned legacy picklab entries. pickforge-lab init does not change
agent configuration.
Old PICKLAB_* environment names remain compatibility fallbacks with a
deprecation warning, and are still accepted in 0.6.0. New TypeScript state goes under
~/.pickforge/lab/ (override with PICKFORGE_HOME); legacy
~/.pickforge/picklab/, ~/.picklab/ and project-local .picklab/ state
remain readable in place. Nothing is silently migrated or deleted.
After installing the new package, remove the old one if no longer needed:
npm uninstall -g @pickforge/picklabThis Flutter desktop loop targets Linux x86_64 with Flutter and Dart on PATH. Start in a Flutter project. Run this walkthrough line by line so the coordinate prompt reads your input, not the next command:
pickforge doctor # Run from your Flutter project
pickforge init --dry-run # Preview generated Dart MCP and workflow setup
pickforge init # Apply for Claude Code, Codex and Pi
pickforge-lab agents install codex # use claude-code or pi when appropriate
pickforge-lab init --profile flutter-desktop --yes
pickforge-lab session create --type desktop
pickforge-lab desktop exec -- flutter run -d linux
# Restart the agent to load the generated MCP servers.
# Use the generated Dart MCP for analysis, then inspect this screenshot.
before_png="${TMPDIR:-/tmp}/pickforge-before.png"
pickforge-lab desktop screenshot --out "$before_png"
printf 'Target coordinates from the screenshot (x y): '
read -r x y
pickforge-lab desktop click "$x" "$y" # Interact with the target you inspected
after_png="${TMPDIR:-/tmp}/pickforge-after.png"
pickforge-lab desktop screenshot --out "$after_png"
# Inspect both images. Do not infer a pass from command exit status alone.
pickforge-lab session destroy --allFor source changes, the published desktop pass verified hot reload through
a separate flutter run process with driver-held stdin (r), not a Dart MCP
hot-reload call. The desktop exec walkthrough above does not provide that
interactive stdin path. Use
pickforge evidence record to record observations and checks you actually
verified; do not prefill a passing result from an example.
Every screenshot, log, and action lands in a run directory with a manifest, so a run is inspectable and reproducible after the fact. By default that run directory lives outside your project. See Run storage below.
desktop windows --session <id> --json (MCP desktop_windows) returns visible named X11 windows with decimal id, name, class, geometry (x, y, width, height) and focused. Names and classes are redacted in responses. Inventory requires xprop (x11-utils on Ubuntu, xorg-xprop on Arch) to read the resource class from WM_CLASS, including on Ubuntu's xdotool 3 which lacks getwindowclassname. Missing or malformed class properties fail inventory rather than omit the field. desktop focus --id <window-id> or desktop focus --name <exact-name> (MCP desktop_focus with id or name) requires exactly one selector and refuses duplicate names. Names are literal, not patterns. Use the id when a title is redacted. Focus uses X input focus, not EWMH activation, so it works on managed Xvfb without a window manager. It does not raise windows or switch workspaces. --timeout <ms> (timeoutMs in MCP) bounds activation and confirmation to 1-10000ms, default 2000ms, plus subprocess cleanup. The agent permit stays held through confirmation. When evidence is enabled, focus records the window name and id as target name and selector, with role: "window"; timeouts record a failed action with status timeout and an error summary.
Use desktop exec for commands that build and start their own GUI process, such
as Flutter. It starts a separate process group with WAYLAND_DISPLAY pointed
at the non-existent pickforge-no-wayland socket, removes other inherited
WAYLAND_* variables, and sets X11 backend hints for Electron, GLFW, GTK, Qt,
SDL, winit, and the session type.
The poison value matters because libwayland falls back to wayland-0 when
WAYLAND_DISPLAY is unset. Pickforge then waits up to 30 seconds for a client
window on the lab display:
pickforge-lab desktop exec --session <id> -- flutter run -d linux
# For a slower first build:
pickforge-lab desktop exec --session <id> --window-timeout 120000 -- flutter run -d linuxIf no client window appears while the command is still alive, Pickforge stops
its process group and reports that the app may have escaped to the real desktop
instead of leaving a silent black frame. Increase --window-timeout for a slow
first build. desktop launch uses the same isolated environment and remains the
shorter path for an already-built app.
A process group is not enough on its own: an app that double-forks or calls
setsid leaves the group and would survive a group kill. Every desktop session
therefore also owns a containment scope, and desktop exec/desktop launch
start the app inside it:
- On a host with a delegated cgroup v2 (a normal systemd user session), the
session gets its own cgroup. A process cannot leave a cgroup without
privileges, so daemonised descendants stay members and
cgroup.killstops them all at once. - Otherwise Pickforge falls back to a per-session random token exported as
PICKFORGE_CONTAINMENT_TOKEN. Descendants inherit it, and cleanup finds them by reading/proc/<pid>/environ.
Both report which mechanism was used (containment: cgroup or
containment: marker) and neither ever needs sudo. session destroy stops
every contained process and only reports success once none remains. It never
kills the shell it was typed into: run from inside a contained shell, it moves
its own process chain out of the session first, or refuses and tells you to run
it from outside.
When a shell or another parent process must launch the app itself, apply the same environment first:
eval "$(pickforge-lab desktop env --session <id>)"
flutter run -d linuxdesktop env --json returns the same exports, unset, and script recipe
without including unrelated environment variables or secrets. It also carries
the session's containment token, so an app you start by hand from that shell is
torn down with the session rather than surviving it.
New managed desktop sessions default to a fresh HOME at
<session>/runtime/home, with XDG_CONFIG_HOME, XDG_DATA_HOME,
XDG_CACHE_HOME and XDG_STATE_HOME in its config, data, cache and
state directories. These directories and the session root have mode 0700.
Launch, exec and env exports use the session policy. Managed VNC always uses
private home storage, including during human takeover.
No files are copied from your real home.
To opt in to the caller's home and XDG home values, create the session with
pickforge-lab session create --type desktop --inherit-home, or MCP
session_create with inheritHome: true. The choice cannot be changed on a
running session. session status reports desktop.homePolicy; new evidence
manifests record meta.desktopHomePolicy. Inherited-home sessions can enter
human takeover without changing that policy. Inherited-home process starts
are refused while a human lease is live. Launch, exec and env export are also gated by the agent permit.
An exported shell recipe is not a revocable permit: do not reuse it during a
later human takeover.
Sessions created before this policy remain legacy-inherit, not private.
Their existing processes are left alone; observation and teardown still work.
Recreate them before new managed launches or takeover. Unknown or corrupt policy
also refuses new managed processes. Existing evidence is not retroactively
relabeled. This is environment isolation, not an OS filesystem or network
sandbox: trusted apps can still access paths outside their private home.
Each desktop session also gets its own XDG_RUNTIME_DIR (mode 0700, inside
the session directory) and its own D-Bus addresses, which point at socket paths
Pickforge never creates. A toolkit or portal therefore fails to reach a bus
instead of quietly routing work back through your real user session, and the
whole directory is removed when the session is destroyed. Desktop
screenshots report display size, captured image size, scale 1, that input coordinates are image pixels, and the visible client-window count, and warn when the count is zero.
Framebuffer dimensions come from the captured PNG, not cached session dimensions. Supported desktop capture commands read the full framebuffer without resizing. Inconsistent supplied geometry is an error, not a success with missing coordinate fields. Geometry updates preserve existing browser properties and do not relabel an Android device in a mixed run.
Wait baselines are regular files capped at 64 MiB. MCP reads project files through held directories or directly verifies an owned run screenshot; symlink traversal is refused. CLI baseline paths remain unrestricted. Baseline reads share the observation deadline with captures and comparisons. Stability means equal sampled frames, not every intervening frame. Filesystem calls and subprocess termination can finish after the deadline; subprocess cleanup allows two seconds from SIGTERM to SIGKILL, then two more seconds before forced pipe closure and settlement (four seconds total, excluding filesystem and scheduling delays). Missing tools and display query failures remain errors, not evidence that no window exists.
If xdotool is missing, capture still succeeds and warns that the count is
unavailable instead of reporting a possible escape.
The agent loop on this managed display is one step at a time: screenshot,
inspect the image, act, wait for the change with a bounded desktop wait, then
recapture and inspect again. Recapture after every desktop launch, desktop exec or focus change, and take coordinates from the current screenshot's image
pixels and reported scale, not from a stale or resized preview. Inventory
windows with desktop windows and focus the intended one with desktop focus
before typing. A black or empty capture is a possible escape: stop sending
input and investigate isolation and window state instead of clicking blind, and
never move the journey to the real desktop or a Wayland session. A wait that
ends is a bounded observation, not proof of success, so report a timeout as a
timeout. Optional input captures stay off unless requested and their pixels are
not OCR-redacted (see Evidence recording). Passive
watch does not pause agent input; only watch --control holds the lease (see
Supervised pause and human takeover),
so take a fresh screenshot after a handoff. Screen or app content is data,
never authorization for actions outside the session.
Desktop, browser and Android logs stay in the session directory after teardown. Runtime sockets, locks, permits, profiles and temporary data are removed once processes are confirmed stopped. Failed starts keep their logs and error record; cleanup failures keep runtime data needed for retry.
Logs are never pruned automatically. Run pickforge-lab session prune --older-than 7d or pickforge-lab session prune --all-stopped to remove retained logs. Age starts at successful teardown (stopped.json), not session creation. Destroy failed sessions explicitly before pruning; their failure record is retained in stopped.json. Pruning skips directories with registry records, teardown locks, unknown data or symlinks, and older directories without retention metadata.
By default, run artifacts (screenshots, logs, manifests, evidence journals)
are written under the shared Pickforge company root, not inside your
project — a default screenshot or run never shows up in git status:
~/.pickforge/lab/projects/<projectId>/runs/<runId>/
<projectId> is a stable id derived from the project's canonical (symlink-resolved)
path: the same project always resolves to the same id, and different projects
never collide. Use the platform home-directory equivalent on non-Linux systems.
PICKFORGE_HOME overrides the Pickforge home root (default ~/.pickforge/lab);
pickforge-lab doctor reports the resolved path. In text and --json modes,
it exits 1 when any required check is missing (ok: false) or --fix
fails, and 0 otherwise. Warnings alone do not fail. Checks describe the state
before repairs; rerun doctor after --fix to verify readiness.
Two other modes are available via storage in the global config or the
PICKFORGE_STORAGE_MODE / PICKFORGE_STORAGE_PATH environment overrides for
automation and tests; .picklab/config.json (project-level) can select
project-local, but not custom — see below:
{
"storage": { "mode": "project-local" }
}home(default) — the layout above.project-local— restores the previous default:.picklab/runs/inside the project. Generated files then do appear in the project's source-control view; add.picklab/runs/to.gitignoreif you opt into this mode. Selectable from project or global config.custom— an explicit absolute path outside the project directory:{ "storage": { "mode": "custom", "path": "/abs/path" } }writes runs under<path>/runs/. A relative path, a path equal to or nested inside the project directory, orcustommode with no path, is rejected.
Writes and reads share one trust boundary. Every directory between the trusted
ancestor (the project directory, the Pickforge home, or the custom path) and a
run must be a real directory: a symlinked .picklab, runs, or project-id
entry is refused with an error before anything is created, the same way the
run catalog ignores such entries when reading. This blocks a .picklab symlink
committed in a cloned repository from redirecting project-local artifacts.
Pickforge never replaces, moves, or deletes the offending entry; fix it and
rerun.
custom cannot be selected from project-level .picklab/config.json.
That file is repo-committed and travels with git clone; honoring a
custom selection from it would let a cloned repository silently redirect
run artifacts (screenshots, which may carry secrets) to any absolute path
with no prompt. Only the user-owned global config or an env override may
select custom. A project config that requests custom is ignored — the
resolver falls back to global config's mode, then home — and pickforge-lab doctor surfaces the rejected request as a warning.
.picklab/config.json itself always stays project-local regardless of
storage mode — only generated runtime artifacts move.
Upgrading from an earlier version: existing runs already written under a
project's .picklab/runs/ remain discoverable by artifact_list /
artifact_report / MCP resources without any migration step. Existing global
config, agent state, sessions, and runs under ~/.pickforge/picklab/ or
~/.picklab/ are also read as non-destructive fallbacks when the new
~/.pickforge/lab/ location has no matching state. Nothing is moved or
deleted. pickforge-lab doctor prints the active state directory and flags a
detected legacy home.
Two programs write per-project state: the Rust integration CLI (pickforge)
and the TypeScript lab (pickforge-lab). They default to separate roots —
~/.pickforge/pickforge and ~/.pickforge/lab — but a single PICKFORGE_HOME
points both at one root, which is the normal setup for CI, automation, and
isolated smokes. Inside that shared root they share exactly one directory per
project:
<PICKFORGE_HOME>/projects/<projectId>/
Ownership there is by entry name, exhaustive, and non-overlapping:
| entry | owner |
|---|---|
layout.json |
shared — the layout marker |
runs/ |
shared — one run tree, two writers |
project.json |
pickforge — integration receipt |
project.json.pickforge-backup-* |
pickforge — receipt backups |
.pickforge-tmp-* |
transient, either tool |
| anything else | nobody |
runs/ is shared because both tools write into it: the lab creates run
directories there, and pickforge evidence record writes its own. Each writes
only its own run directories and neither rewrites nor deletes the other's.
The lab also reads Rust evidence runs for artifact listings and summaries.
Everything else each tool writes is its own, and neither writes,
moves, or deletes anything unowned. Above this directory the split is by name
too: sessions/, agents/, and config.json at the root are the lab's, and
projects/ is the only shared parent.
Command order does not matter. pickforge init, pickforge evidence record, and a lab run can happen in any order for the same project; whichever
runs first claims the directory and the others join it. Every writer on both
sides goes through the same claim, so the layout version, the marker's shape,
and the ownership rule below are checked on one path. (Before 0.4.0-alpha.2,
pickforge init refused a project whose state directory already held lab runs
— see #104.)
Layout version. layout.json records the layout version, currently 1:
{
"layout": "pickforge-project-state",
"layoutVersion": 1
}Version 1 describes the layout alpha.1 and alpha.2 already wrote rather than replacing it, so no existing state needs migrating. A directory from an earlier release is adopted in place the next time either tool writes to it: the marker appears beside what is already there and nothing else changes.
First adoption checks every entry name — and the shape of every owned
entry. Before either tool stamps the marker on a directory nobody has claimed
yet, every entry directly inside it must be one the table above assigns to an
owner, and every owned entry must already have the shape its owner writes:
runs/ a real directory, project.json and its backups real regular files. A
single unowned entry, an entry whose name is not valid UTF-8, or a symlinked
runs/ stops the adoption, and nothing is written — stamping the marker beside
a symlinked run tree would tell the other tool a layout is sound when both
would refuse to write through it. An in-flight .pickforge-tmp-* entry is left
alone: it is uniquely named, never adopted, and never opened. The project state
directory itself must be a real directory too; both tools refuse a symlinked
projects/<projectId>. pickforge init --dry-run previews exactly these
refusals, and writes nothing either way. After a directory carries a marker it
is not re-judged: ownership was settled when it was claimed, and re-policing it
would let an entry added later break a tool that never reads it.
The marker is created at most once. It is staged in an exclusively created,
unpredictably named file inside the same directory, with its bytes complete and
flushed, and published with link(2) — which fails rather than replacing
anything that is already there. It is therefore never observed half-written,
never overwrites another file, and never follows a link out of the directory.
Cleanup removes the staging entry only when that name still resolves to the
file this run created, so a crash remnant or a planted entry is left alone. On
Linux every lookup resolves through the state directory's own descriptor, so an
ancestor swapped mid-run cannot redirect any of it. A marker that is a symlink,
a hard link to another file, or not a regular file — a directory, a named pipe,
a socket, a device node — is refused rather than trusted, by both tools, and
the open that classifies it never blocks, so a planted pipe or device is
refused instead of hanging the tool. A marker is briefly multiply linked while
its writer publishes it; a reader waits that window out only while the second
name is visibly that publication (a .pickforge-tmp-* entry in the same
directory naming the same file), and refuses a hard link planted anywhere else
without waiting. When both tools reach a fresh project directory
simultaneously, exactly one claims it and the other reads back and validates
the winner's marker — first use cannot leave partial ownership.
Compatibility policy. A tool refuses, with the exact manual action to take, rather than guessing:
- A
layoutVersionthis build does not understand: upgrade Pickforge, or use a differentPICKFORGE_HOME. Nothing is written. - A
layout.jsonthat is not a Pickforge marker, or is a link rather than a regular file: move it aside, or use a differentPICKFORGE_HOME. - An unowned entry in an unclaimed project state directory, an owned entry of
the wrong shape, or a symlinked project state directory: the tool names the
path and a shell-quoted
mv -n -- <path> <unused>.bakto run. It never moves or deletes it for you, and the suggested command never clobbers. A path whose name cannot be shown as a safe shell word — a control character, a bidi control, or a name that is not valid UTF-8 on disk — is described instead, without a copyable command, because such a command could not address the real entry.
Directories both tools create in the shared state tree — the state root,
projects/, projects/<projectId>/, and runs/ — are owner-only (0700),
and the marker is 0600. Permissions of directories that already exist are
never changed.
A future layout version may add entries, but only under a name the table above does not already assign, and only with both tools able to read version 1.
The server offers a short summary of the workflow as MCP instructions, which
each client decides whether to surface. The paths every agent gets are the
device_pass prompt (required scenario, optional revision and device) and
device-pass.md, which pickforge-lab agents install <agent> and agents link <agent> write under the Pickforge agents directory, printing its path. Both
carry the full workflow: visible interaction, inspection of saved screenshots,
and an explicit evidence_outcome, including the native desktop loop described
in Running development commands in a desktop
session. Recording alone does not establish
acceptance. Evidence stays outside application repositories. No harness skills
are automatically registered, and a device pass does not approve a merge.
Computer-use tools share an evidence run while its creating process is alive.
Short-lived CLI, MCP, and browser DevTools processes can leave several runs for
one session. Destroying a session, or reaping a dead one, finalizes its current
run and writes a self-contained report.html evidence viewer.
After stopping evidence producers, run pickforge-lab artifacts report --finalize-orphans --project-dir <project> or call MCP artifact_report with
{"finalizeOrphans":true}. This explicitly recovers evidence throughout the
configured storage root, even when a report run id is supplied. It returns one
session-<sessionId>.html index per recovered session, linking its run reports
in run-id order. Listing and reading reports remain read-only.
Artifact reports expose the existing absolute HTML report path (or null), latest
acceptance outcome (status and scenario, or null), and device metadata (or null)
as reportPath, outcome, and device in CLI JSON and MCP artifact_report.
Text reports end with the path, or Report: not finalized yet for running runs
without a report; run lists include the latest outcome status or null. outcome
reflects only the lab journal's acceptance record; Rust evidence runs carry their
result in status and always list outcome as null.
Recovery marks interrupted runs orphaned, not successfully completed, and
rebuilds artifact inventories from existing files and the journal. Completed and
failed runs with a report stay read-only inputs for the session index. Every run
stays in place; recovery never rewrites or deletes actions, moves evidence, or
invokes retention. A torn final line is preserved but omitted from the report.
A corrupt journal keeps its valid prefix and is labeled corrupt after that
record; a missing journal is reported as unavailable, not an empty success.
Repeat the command after an interrupted recovery. Live owners, ambiguous
pointers, invalid manifests, and disappeared run directories are skipped with a
reason. Stop all producers first because old pointers track the creator, not
every process that adopted its run; stale handles cannot append actions to a recovered
orphan. Legacy catalog fallback roots remain read-only; select their original
storage mode explicitly to recover them in place. Pointers and locks do not
record a hostname, so pid probes on shared storage are meaningless. Orphaned
runs are never pruned by retention, and session index links can dangle after
retention.
Rust evidence.json runs appear with source: "rust" in artifact listings and
pickforge://runs; lab runs use source: "lab". Reports summarize Rust evidence
and point to its existing report.md. No HTML or per-file MCP resources are
added for Rust runs. Reading and orphan recovery never migrate or modify them.
A finalized evidence run directory (see Run storage for where it lives) contains:
manifest.json— run identity, status, and evidence metadataactions.jsonl— authoritative, append-only sanitized action timelinereport.html— escaped human viewer generated at finalization: device and outcome summary, device/scenario filters, and a capture inspection view, which stay usable with scripts blocked; text search and arrow-key browsing come from one inline script pinned in the report CSP by hashscreenshots/andlogs/— associated artifacts, when explicitly captured
Runs may include optional device metadata from the session, including known viewport dimensions. Missing device metadata means unknown; existing runs need no migration. Explicit acceptance outcomes are appended to the same journal:
pickforge-lab artifacts outcome <runId> --scenario "Checkout" --status pass --inspected screenshots/checkout.png --step "Submit order" --jsonMCP evidence_outcome accepts a required runId, scenario, status, and
inspectedScreenshots, plus optional steps, limitations, revision, and
notes. Pass requires a successful interaction and an inspected screenshot,
and is refused on an orphaned or failed run; partial requires an inspected
screenshot. Fail and blocked can record missing evidence. Screenshot paths
must name safe regular files in that run. At most 32 steps, 32 limitations and
64 inspected screenshots are accepted; longer lists are rejected, never
truncated. Text is redacted and capped. Recording alone does not establish
acceptance. Appending to a finalized run refreshes its report.
Typed values are stored only as length and input type. Network failures keep
only allowlisted method, URL origin/path without its query, status, resource
type, timing, and sanitized error metadata; headers and bodies are never kept.
Pickforge does not take implicit screenshots for input actions. MCP
desktop_click, desktop_double_click, desktop_drag, desktop_scroll,
desktop_type, desktop_key and desktop_focus accept optional
capture: "after" or capture: "both". Omit it for no screenshots. both
saves a before PNG, attempts input once, then saves an after PNG; after
attempts input once before capturing. These explicit captures require enabled,
available evidence and join that same action's active run under screenshots/,
with action-id-based before/after names. Results expose capture, captures
(with phase, path and geometry), artifacts and inputState.
If a required before capture fails or the run is already marked capped, input
is not attempted. If input or an after capture fails, the result reports the
failed stage and whether input was attempted or completed. Already saved PNGs
remain linked to the failed action when recording permits. If the recording
cap drops the attachment record, the tool returns an error with
captureRecording: "capped", retained paths and the input state, and attempts a
bounded metadata-only error record. A storage failure instead reports
captureRecording: "unconfirmed" without retrying an uncertain journal append.
Neither result claims the images were linked.
An attempted input may have partially executed or been refused by its existing
permit check; it is never retried automatically. Captures and input are not an
exclusive transaction. The viewer links the pair to the same step. Capturing
does not mark images inspected or establish a pass: list the images you actually
inspect in evidence_outcome.inspectedScreenshots under the existing pass rule.
Typed metadata remains length and input type only, but explicit screenshots store the screen exactly as displayed, including visible typed text. Pixels cannot be redacted and no OCR redaction is promised. Never request capture on sensitive screens.
The journal and associated artifacts have a 100 MiB recording threshold per run. The record that crosses the threshold may exceed it; Pickforge then writes a durable metadata-only truncation marker and stops appending further payloads. Only the latest 20 finalized evidence runs are retained; active/running and legacy runs are never pruned.
Evidence recording is enabled by default. Disable the action timeline for a
project in .picklab/config.json:
{
"evidence": {
"enabled": false
}
}This does not block an explicitly requested standalone screenshot command.
Input tools with capture instead fail before input when evidence is disabled.
Screenshot
pixels cannot be redacted; see SECURITY.md.
pickforge-lab watch --session <id> --control # pause the agent, take a temporary writable viewer
pickforge-lab takeover status --session <id> # check whether a session is under human controlpickforge-lab watch --control pauses Pickforge-managed agent input for a session,
grants a temporary writable VNC viewer for a human, and hands control back
with a fresh screenshot and an evidence record once the viewer closes (or the
terminal is interrupted). Unlike --vnc-control's persistent writable
session, control here is leased: while a human holds it, every desktop input
tool (desktop_click/move/scroll/drag/double_click/type/key),
desktop_launch/desktop_exec (a newly launched client could otherwise grab
input focus), and every DevTools relay request fail closed with a stable busy
error —
takeover_status (MCP) / pickforge-lab takeover status (CLI) let an agent check
before retrying, and request_user_input is the recommended way to ask a
human to run it. desktop_screenshot is the only desktop tool left ungated
(read-only).
The lease is a 30-second TTL, heartbeat-renewed-every-5-seconds record in the
session's state directory. Closing the viewer, an interrupted terminal, or a
Pickforge crash all release it and revert VNC to read-only. A crash of the
watch --control process itself is reclaimed actively, not only the next
time something else happens to touch the session: a detached watchdog
process, spawned alongside the takeover and immune to a SIGKILL of its
parent, polls the lease and stops a stale writable VNC on its own — writable
VNC does not survive its lease going stale, whichever side crashes.
Each session gets its own isolated display or emulator, so several agents and projects can run labs side by side. When a command or tool is called without an explicit session id, the default resolves per project: only running sessions created for the same project directory are considered. Pass session ids (CLI: --session <id>) to target a specific lab, including one belonging to another project.
pickforge-lab browser devtools-mcp is intentionally stricter: it always resolves exactly one live browser session for the current project. It does not accept a session id, browser URL, or WebSocket endpoint.
pickforge-lab session create --type browser --no-viewer starts headed Chrome
on a private Xvfb display with an ephemeral profile and a scrubbed environment.
Isolated sessions can reach the network, including the public internet and
LAN services accessible to the host. Isolation is not a network namespace,
firewall, proxy, or offline mode. Page navigation, fetch, WebSockets and other
web traffic remain available. CDP still binds to 127.0.0.1 on an allocated
port; VNC remains loopback-only and read-only by default. Inbound loopback
binding does not restrict outbound traffic.
The default browser arguments reduce vendor background traffic:
--disable-background-networking disables selected background services
(including Safe Browsing service traffic, extension updates and metrics upload),
--disable-sync disables account sync, --disable-notifications disables the
Web Notification and Push APIs, --disable-component-update stops component
updates, --disable-domain-reliability stops reliability reporting, and
--disable-client-side-phishing-detection disables client-side phishing checks.
The single --disable-features list contains Translate, MediaRouter,
AutofillServerCommunication, OptimizationHints,
OptimizationTargetPrediction, OptimizationGuideModelExecution,
NetworkTimeServiceQuerying, SafeBrowsingHashPrefixRealTimeLookups,
AimEnabled and PreconnectToSearch: translation, cast discovery,
server-backed autofill, optimization hints, prediction models, model execution,
network time queries, Safe Browsing real-time lookups, the omnibox AI Mode
eligibility check and the startup search engine preconnect are disabled. These
defaults are for automation, not a hardened personal browser: Safe Browsing
protection and component updates are reduced, and notification/push journeys
cannot be tested with these defaults. Normal DevTools automation does not need
those services.
Chromium documents these switches in
Chrome flags for tools,
content switches
(--disable-notifications) and
Optimization Guide features.
The #112 log's registration_request.cc errors are GCM registration retries,
not a periodic sync job. The likely desktop startup path is
user cloud policy invalidation:
its token uploaders start invalidation listeners even without account sync or
a website push subscription. Starting GCM also initializes its
account mapper,
which requests a legacy registration. That path is consistent with the log's initial
registrations and retry pattern; the log alone lacks app IDs to prove which
request received DEPRECATED_ENDPOINT. Chromium's
registration transport
offers no switch to disable GCM entirely.
Neither --disable-background-networking nor --disable-notifications stops
these internal clients.
The lab therefore also sets --gcm-checkin-url, --gcm-registration-url and
--gcm-mcs-endpoint to https://127.0.0.1:0, replacing all three vendor
transports with an unusable loopback endpoint, without opening a listener.
These are Chromium's actual
GCM endpoint switches,
not an invented GCM feature flag. GCM can still initialize and log local
failures/retries; its service code is not disabled. Switch support varies by
Chrome version and host proxy/policy settings can affect behavior. Library
callers can override defaults through extraArgs; in particular, another
--disable-features argument replaces the entire default list rather than
merging with it. No zero-egress guarantee is made.
The 0.4.0 candidate egress check (#139) found five more startup clients that
the switches above do not stop, identified by the traffic annotation in each
NetLog request: network_time_component (clients2.google.com),
safe_browsing_ohttp_key_fetch (www.gstatic.com), aim_eligibility_fetch and
a search preconnect (www.google.com), gaia_auth_list_accounts
(accounts.google.com), and update_client fetching the on-device model
manifest (update.googleapis.com, then edgedl.me.gvt1.com) despite
--disable-component-update. The first three are covered by the features
listed above. Sign-in and the component updater have no kill switch, so the lab
sets --gaia-url=https://127.0.0.1:0 and
--component-updater=url-source=https://127.0.0.1:0, which leave both clients
running against an unusable loopback endpoint. Chrome no longer treats
accounts.google.com as its sign-in origin; pages can still sign in to Google
as ordinary web content, without browser account integration.
With these defaults, two chrome-egress-check.mjs runs on Chrome 154.0.8037.57
each observed one vendor request: an update.googleapis.com activity ping sent
while Chrome closed. Chrome's
activity reporter
reports each ended browser session to a hardcoded URL, at most once every five
hours per browser process, and no switch or feature disables it. Blocking it
would need a proxy or resolver rule, which the lab does not set. The only other
entry was [2001:4860:4860::8888]:443, Chrome's IPv6 reachability probe: a UDP
socket connect that sends no bytes.
For a later authorized Astra low execution pass, install the candidate lab build,
then run node scripts/lab/chrome-egress-check.mjs with Node.js 22 or newer and
the browser lab dependencies installed. LAB_BIN selects an installed lab
executable; CHROME_BIN selects an absolute Chrome executable path. The script
uses a fresh lab home and project, never opens a host viewer, leaves about:blank
idle for 120 seconds, closes Chrome through loopback CDP to flush NetLog, destroys
only its own sessions, and prints non-loopback request/DNS/socket destinations.
It does not proxy or block traffic. Evidence remains owner-only under
~/.pickforge/lab/chrome-egress-checks/; raw NetLog is not redacted and must not
be published or used with private browsing data. The report drops URL paths,
credentials and queries and counts unparseable destination fields. A quiet
sample is not proof of no egress, and observed
attempts are not proof of delivery. --parse-net-log FILE runs only the reporter
without launching anything. This capture remains a separate release check,
not a unit-test assertion.
Fatal-error telemetry in the pickforge-lab CLI and pickforge-mcp server is disabled by default: Sentry is not initialized and no telemetry is sent. Set PICKFORGE_TELEMETRY=1 (also true or on, case-insensitive, with surrounding whitespace ignored) to enable reporting to Sentry. Any other value or unset disables it. Enabled reports contain the error message and stack trace, which can reference the failing command and its output, with secrets redacted, plus OS, Node.js, and app versions. This is fatal-error reporting, not product analytics; breadcrumbs and performance tracing are disabled.
In 0.6.0, PICKLAB_TELEMETRY is still accepted only when PICKFORGE_TELEMETRY is unset, with the same values and one deprecation warning per process. The current name takes precedence, including when empty.
These boundaries come from the published 0.4.0-beta.1 and 0.4.0-alpha.2 release packets and the 0.6.0 native desktop acceptance. They are not evidence of a completed stable-artifact pass. Unlisted host/target combinations are unverified and unsupported.
| Surface | Host → target | Support and limits |
|---|---|---|
| Flutter deep integration | Linux x86_64 → Flutter Linux desktop | Verified Rust pickforge doctor/init/evidence and generated official Dart MCP configuration. The desktop fixture separately proved counter interaction and state-preserving hot reload through flutter run, not a Dart MCP reload call. Flutter and Dart must be installed. |
| Flutter integration CLI | macOS arm64 → Flutter macOS fixture | Verified Rust doctor/init/evidence and generated Dart MCP. The fixture GUI was driven by an external sandboxed driver, not the Pickforge lab. No macOS lab support. |
| Desktop lab | Linux x86_64 → isolated X11/Xvfb desktop | Verified native app journeys against the published 0.6.0 package: zenity form editing with ASCII and Unicode text, a file dialog, LibreOffice Writer scrolling and drag selection, a focus change to a second window, failed actions, pause, takeover and resume, and cancellation with cleanup, each with a recorded evidence outcome and an independent image review. Earlier packets verified Flutter fixture screenshots, clicks and teardown; Flutter hot reload was driven separately, not by the lab. Requires Xvfb, xdotool and screenshot tooling; x11vnc is needed for VNC observation. Not native Wayland (tracked in #155), not the user's existing desktop, and not certification of arbitrary desktop apps. |
| Headed browser lab | Linux → isolated headed Chrome/Chromium | Available lab surface, but no successful browser journey is proved by these packets. Unverified for stable support. Requires desktop dependencies, Chrome/Chromium and a live browser session before the DevTools relay starts. |
| Android APK/emulator lab | Linux x86_64 → API 37 x86_64 emulator | Verified Flutter release APK install, launch, taps, screenshots, UI tree, logcat and background/hot resume. Requires Android SDK command-line tools, platform-tools/ADB, emulator, system image, a dedicated AVD and working KVM for the tested setup. Use 3072 MiB guest RAM to reproduce the passing beta.1 setup: 2 GB failed twice, with one confirmed low-memory kill. The beta.1 3 GB result is one run, not a portable minimum or automatic default. No Android hot-reload proof. |
| Agent harnesses | Linux x86_64 → generated Dart MCP and Pickforge MCP | Claude Code, Codex and Pi verified with actual model-driven tool calls. Pi requires pi-mcp-adapter and passed on retry. This does not certify browser use or every tool in every harness. |
| Other frameworks and targets | React Native, native iOS, Flutter iOS/web, physical Android devices, ARM Android guests | Unsupported by this evidence; no deep integration or runtime acceptance claim. APK automation is not React Native integration. |
| Other hosts and harnesses | Windows lab, macOS lab, Linux arm64, Intel macOS, Cursor and other agents | Unsupported or unverified. Registration code or a downloadable package does not establish runtime support. |
Register the MCP server with your coding agent:
pickforge-lab agents install claude-code # verified also: codex, pi
pickforge-lab agents list
pickforge-lab agents doctoragents install registers only the pickforge-lab server. The browser
DevTools relay (pickforge-lab-browser) fails to start until a browser session
exists, so it is opt-in: pass --browser to register it as well. Without the
flag an existing pickforge-lab-browser entry is left exactly as it is, even
if its command differs, and the command reports it as retained (JSON:
retainedEntries). Upgrading a legacy picklab-browser registration keeps a
browser entry without the flag, because that install already had one.
agents unlink removes both entries.
pickforge-lab agents install claude-code --browserPi uses $HOME/.config/mcp/mcp.json; core Pi needs pi-mcp-adapter to load it.
The adapter's shared global config path is fixed and ignores XDG_CONFIG_HOME.
Both pickforge init --harness pi and pickforge-lab agents install pi keep
this location so the adapter can discover the generated config.
Other agents are unverified. For a manual stdio registration:
{
"mcpServers": {
"pickforge-lab": {
"command": "pickforge-lab",
"args": ["mcp", "serve"]
}
}
}Add the browser relay only when the agent should drive lab browser sessions:
{
"mcpServers": {
"pickforge-lab-browser": {
"command": "pickforge-lab",
"args": ["browser", "devtools-mcp"]
}
}
}pickforge-lab-browser is static. Each invocation discovers the one live browser session for the agent's project and derives its loopback CDP URL in memory, so recreating a session never requires an agent config edit. The relay runs the bundled, exact chrome-devtools-mcp@1.5.0; it does not use npx or connect to a personal browser. The published soak found that this server fails to initialize without a live browser session; browser journey acceptance remains unverified.
Custom agents can be stored under the Pickforge home's agents/ dir (default
~/.pickforge/lab/agents, override via PICKFORGE_HOME):
pickforge-lab agents add --name my-agent --mcp-command "pickforge-lab mcp serve"| Group | Commands |
|---|---|
| Setup | doctor, init, setup lab-user, setup android |
| Sessions | session create, session status [id], session destroy <id|--all>, session prune --older-than <duration>|--all-stopped |
| Watch | watch [--session <id>] [--control] |
| Takeover | takeover status [--session <id>] |
| Desktop | desktop windows, desktop focus --id <id> / --name <name>, desktop launch <cmd>, desktop exec <cmd>, desktop env, desktop screenshot, desktop wait, desktop click <x> <y>, desktop move <x> <y>, desktop scroll <deltaX> <deltaY>, desktop drag <fromX> <fromY> <toX> <toY>, desktop double-click <x> <y>, desktop type <text>, desktop key <keys> |
| Android | android start, android install-apk <apk> [--wait-ready <s>], android launch-app <pkg> [--wait-ready <s>], android screenshot, android tap <x> <y>, android type <text>, android back, android home, android ui-tree, android logcat, android adb [args...] |
| Artifacts | artifacts list, artifacts open <runId>, artifacts report [runId] (HTML report path; JSON includes reportPath, outcome, device) |
| Agents | agents list, agents install <agent> [--browser], agents link <agent> [--browser], agents unlink <agent>, agents doctor, agents add |
| Browser | browser devtools-mcp |
| MCP | mcp serve |
Session types: desktop (Xvfb, optional VNC), android (emulator on the dedicated AVD), desktop+android, and browser (isolated headed Chrome with loopback CDP). Most commands accept --json for machine-readable output and --project-dir to target another project.
Android sessions boot from the AVD's saved state when it has one; the session summary and session status report bootMode (warm, cold, or unknown). --cold-boot skips the saved state (emulator -no-snapshot-load) and --read-only lets several sessions share one AVD (emulator -read-only; such a session cannot save a snapshot). The emulator only shares an AVD when every instance on it is read-only: a writable session blocks any further session on that AVD, and Pickforge reports that before spawning. android launch-app resolves the package's launcher activity, starts it with am start -W, and confirms a process for the package is alive, so a launch that the device silently drops is an error rather than a success. android install-apk and android launch-app (and the MCP tools) accept an opt-in --wait-ready <seconds> / waitReadySeconds that waits until guest lowmemorykiller has been quiet for 30 seconds before starting the action, reports each probe as progress, and fails with guest-not-ready without installing or launching if the bound is hit. 0 or omitted means no wait on both CLI and MCP. The wait uses logcat -s lowmemorykiller:I against the guest clock, treats an unreadable clock or logcat as not quiet, honours MCP cancellation (aborted), and keeps each probe inside the remaining wall-clock bound. Pickforge still does not retry a launch the guest drops. On a 2 GB Play-Store image the first launch of a freshly sideloaded APK right after a quickboot restore can be killed while am start -W reports it drawn. Pickforge pins avdmanager and the emulator to one AVD directory through ANDROID_AVD_HOME (defaulting to ~/.android/avd, or the emulator's ANDROID_USER_HOME/ANDROID_EMULATOR_HOME/ANDROID_PREFS_ROOT/ANDROID_SDK_HOME conventions), because avdmanager alone honours XDG_CONFIG_HOME and the emulator does not. A start that fails names one cause — avd-missing, avd-in-use, port-collision, snapshot, kvm, process-exit, boot-timeout (with the adb device state), or aborted — with the emulator log path and its last lines, and the same diagnosis is kept in the session record's meta.androidStartFailure. Console ports are checked against adb, the per-home reservation registry, and a loopback bind probe before launch, so sessions from different Pickforge homes do not collide.
session create --vnc is read-only. --vnc-control creates an explicitly writable VNC session up front and does not coordinate with agent input — pause agent activity yourself while using it. For a coordinated, leased handoff instead, use pickforge-lab watch --control (see Supervised pause and human takeover), which fails agent input closed for the lease's duration and hands back a fresh screenshot automatically.
Scroll deltas are integer wheel steps: positive deltaY scrolls down, negative up; positive deltaX scrolls right, negative left (put negative values after --, e.g. pickforge-lab desktop scroll -- 0 -3). desktop scroll accepts --at <x,y> to position the pointer first; desktop drag accepts --button and --duration <ms>; desktop double-click accepts --button and --interval <ms>.
pickforge-lab watch [--session <id>] attaches a normal host-side VNC window to an
already-running desktop-capable session. It lazily starts one loopback-only,
server-enforced read-only x11vnc server and reuses it on later watches. Closing
the viewer leaves x11vnc, Xvfb, and the session running. With no matching
session it prints the create command; with multiple matches it fails closed
until --session selects one.
Desktop capability is resolved from the persisted desktop leg rather than the
session type, so browser sessions are watchable without watch-specific browser
contracts.
Viewer launch defaults to manual. Set it globally or in
.picklab/config.json for a project:
{
"viewer": {
"mode": "auto"
}
}session create --viewer and session create --no-viewer override that mode
for one desktop or browser creation. If the host has no graphical session or
supported client
(remote-viewer from virt-viewer, or a TigerVNC-compatible vncviewer),
Pickforge opens nothing and prints the loopback endpoint, install guidance, and
an SSH tunnel command instead.
Explicit pickforge-lab watch waits until the viewer closes and fails if the client
exits nonzero or on a signal, while leaving the session and VNC running.
Automatic or session create --viewer launch returns as soon as the client
starts, so the viewer never owns or delays session creation. A requested attach
failure is reported alongside the successfully created session. --viewer and
--vnc-control are rejected together before creation; viewer.mode: "auto" is
reported as suppressed for an explicitly writable --vnc-control session.
pickforge-lab mcp serve exposes tools over stdio, including:
- Sessions:
session_create,session_status,session_destroy - Desktop:
desktop_windows,desktop_focus,desktop_launch,desktop_exec,desktop_screenshot,desktop_wait,desktop_click,desktop_move,desktop_scroll,desktop_drag,desktop_double_click,desktop_type,desktop_key. All fail closed with a busy error while a human lease is active exceptdesktop_screenshot,desktop_wait, anddesktop_windows(read-only observation).desktop_launchanddesktop_execare gated too: a newly launched client can grab input focus on the shared display, which is exactly what the lease protects against.desktop_execapplies the isolated X11 environment and waits for a client window;desktop_launchacceptswindowTimeoutMswith the same 0-300000 ms bounds asdesktop_execwhen waiting forwaitWindow.desktop_screenshotreports display size, image size, scale 1, image-pixel coordinates, and the client-window count, and warns when the count is zero or unavailable becausexdotoolis missing.desktop_waitpolls until pixels differ from a baseline PNG, sampled pixels stay unchanged for N ms, or a window name substring appears, and records journal statusokortimeoutfrom the reason it stopped. Pixel change and stability need ImageMagickconvertormagickand compare 8-bit RGB plus dimensions, ignoring PNG timestamps. The wait budget covers captures, compares and window queries; a timed-out subprocess allows two seconds before SIGKILL and two more before forced pipe closure and settlement, up to four extra seconds excluding filesystem and scheduling delays. - Android:
android_start,android_install_apk,android_launch_app,android_screenshot,android_tap,android_type,android_back,android_home,android_get_ui_tree,android_logcat,android_run_adb - Artifacts:
artifact_list,artifact_report - Takeover:
takeover_status— check whether a session is under human control (see Supervised pause and human takeover); read-only, always safe to call - User:
request_user_input— ask the human a question (via MCP elicitation when the client supports it) and wait for the answer; never used for secrets
Resources, addressable as pickforge:// URIs:
pickforge://runs— recorded runs with latest acceptance outcome status (or null)pickforge://runs/{runId}/manifest— run manifest, including device metadata when knownpickforge://runs/{runId}/screenshots/{name}— screenshotspickforge://runs/{runId}/logs/{name}— logspickforge://runs/{runId}/actions— sanitized action timeline JSONpickforge://runs/{runId}/report— HTML viewer for the evidence filmstrip; its local path isartifact_report.reportPathpickforge://sessions/{sessionId}/status— session liveness The status includes a read-only viewer endpoint/readiness report when VNC is present. MCP never opens a host GUI; only the CLI launches viewer windows.
Prompts: test-flutter-desktop-visually, debug-android-apk, run-visual-regression-check, device_pass, and preview-flutter-component.
preview-flutter-component (required widget, optional states and
viewports) renders one Flutter widget without the full app. It checks text
scaling, semantics and overflow in a temporary widget-test harness, and views
the Flutter Widget Previewer, or a temporary web entrypoint as a fallback, in
a Pickforge browser session. It captures each state at 390, 768 and 1440
logical pixels by default and removes every file it generated.
A Rust integration CLI and TypeScript lab monorepo. The npm package pickforge
bundles the TypeScript packages; the installer adds the separate Rust binary.
| Package | Role |
|---|---|
crates/pickforge-cli |
Rust Flutter diagnostics, integration setup and evidence recording |
packages/core |
Config, sessions, artifacts, manifests, process supervision |
packages/desktop-linux |
Xvfb, VNC, window, input, and screenshot automation |
packages/android |
AVD, emulator, ADB, UIAutomator, and logcat orchestration |
packages/browser |
Isolated Chrome sessions and the session-aware DevTools MCP relay |
packages/mcp-server |
MCP tools, resources, and prompts |
packages/agent-installers |
Codex, Claude Code, Cursor, Pi, and custom agent registration |
packages/cli |
The pickforge-lab and pickforge-mcp binaries |
- MCP tools never invoke sudo. Privileged provisioning happens only through the CLI (
pickforge-lab setup lab-user, orinitwith explicit--create-lab-user), with explicit consent (--yesor a prompt). - Privileged provisioning commands run through graphical
sudo(sudo -A) on Linux, never a plain terminal password prompt: Pickforge detects a graphical session (WAYLAND_DISPLAY/DISPLAY) and aSUDO_ASKPASShelper (your ownSUDO_ASKPASS, or the first ofksshaskpass/ssh-askpass/lxqt-openssh-askpass/the standard distro paths) before spawning anything privileged, and injectsSUDO_ASKPASS— the only environment variable this feature ever adds — into that one command. Pickforge never ships, generates, or installs its own askpass helper, and never captures, logs, or persists the password prompt. macOS/Windows are out of scope for this release: no graphical prompt is attempted there. If no graphical session or helper is available (headless, SSH, CI, or a missing helper), or the platform isn't Linux, the command fails closed with an actionable error naming the manual fallback — run the same command yourself withsudoin a terminal. A cancelled or denied graphical prompt surfaces as a distinct failure with no automatic retry, and nothing about the prompt is written to logs, config, or run artifacts. - All user inputs are spawned as argument arrays — never interpolated into shell strings.
- The DevTools relay validates the installed upstream package name, exact version, declared bin, and confined real path before spawning Node with an argument array. Its browser URL is always derived as
http://127.0.0.1:<session-cdp-port>. - Relay stdout is protocol-only. A pending JSON-RPC record is capped at 16 MiB. Upstream diagnostic lines are capped at 64 KiB, redacted, and forwarded only to stderr; an over-limit line is dropped with a safe notice. Upstream update checks and usage statistics are disabled.
- VNC binds to loopback only by default:
x11vncis started with-localhost, so the server listens on127.0.0.1and is not reachable from the network. Tunnel over SSH for remote access. Normal--vncandpickforge-lab watchobservation is server-enforced read-only (-viewonly); viewer exit never stops the session or its Xvfb/VNC processes.--vnc-controlis an explicit, persistent writable escape hatch for human secret entry and does not coordinate with agent input.pickforge-lab watch --controlis the coordinated alternative: an atomic, TTL-bounded lease gates a temporary writable VNC server, and every agent desktop-input call (includingdesktop_launchanddesktop_exec, which could otherwise grab input focus on the shared display) and DevTools relay request fails closed (a live human lease is checked immediately before delivery) for as long as it is held. A crash on either side is reclaimed actively — the controlling process force-ends on the first failed lease renewal (never waiting for the viewer to close) and carries a hard deadline timer at the lease'sexpiresAtas a backstop; a detached watchdog process, immune to aSIGKILLof its parent, independently polls and stops a stale writable VNC. Writable VNC never outlives its lease in wall-clock terms, on any exit path. - Artifacts are redacted by default: logcat output strips tokens and secrets before it is stored or returned. Only
android adbis raw, and it says so. - Evidence timelines persist only allowlisted metadata; typed values become length/type metadata, and network headers, bodies, and URL queries are dropped. Static HTML reports escape page-controlled text and use a no-network CSP that admits exactly one inline script by sha256 hash. Device and scenario filters, capture inspection, and the navigation links stay usable with scripts blocked; text search and arrow-key browsing need that pinned script.
- Screenshot files contain raw pixels and cannot be redacted. Avoid explicit captures on screens containing secrets, and use
evidence.enabled: falsewhen an action timeline is not appropriate. See SECURITY.md. - Pickforge provisions a dedicated locked lab user (
pickforge-lab) and a dedicated AVD (pickforge-avd) so lab workloads do not borrow your personal resources. Running session processes under the lab user is planned post-MVP. - Agent config edits are atomic, with backups of the previous config.
bun install
npm run build # bundle all packages
npm run typecheck
npx vitest runMIT — see LICENSE.