Beta built-in tools for the Microsoft Agent Framework. A home for first-party
Python tools that plug into any chat client's shell / function surface. The
first tool is LocalShellTool.
pip install agent-framework-tools --preimport asyncio
from agent_framework import Agent
from agent_framework.openai import OpenAIChatClient
from agent_framework.tools import LocalShellTool
async def main() -> None:
client = OpenAIChatClient(model="gpt-5.4-nano")
async with LocalShellTool() as shell:
agent = Agent(
client=client,
instructions="You are a helpful assistant that can run shell commands.",
tools=[client.get_shell_tool(func=shell.as_function())],
)
result = await agent.run("Print the current working directory.")
print(result.text)
asyncio.run(main())- Persistent (default): a single long-lived shell session.
cd,export, and shell functions persist across tool invocations. - Stateless (
mode="stateless"): each command runs in a fresh subprocess.
LocalShellToolis not a sandbox. It runs commands directly on the host with the agent process's permissions. Human approval provides a check before agent-requested commands run, but does not isolate the shell. Production use needs separately enforced isolation and restricted permissions. For untrusted input use a sandboxed executor, such asDockerShellToolfor shell commands — see alsoagent-framework-hyperlight.
Defenses (in priority order):
- Approval-in-the-loop — when an agent calls the tool returned by
as_function()withapproval_mode="always_require"(the default), commands require approval through auser_input_requestbefore they run. Disabling this requiresacknowledge_unsafe=True. Direct calls torun()do not request approval, regardless of the configured approval mode. - Process-tree termination on timeout via
psutil, so child processes (make, watchers, network tools) cannot survive the timeout. - Output truncation to 64 KiB (head + tail with marker).
- Audit hook (
on_command=…) for SIEM / append-only logs. - Optional command-pattern filter via
ShellPolicy(denylist=[...], allowlist=[...]). Empty by default. This is a UX pre-filter, not a security boundary — operators are expected to supply patterns that match their workload (and they can be defeated by trivial obfuscation such as\rm -rf /,${RM:=rm} -rf /,python -c "…", encoded payloads, or PowerShell-native equivalents). Isolation must be enforced separately, for example by the sandbox tier (DockerShellTool); neither approval nor command filters isolate the shell. Seetests/test_security.pyfor the documented residual risk surface.
Command-text filtering with ShellPolicy and required human approval:
Warning
The filters below can allow embedded commands, including $(...) and
backticks, and do not enforce read-only access. Review the full command before
approving it. Approval does not isolate the shell: commands still run with
the application's permissions.
from agent_framework.tools import LocalShellTool, ShellPolicy
shell = LocalShellTool(
policy=ShellPolicy(allowlist=[r"^ls\b", r"^cat\b", r"^git status$"]),
approval_mode="always_require",
)Expose this tool to the agent through shell.as_function() and handle its
approval requests with human review. Calling shell.run() directly does not
request approval.
- Windows:
pwsh -NoProfile -Command -(falls back topowershell.exe). - Linux / macOS:
/bin/bash --noprofile --norc(falls back to/bin/sh). - Override via the
shell=constructor argument or theAGENT_FRAMEWORK_SHELLenvironment variable.
A model talking to a PowerShell session will sometimes default to bash
syntax (export FOO=bar, ls -la, > /dev/null) and vice versa.
ShellEnvironmentProvider is an AIContextProvider that probes the live
shell once per session — family, version, OS, working directory, and a
configurable list of CLI tools (git, node, python, docker by
default) — and injects a system-prompt block describing the shell idiom
to use and the available CLIs.
from agent_framework.tools import (
LocalShellTool,
ShellEnvironmentProvider,
ShellEnvironmentProviderOptions,
)
shell = LocalShellTool()
provider = ShellEnvironmentProvider(
shell,
ShellEnvironmentProviderOptions(probe_tools=("git", "uv", "node")),
)
agent = Agent(
client=client,
tools=[client.get_shell_tool(func=shell.as_function())],
context_providers=[provider],
)Probe failures from expected error types (timeouts, policy rejections,
spawn failures) are recorded as None fields in the snapshot rather
than raised; a missing CLI never fails the agent. A failed first probe
does not poison the cache — the next call retries.
When commands originate from untrusted input (e.g. the model is acting on
prompt-injected document content), prefer DockerShellTool. With the
default isolation flags and a trusted container runtime, the container
is the intended security boundary and approval gating becomes optional.
import asyncio
from agent_framework.tools import DockerShellTool
async def main() -> None:
async with DockerShellTool(
image="mcr.microsoft.com/azurelinux/base/core:3.0",
approval_mode="never_require", # container is the boundary
) as shell:
result = await shell.run("uname -a && id")
print(result.stdout)
asyncio.run(main())Defaults applied to every container:
--network none— no host or external network.--user 65534:65534— runs asnobody:nogroup.--read-onlyroot filesystem; only mounted host paths are writable.--cap-drop ALLand--security-opt no-new-privileges.--memory 512m,--pids-limit 256, ephemeraltmpfs /tmp.
To expose a host directory, pass host_workdir="/path" (mounted
read-only by default; mount_readonly=False to allow writes). Swap the
container runtime with docker_binary="podman".
| Use case | Tool | Sandbox |
|---|---|---|
| Run code (untrusted) | HyperlightCodeActProvider.execute_code (agent-framework-hyperlight) |
Hyperlight WASM microVM |
| Run shell (untrusted) | DockerShellTool |
OCI container (network-off, non-root, capabilities dropped) |
| Run shell (trusted dev) | LocalShellTool |
None; human approval for agent calls by default |
agent-framework-hyperlight is a code sandbox (a single WASM guest
loaded into a microVM, called via a hostcall ABI — there is no kernel,
userland, or shell binary inside). It is the right tier for executing
generated code. For sandboxing shell commands, the realistic tier is
OCI, which DockerShellTool provides.