Skip to content

Latest commit

 

History

121 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Secure single-server AI agent environment

AI coding agents become useful when they can read a workspace, run tools, and call a model. Those are also the powers that make them risky and expensive to operate: tools can touch unrelated data, model requests can bypass the approved path, provider keys can spread across harness homes, and one long task can consume a shared model.

This repository shows how to run those agents on one administrator-managed RHEL server without giving them that unrestricted power. The harness keeps the developer experience. OpenShell constrains tools and network access. Praxis owns model routing and shared usage. A separate vLLM server can supply an approved private model. bootc makes the host reproducible and rollback-capable.

Start with the architecture value walkthrough. It follows the validated OpenCode → Praxis → vLLM path, shows the commands that prove each boundary, and states what is not yet qualified.

The environment brings together harnesses, OpenShell, Praxis, and separate private or cloud inference. bootc packages the host setup into an updatable OS image. The preferred example runs OpenCode through Praxis against Qwen3-8B in vLLM on a separate CPU or single-NVIDIA-L4 server.

How the pieces fit

Component Responsibility
Harnesses Provide the coding-agent experience: prompts, model interactions, and tool calls. Recipes cover several harnesses; supported integrations differ.
OpenShell Runs the harness and tools inside a sandbox with declarative filesystem and network policies.
Praxis Routes model requests and applies shared request and token limits. Cloud profiles keep provider credentials at the gateway.
Inference backend Supplies the model: optional vLLM for Qwen3-8B on a separate private server, or a cloud provider through a separate Praxis profile.
bootc Packages service setup and the selected harness configuration into a reviewed RHEL OS image, with image-based updates and OS rollback.

Two paths meet at the harness: tools execute within OpenShell's policies, while model requests travel through Praxis. The diagram shows the preferred separate vLLM path and the separate cloud-profile option:

flowchart LR
    subgraph Host["bootc-managed RHEL host"]
        subgraph Sandbox["OpenShell sandbox: filesystem and network policies"]
            H["OpenCode harness"]
            T["Agent tools and workspace"]
            H -->|Tool execution| T
        end
        P["Praxis container<br/>Shared request and token limits<br/>127.0.0.1:8080"]
        H -->|Policy-permitted model requests| P
    end
    V["Separate vLLM server<br/>Qwen3-8B: CPU or NVIDIA L4<br/>Private AWS address:8000"]
    P -->|Private vLLM profile| V
    P -. Separate cloud profile .-> Provider["Cloud model provider"]
Loading

The vLLM profile permits sandbox inference traffic only to Praxis. Direct vLLM and cloud-provider access are denied, and there is no cloud fallback. OpenCode is the recorded harness for this path; Codex and OpenClaw adapters are not enabled.

Pinned workload containers are pulled on first boot and cached across reboots. Credentials, model caches, and workspace data stay outside the OS image. OS rollback restores the host deployment; it does not restore application data or sandbox workspaces.

Choose a deployment

Workflow Where the harness and tools run Guide
RHEL with Qwen and optional cloud providers Ordinary user accounts on all-in-one, or remote clients; CPU/GPU selected independently Install Qwen inference, then add providers
Sandboxed agents with separate inference OpenShell on a bootc-managed server; Praxis routes to a private CPU or NVIDIA L4 vLLM server Qwen3-8B example
Sandboxed harness exploration OpenShell on the server, with harness-specific policies and provider setup OpenShell recipes
Shared host with a cloud gateway Harnesses run directly under OS accounts on RHEL; Praxis owns provider credentials All-in-one gateway
Remote clients with a central gateway Harnesses and tools stay on client machines; requests reach Praxis over HTTPS with caller JWTs Remote gateway

For OS image creation, start with the RHEL 9 bootc guide. The current target is x86_64, with a shared base and separate Codex, OpenCode, and OpenClaw OS variants. A harness image being available does not mean every Praxis/backend combination is supported; consult the integration matrix and the local example's validation report.

Validated today

This is an experimental deployment and validation repository. The strongest recorded end-to-end example is OpenCode → Praxis → vLLM on bootc. AWS tests with OpenShell 0.1.2-rhaiv.0 passed on CPU and NVIDIA L4, covering real Qwen3-8B inference, streamed responses, independently verified tool execution, explicit bypass denials, cached reboot, and disable/re-enable behavior. The local inference report records exact pins and limits, including the GPU instance's cleanup issue.

The preferred topology now places vLLM on a separate server and keeps only OpenShell, Praxis, and the harness on the single server. The AWS helper discovers that server's private address and grants access by security group; its fresh real-inference qualification remains separate work.

The mutable AWS workflow also passed native mocked Qwen/OpenAI/Anthropic tests. With vLLM 0.30, all-in-one real Qwen tasks passed for Codex, Claude and OpenCode on both CPU and GPU with the current Praxis image. See the compatibility matrix for exact pins, remote-gateway baselines, protocol limits and separate OpenShell results.

Earlier AWS testing also exercised bootc builds, Codex/OpenCode boot and CLI execution, OS upgrades, a harness switch, and rollback. See the host validation record. The separate sandbox-to-Praxis cloud-provider path remains under qualification; Codex and OpenClaw reject Praxis configuration.

The current trust and usage boundaries are explicit:

  • OpenShell assumes a trusted single operator. Local management is not a multi-tenant authorization boundary.
  • Praxis quotas are shared token allowances, not per-user limits or USD budgets. bootc uses in-memory quotas; the mutable Praxis deployment offers Valkey for persistent token usage.
  • CPU inference is functional but slow on the tested eight-vCPU host. Additional hardware, sustained load, and coding quality are not qualified by the smoke tests.

See the OpenShell trust model and quota semantics for the control boundaries.

Where this is going

The broader goal is an environment where people can run useful agent tasks with approved models, controlled workspaces, predictable shared usage, and repeatable operations. Local inference now provides a tested foundation for that work. Durable bootc quotas, routing and failover, inference guardrails, individual and team authorization, retained collaborative sessions, and usage visibility remain planned or partially implemented capabilities.

The scope and roadmap separates those goals from current acceptance. To contribute or validate a new combination, use the testing guide.

Upstream projects

This repository supplies deployment configuration, lifecycle scripts, harness recipes, and acceptance tests. It consumes Praxis experimental, OpenShell, and vLLM; it does not implement those runtimes or the harnesses themselves. The older Praxis Ruby framework is a separate project and is not the gateway used here.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages