Skip to content

Repository files navigation

Agentic Platform Engineering Extravaganza

License: MIT Policy engine: OPA 1.19.1 Conftest 0.69.0 Score k8s 0.16.0 MCP 2025-06-18 Verify: 15 checks Cluster: kind v1.36.1 Open in GitHub Codespaces Deploy to Vercel

Hand a capable model one sentence — "a PCI-scope payments service with a Postgres database, EU residency, staging and prod" — and it returns Kubernetes manifests that apply cleanly and fail 42 policy checks across 12 rules. Route the same sentence through a golden path, a provisioner set, a policy bundle and an authorization model, and it converges to zero in three iterations, inside budget, with a human still holding the production approval.

The agent is not the platform. This repository is that comparison, executable.

The argument in four acts

A real terminal recording of ./run.sh demo --acts 2,3,5,7 --scorecard — a command you can type yourself. Every cast in recordings/ is a PTY capture with wall-clock timings; nothing is re-enacted and no line is edited in afterwards. agg plays it back faster than life and trims dead air over 1.5s; those are the only two liberties, and both are declared in src/build_casts.py.

Every policy verdict below came from running the real conftest binary against the real Rego in policy/. Every manifest came from running the real score-k8s binary against the provisioner set in platform/. Run it twice and the numbers are identical.


Table of contents


Why this exists

The interesting question in 2026 is not whether an AI agent can write a Kubernetes manifest. It obviously can. The question is what happens to your estate when a few hundred engineers each have one, and the answer depends almost entirely on what the agent is allowed to render into.

Three numbers frame it:

68% / 50% of developers save 10+ hours a week with AI — and lose 10+ hours a week to organisational inefficiency. Developers spend 16% of their time writing code.
Atlassian, State of DevEx 2025, n=3,500, fielded March 2025
"negligible" DORA's 2025 finding on what AI does for organisational performance when platform quality is low. Platform engineering is the moderator, not a nice-to-have.
DORA, State of AI-assisted Software Development 2025, ~5,000 respondents
57% / 61% of engineers still wait on a human or a ticket to get an environment; 61% call environment provisioning a major roadblock.
Rafay-commissioned survey, 500+ practitioners, 2023

The bottleneck moved to exactly where platform engineers already live. That is the opportunity — and the risk, because an agent inherits whatever the platform gives it. A golden path makes an agent boringly correct. Its absence makes an agent a very fast junior engineer with production credentials and no supervision.

This demo argues that agentic platform engineering needs more platform engineering, not less — and then proves the claim by running the tools rather than describing them.


Quick start

No accounts, no API keys, no cluster, no cloud. Python 3.10+ and about ninety seconds.

In a browser — GitHub Codespaces

Open in GitHub Codespaces

The dev container installs PyYAML, fetches the pinned upstream binaries and runs the acceptance suite before it hands you a prompt, so the first thing you see is 14/14 passing. Then:

./run.sh demo

On your machine

git clone https://github.com/adventurewave-labs/agentic-platform-engineering-extravaganza
cd agentic-platform-engineering-extravaganza

./run.sh setup     # fetch the pinned upstream binaries (conftest, opa, score-k8s, kube-linter)
./run.sh demo      # the full run: eight acts and a scorecard
./run.sh verify    # 15 acceptance checks against real tool output
./run.sh test      # 120 unit tests on stdlib unittest, no extra dependencies

run.sh checks your Python version and installs the one dependency it needs if it is missing, rather than raising a traceback out of an import.

Other entry points:

./run.sh act 5                 # just the policy gate and the remediation loop
./run.sh demo --acts 2,3,5,7 --scorecard   # the argument, without the setup
./run.sh tools cost-reviewer   # the MCP tool list a read-only agent receives
./run.sh gate outputs/final-manifests.yaml
./run.sh mcp 8099              # the platform MCP server, streamable HTTP
./run.sh drift                 # the day-2 drift agent
./run.sh site                  # the showcase page on :8080
make help                      # everything, with descriptions

The showcase page reads in English and Spanish — the switch is in the nav, it follows navigator.language on a first visit, and it remembers your choice. Terminal output, YAML and shell commands stay in the original on purpose: those are verbatim conftest, score-k8s and kube-linter output, and translating them would be the one fabricated thing on a page that argues nothing here is fabricated. make site-build refuses to render when the English and site/i18n.es.json have drifted apart, so the two languages cannot end up making different claims.

Or with Docker:

docker compose run --rm demo
docker compose run --rm verify
docker compose up site         # http://localhost:8080
docker compose up mcp          # http://localhost:8099/mcp

The eight acts, and a scorecard

ActWhat happens
I
the problem
One sentence becomes six tickets across six teams: platform, security, cloud-infra, network, FinOps, release engineering. Eleven working days. About four hours of that is work. The rest is queue.
II
no platform
Give the sentence to a capable model and ask for manifests. What comes back is not stupid — it is plausible. It parses, it applies, it would pass a distracted review. Conftest and kube-linter return 42 denials:
NW-K8S-003 ×10no owner, no cost centre, no data classification — nobody can even tell who this belongs to
NW-K8S-002 ×6no resource requests or limits on any container
NW-K8S-004 ×6may run as root, writable root filesystem, capabilities not dropped
NW-K8S-005 ×4no liveness or readiness probes
NW-K8S-001, -008 ×2 each:latest tags, from an unapproved registry
NW-K8S-006 ×2a literal database password in the pod spec
5× KL-* ×2 eachthe same failures, independently confirmed by kube-linter

Worth noticing what is not in that list. The PCI rules never fire, and neither does the production availability floor — because those rules key off labels (northwind.io/data-classification, northwind.io/environment) that the unguided output does not carry. An unlabelled workload does not merely fail your compliance checks; it is invisible to them. Getting NW-K8S-003 right is what lets four further rules — NW-K8S-007, NW-PCI-001, NW-PCI-005, NW-PCI-006 — see the workload at all. That is not a rhetorical number: it is the difference this repository's own playground data records between the labelled and unlabelled runs.

Nothing in that prompt told the model what PCI means at Northwind, which registry is approved, or that this cost centre has $400 a month. That knowledge is not in any model. It is in a platform.

III
the platform API
The agent connects over MCP as a principal, and receives 12 of 14 tools. It queries the catalog, reads the agent fleet as AiResource entities, and lists the golden paths. The two tools it does not receive are the point.
IV
render
resources: {db: {type: postgres}} — the entire database request a developer writes — becomes a Crossplane v2 namespaced composite with encryption, a private endpoint, sized backups and a cost-centre tag, because platform/northwind.provisioners.yaml says that is what "postgres" means here. Swap that file and every service in the estate gets the new definition on its next render.

The platform also injects, unasked: hardened security context, dropped capabilities, read-only rootfs, probes, topology spread, a PodDisruptionBudget, a default-deny NetworkPolicy, no automounted service-account token, and ownership labels on every object. Every Act II violation with exactly one correct answer is now structurally impossible, and no model was consulted about any of it.
V
the gate
What the platform cannot decide for you is judgement: how big, which region, how long to keep backups, how much to spend. The agent read "high volume at peak" and picked db.r6g.xlarge, multi-AZ, 500 GiB — $960.82/mo against a $400 envelope — and left the backup window at the 7-day default inside PCI scope.

Three real conftest evaluations later it is at zero. Each denial carries its own remediation after ->; the agent reads that, changes the input, and the whole artefact is regenerated. It never edits rendered output to make a check go green — that is the failure mode that makes automated remediation worse than none, because the finding returns silently on the next render.
VI
money
$960.82 → $359.15/mo. $602/mo on one service, caught before anything was provisioned. And the backup window went the other way, 7 days to 30, because the cost gate is not allowed to win an argument with the PCI gate.
VII
propose, then stop
The agent opens a pull request whose body carries the policy verdict and the cost delta — the reviewer sees evidence, not a diff to audit by hand. Then it tries to approve its own production promotion, and the platform declines. Not the model declining: the platform.
VIII
day two
Four weeks on. 18 replicas instead of 6 (kubectl scale at 02:14, never reconciled; 12 extra × $16.43 = $197/mo, priced from the same rate card as everything else here). The PCI backup window silently cut to 7 days by a legacy Terraform pipeline. The NetworkPolicy gone for fifteen days after a namespace was recreated. All attributed, all proposed as a diff — the drift agent holds no tool that can mutate a cluster, so when it is wrong, and it will be, git is still the truth.
—
scorecard
Eleven working days of queue → a couple of seconds of compute. 6 handoffs → 1. 42 violations → 0. $602/mo recovered. Zero cluster mutations by any agent, by construction.

The scorecard deliberately does not print a speed-up multiple. Eleven days against two seconds is a number in the hundreds of thousands, and it would be arithmetically true and rhetorically worthless — those eleven days were never work. What actually changed is that the waiting is gone and the review got better evidence.

Act V — the policy gate, and the agent closing it

Real OPA denials, read and acted on. Three iterations, zero at the end. Nothing here is scripted: the same conftest binary you would run in CI produces every line.

The policy gate and the remediation loop

Act III — the tool an agent is never told about

The same MCP server, two identities, two different worlds. platform.approve_promotion is not refused at call time; it is absent from the list.

MCP tool list filtered by calling identity

Act IV — three lines of Score become a Crossplane composite

resources: {db: {type: postgres}} in, encrypted-and-tagged managed database out — because platform/northwind.provisioners.yaml says that is what "postgres" means here.

score.yaml rendered through the Northwind provisioner set

Act VIII — day two, with attribution

18 replicas where git says 6. The PCI backup window silently cut to 7 days. The NetworkPolicy gone for fifteen days. Detected, attributed, and proposed as a diff — never applied.

Drift detection with attribution

The complete eight-act run is gifs/wow.gif (about 1 MB). Every .cast is in recordings/ and plays with asciinema play recordings/wow.cast — and because these are genuine PTY captures, the playback timing is the timing the demo actually had.


The idea worth stealing: authz-gated MCP

If you take one thing from this repository, take this.

Every tool the platform exposes declares a permission, and tools/list is filtered by the calling identity before the model ever sees it:

$ ./run.sh tools platform-agent
identity: agent:platform-agent  (12/14 tools visible)

    catalog.get_entity                 catalog:read
    catalog.query                      catalog:read
    catalog.refresh_entity             catalog:refresh
  ! catalog.register                   catalog:write
    platform.estimate_cost             finops:read
    platform.evaluate_policy           policy:evaluate
  ! platform.open_pull_request         delivery:propose
    platform.plan_promotion            delivery:plan
    platform.render_workload           platform:render
  ! scaffolder.execute                 scaffolder:execute
    scaffolder.get_template_parameters scaffolder:read
    scaffolder.list_templates          scaffolder:read

  withheld from this identity:
    x platform.approve_promotion       identity lacks delivery:approve
    x platform.detect_drift            identity lacks observability:read

! marks a tool that writes. Run it yourself — that block is copied from the command's actual output, not retyped.

An agent without delivery:approve is not instructed to avoid approving production. It is never told the capability exists. A system prompt that says "never approve production" is a request that can be argued with. An empty tool list is a fact that cannot.

Two layers, deliberately:

  1. Identity → actions. Twelve platform actions in src/catalog.py, granted per principal. platform-agent holds delivery:propose and never delivery:approve.
  2. AiResource.spec.allowedTools. A git-tracked, owned, reviewable catalog entity that narrows an agent further. An agent's blast radius is a YAML file a human reviews and an auditor can diff — not a prompt edit, not a toggle in a vendor console.
apiVersion: backstage.io/v1alpha1
kind: AiResource
metadata:
  name: platform-agent
spec:
  type: agent
  owner: group:default/platform-team
  identity: platform-agent
  # The blast radius: twelve entries, and note which one is absent.
  allowedTools:
    - catalog.query
    - catalog.get_entity
    - catalog.refresh_entity
    - catalog.register
    - scaffolder.list_templates
    - scaffolder.get_template_parameters
    - scaffolder.execute
    - platform.render_workload
    - platform.evaluate_policy
    - platform.estimate_cost
    - platform.plan_promotion
    - platform.open_pull_request
    # platform.approve_promotion is not here, and never will be.

Four identities ship, and they receive genuinely different worlds — try each:

./run.sh tools platform-agent    # 12/14 — can propose, cannot approve
./run.sh tools drift-agent       #  5/14 — read-mostly, cannot mutate a cluster at all
./run.sh tools cost-reviewer     #  3/14 — strictly read-only
./run.sh tools release-manager   # 14/14 — a human in the owning group

Credit where it is due: this pattern is lifted from OpenChoreo (Apache-2.0, CNCF Sandbox), whose control plane routes every MCP tool call through the same policy decision point that governs human API access — see pkg/mcp/tools/*.go. It is the only OSS internal developer platform I found that does this, and it deserves far more attention than it gets.


The policy bundle

Seventeen controls across three bundles — 34 deny/warn rule bodies — all real Rego, all evaluated by conftest.

Bundle Rules Covers
policy/kubernetes/workload.rego NW-K8S-001…008 immutable image refs, resource envelopes, ownership and cost labels, restricted pod security, probes, no plaintext credentials, production availability floor, approved registries
policy/kubernetes/pci.rego NW-PCI-001…006 default-deny networking both directions, encrypted storage, 30-day backup floor, no public database endpoints, EU data residency, audit sink, no automounted service-account token
policy/cost/budget.rego NW-FIN-001…006 cost-centre envelope, no production-grade instances outside production, no unreviewed single line item over 70%, cost-allocation tags, no multi-AZ in non-prod, spend-spike warning
policy/score/workload_spec.rego NW-SCORE-001…005 required platform annotations, supported resource types only, credentials by reference, declared resource envelope, residency declared for PCI

Every message is shaped so an agent can act on it:

[NW-PCI-003] SQLInstance/nw-payments-ledger-prod: backupRetentionDays=7 is below the
             30-day PCI floor -> set spec.parameters.backupRetentionDays>=30

The -> half is what makes the loop agentic rather than merely a linter. The agent parses it, maps it to a Score parameter, and re-renders. If a denial cannot be mapped back to an input, the agent escalates rather than papering over it — see unresolved() in src/agent.py.

Run the bundle against anything you like:

./bin/conftest test --policy policy/ your-manifests.yaml

Point your own agent at it

The MCP server is a real one — MCP 2025-06-18, stdio and streamable HTTP, and no MCP SDK: the protocol is implemented directly, against PyYAML and the standard library.

# stdio
claude mcp add northwind -- python3 src/platform_mcp.py

# or HTTP
./run.sh mcp 8099
claude mcp add --transport http northwind http://127.0.0.1:8099/mcp

See examples/mcp.json for Cursor, Copilot and other clients.

Then ask it for a service. It will list the golden paths, read the parameter schema, render, evaluate, remediate, and stop at the pull request — because that is the last thing its identity is permitted to do.

The fourteen tools
Tool Permission Does
catalog.query catalog:read search entities by kind, owner, tag or text
catalog.get_entity catalog:read fetch one entity in full
catalog.refresh_entity catalog:refresh re-queue and read back fresh state — this is what closes the scaffold→verify loop
catalog.register catalog:write register or update an entity
scaffolder.list_templates scaffolder:read the golden paths on offer
scaffolder.get_template_parameters scaffolder:read JSON Schema for a template
scaffolder.execute scaffolder:execute run a golden path; renders Score, manifests, Argo CD, Kargo
platform.render_workload platform:render Score spec → manifests via score-k8s
platform.evaluate_policy policy:evaluate run the real policy bundle
platform.estimate_cost finops:read Infracost-shaped estimate per environment
platform.plan_promotion delivery:plan what moves, what gates it, who must approve
platform.open_pull_request delivery:propose propose the change
platform.approve_promotion delivery:approve no agent identity holds this
platform.detect_drift observability:read observed vs desired, with attribution

catalog.refresh_entity mirrors the catalog refresh action Backstage added in the v1.54.0 line (catalog:refresh-entity), which exists precisely so an agent can scaffold something and immediately read the catalog's view of it instead of racing the processing loop. Without it the agentic golden path does not close.


The stack, and why each piece is here

Everything is free, open source and currently maintained. ./bin/setup.sh fetches official release artefacts at pinned versions — nothing is vendored or reimplemented.

Project Role here License Version Agentic surface today
Conftest / OPA evaluates every verdict in this demo Apache-2.0 0.69.0 / 1.19.1 the gate itself
score-k8s / Score developer-facing spec; renders the manifests Apache-2.0 0.16.0 small schema'd YAML a model gets right
kube-linter independent second opinion Apache-2.0 0.8.3 verification
Crossplane v2 the platform API the provisioner renders into Apache-2.0 v2.4.0 ⚠️ no official MCP server
Backstage catalog + Software Template shape; AiResource Apache-2.0 v1.54.3 first-party MCP Actions backend
Argo CD reconciles the merged change Apache-2.0 v3.5.1 argoproj-labs/mcp-for-argocd
Kargo promotion path with a human gate Apache-2.0 v1.11.2 ⚠️ MCP proposal closed not_planned
OpenChoreo source of the authz-gated MCP pattern Apache-2.0 v1.2.3 3 MCP servers, 3 in-tree agents
Trivy / OpenTofu optional scanners and IaC toolchain Apache-2.0 / MPL-2.0 0.74.0 / 1.12.6 optional
Checkov optional IaC gate — not fetched by bin/setup.sh; gates.py runs whatever is on $PATH Apache-2.0 unpinned, by design optional

Deliberately excluded, with reasons — because a survey that only lists winners is not a survey:

  • Cyclops (3.3k★) — genuinely nice idea, but last commit July 2025, and its MCP server has no LICENSE file at all.
  • Kusion (1.3k★) — two commits in thirteen months.
  • Port — closed-source core, port-mcp-server archived in favour of a hosted endpoint, AI agents behind a paid tier. Fails the free-and-OSS bar.
  • Community Crossplane MCP servers — 0–2★ each, one archived. Not demo-grade. If you need agentic Crossplane today, drive it through the Kubernetes API or Backstage MCP Actions.

Also worth your time, not used here: kagent (CNCF Sandbox — agents as Kubernetes CRDs, the best story in this space), kubectl-ai (7.5k★, first-class Ollama support), k8sgpt and HolmesGPT (both CNCF Sandbox — note HolmesGPT moved out of robusta-dev/ into its own org), CAIPE + idpbuilder (CNOE's multi-agent platform and the best "whole IDP on a laptop" experience), and Dagger's LLM primitive with two-way MCP.


What is real and what is a fixture

A demo that overstates itself is worse than no demo. The line, drawn honestly:

Component Status Detail
Policy evaluation ✅ real the pinned conftest binary over the Rego in policy/. Every number in this README comes from it.
Manifest rendering ✅ real the pinned score-k8s binary with the provisioner set in platform/.
kube-linter ✅ real an independent linter nobody here tuned. It found four genuine defects in the platform defaults during development — missing containerPort, no anti-affinity, an unresolvable ServiceAccount reference, and an unset unhealthyPodEvictionPolicy. All four were fixed in the platform, not worked around in the demo.
MCP server ✅ real MCP 2025-06-18, stdio + streamable HTTP. ./run.sh verify exercises initialize, tools/list, tools/call and resources/list.
Authorization ✅ real enforced in code at tools/list and tools/call, not by prompt instruction.
The agent's reasoner ⚠️ deterministic by default denials are parsed and mapped to Score parameter changes by rules, so the demo reproduces exactly with no API key. --backend llm swaps in a real model against any OpenAI-compatible endpoint (Ollama, vLLM, OpenRouter, Z.AI, OpenAI itself); the loop is unchanged. That it makes no difference to the outcome is the argument — and T15 is that argument executed: it drives the loop through the LLM code path against a recorded transcript and asserts it lands on the same manifests and the same $359.15.
Intent extraction ⚠️ regex prose → structured request. A model does this better on messy input and worse on reproducibility. One function, replaced by --backend.
Kubernetes cluster ⚠️ not in the demo; real in CI ./run.sh demo applies nothing anywhere — every claim it makes lives in rendering and evaluation. But "valid against a real cluster" is not left as an assertion: .github/workflows/cluster.yaml stands up a pinned kind cluster, installs the platform API as a CRD, and puts the committed manifests through the real API server — schema validation, admission, defaulting — then asserts the controllers acted on them. It also runs the negative control: the same API server accepts the unguided manifests too, all 42 violations of them.
Cloud provisioning ❌ not present the SQLInstance is a real Crossplane-shaped composite; no Composition is installed, no cloud account is touched.
Cost figures ⚠️ static rate card by default a checked-in table in src/costing.py, so there is no account requirement and the artefacts stay byte-reproducible. "Swap in Infracost and the Rego does not change" is wired rather than asserted: NORTHWIND_COST_SOURCE=infracost re-prices the database and load-balancer lines against real cloud rates via src/sources/infracost.py, and every estimate carries a costSource so a rate-card number is never mistaken for a priced one.
Drift observation ⚠️ fixture by default platform/observed-state.yaml. Detection, attribution and the proposed patch are real code over that shape. NORTHWIND_DRIFT_SOURCE=argocd swaps the fixture for src/sources/argocd.py, which reads Argo CD's managed-resources for the desired/observed pair and metadata.managedFields for attribution — and reports a field manager, not a person, because that is all a control plane actually knows.
The GIFs ✅ real recordings genuine PTY captures. A pseudo-terminal is allocated, the command is typed into a real interactive shell, and every byte is timestamped as it arrives — the same thing asciinema rec does, minus the human hand. Wall-clock timings, a real prompt, and the hero reel is a real run of ./run.sh demo --acts 2,3,5,7 --scorecard rather than a full run with lines cut out. agg plays them back faster than life and trims dead air over 1.5s; both are declared in src/build_casts.py and neither can change a character of what was recorded.
Northwind Retail ❌ fictional the company, the teams, the ticket numbers. The pain is not.

./run.sh verify runs 15 acceptance checks and is the thing to trust rather than this table: the toolchain is present, the Rego compiles under opa check --strict, the unguided path is still rejected with exactly the recorded number of findings, the golden path still converges to zero, kube-linter still finds nothing in the platform's output, the agent still changes only inputs — proven by re-rendering, not by reading its own account of itself — the agent still cannot approve production while a human still can, tool visibility is still ordered by privilege, the MCP protocol still answers, the LLM backend still reaches the same answer as the deterministic one, and the committed run record still reproduces exactly (the unguided finding count and its per-policy breakdown, the iteration count, the cost before and after, the rendered object count, and a SHA-256 of the rendered manifests — seven properties, each compared, none merely printed).

Three of those sentences were not true before this round. The check that claimed the Rego compiled ran a command that could not have noticed if it did not; the check that claimed the agent never edits rendered output only ever inspected the agent's own decision records; and the unguided path was held to "at least 20" against an actual figure of 42. Each is now written so that breaking the thing it describes makes it fail — which was tested by breaking them.

./run.sh test runs 120 unit tests underneath that — stdlib unittest, no new dependency. They cover the branches the worked example never reaches: the intent extractor's whole surface, the staging-only FinOps rules, an unmappable denial, a model that replies with prose instead of JSON, a model that invents a field, and the Argo CD adapter's normalisation without an Argo CD. A system check tells you the demo broke; a unit test tells you which function did it.

tests/test_docs.py holds this file to the same standard. Every number the README and the showcase page share — the size of the suite, the denial count, the cost before and after, the versions pinned in the stack table, a GIF size quoted to a decimal place — is checked against the code or the committed artefacts rather than against the other document. It also fails on a dangling README link, and on a chmod in the dev-container bootstrap that names a path which does not exist. Both of those had happened, and neither was caught by anything until it did.


Repository layout

.
├── index.html                          the showcase page (single file, no build step)
├── site/index.template.html            its source; regenerate with `make site-build`
├── site/i18n.es.json                   the Spanish half of the page, and the drift guard
│
├── policy/                             ← the part worth stealing first
│   ├── kubernetes/workload.rego        NW-K8S-001..008
│   ├── kubernetes/pci.rego             NW-PCI-001..006
│   ├── cost/budget.rego                NW-FIN-001..006
│   └── score/workload_spec.rego        NW-SCORE-001..005
│
├── platform/
│   ├── northwind.provisioners.yaml     ← and this second: what "postgres" means here
│   ├── observed-state.yaml             the day-2 drift fixture
│   ├── crds/sqlinstance.yaml           the platform API as a real CRD, for the cluster job
│   └── catalog/
│       ├── org.yaml                    Groups, Users, cost centres, budgets
│       ├── systems.yaml                Components, Resources, the platform API
│       ├── ai-resources.yaml           ← the agent fleet and its blast radius
│       └── templates/
│           └── golden-path-service-postgres/template.yaml
│
├── src/
│   ├── platform_mcp.py                 the MCP server — 14 tools, authz-filtered
│   ├── catalog.py                      catalog + the authorization model
│   ├── renderer.py                     both paths: no-platform, and the golden path
│   ├── gates.py                        thin wrappers over the real binaries
│   ├── agent.py                        intent extraction, remediation, LLM backends
│   ├── costing.py                      Infracost-shaped estimates
│   ├── driftd.py                       the day-2 drift agent
│   ├── sources/argocd.py               live observed state, instead of the fixture
│   ├── sources/infracost.py            real cloud prices, instead of the rate card
│   │                                   (sources/__init__.py states the contract)
│   ├── goldenpath.py                   the eight-act orchestrator
│   ├── ui.py                           terminal presentation
│   ├── build_report.py                 the 15 acceptance checks, timed
│   ├── build_casts.py                  records a real PTY session, renders the GIF
│   ├── build_playground.py             real gate results for the interactive page
│   └── build_site.py                   renders index.html from the template
│
├── outputs/                            committed artefacts from the last run
│   ├── run-record.json                 what `verify` checks against
│   ├── final-score.yaml                the developer-facing contract
│   ├── final-manifests.yaml            7 objects, 0 denials
│   ├── kargo-pipeline.yaml             Warehouse → staging → prod
│   ├── verify-report.json              the 15 checks, with real durations
│   ├── final-cost.json                 the FinOps document the Rego evaluates
│   ├── vibe-policy-report.json         the unguided run's 42 denials, in full
│   ├── drift-report.json               the day-2 findings with attribution
│   └── playground.json                 recorded conftest output for the page
│
├── tests/                              120 unit tests + the fake OpenAI endpoint
│   ├── fake_llm.py                     replays a recorded transcript; no key, no spend
│   ├── test_docs.py                    holds the README and the page to the same standard
│   ├── context.py                      puts src/ on the path; __init__.py alongside
│   └── test_agent.py  test_gates.py  test_costing.py  test_drift.py
│
├── .github/workflows/
│   ├── verify.yaml                     the 15 checks + the unit tests, on a clean checkout
│   ├── cluster.yaml                    a real kind cluster, and the negative control
│   ├── gifs.yaml                       re-renders the GIFs from the committed casts
│   └── permissions.yaml                restores the exec bit on the shell entry points
│
├── examples/mcp.json                   drop-in config for any MCP client
├── recordings/  gifs/  captured/       demo assets, all regenerable
├── .devcontainer/                      one-click GitHub Codespaces
├── bin/setup.sh                        fetches the pinned upstream binaries
├── run.sh  Makefile  requirements.txt  entry points
├── LICENSE  NOTICE                    MIT; upstream licenses enumerated separately
└── Dockerfile  docker-compose.yml  vercel.json

Extending it

Change what "postgres" means. Edit platform/northwind.provisioners.yaml and re-run ./run.sh demo. Every service in the estate gets the new definition on its next render. This is the whole leverage of platform engineering in one file, and it is the leverage an agent inherits for free — the agent never has to know how to provision a database correctly, because the only database it can ask for is the correct one.

Add a policy. Drop a .rego file in policy/. Give the message a [POLICY-ID] prefix and a -> how to fix it suffix and the agent will act on it. Add a branch in remediate() in src/agent.py to map it to a Score parameter; leave it out and the agent will correctly escalate instead of guessing.

Add a golden path. Copy the template directory. scaffolder.list_templates picks it up with no code change.

Change an agent's blast radius. Edit spec.allowedTools in platform/catalog/ai-resources.yaml. Confirm with ./run.sh tools <identity>. Note that this is a pull request, which is the point.

Use a real model.

export NORTHWIND_LLM_BASE_URL=http://127.0.0.1:11434/v1   # Ollama, vLLM, OpenRouter, Z.AI...
export NORTHWIND_LLM_MODEL=qwen2.5-coder:7b
./run.sh demo --backend llm

The loop is identical. If the model is any good, the outcome is identical too — because the platform, not the model, is what determines it.

Wire it to a real cluster. Replace the fixture in platform/observed-state.yaml with output from argoproj-labs/mcp-for-argocd or containers/kubernetes-mcp-server, install a Crossplane Composition matching the SQLInstance XRD, and point score-k8s at your kubeconfig. Nothing in policy/ or src/gates.py changes.


Research notes

The tool survey behind this repository was done in August 2026 against primary sources — GitHub release pages, official docs, CNCF project pages. A few findings worth recording, because they are easy to get wrong:

  • There is no official Kubernetes MCP server. kubernetes-sigs/mcp-server does not exist. containers/kubernetes-mcp-server is the de-facto leader and moved into the neutral containers org in July 2025 specifically for community governance — but per its own maintainer, standard-bearer status is "aspirational rather than established". Do not call it official.
  • kro moved. It is kubernetes-sigs/kro now, a Kubernetes SIG Cloud Provider subproject — not CNCF, and no longer at kro-run/kro.
  • HolmesGPT moved out of robusta-dev/ into its own HolmesGPT GitHub org.
  • akuity/argocd-mcp moved to argoproj-labs/mcp-for-argocd.
  • Kargo has no MCP server and is not getting one soon — issue #6212 was closed not_planned in May 2026 and the accompanying PR closed unmerged.
  • Crossplane has no official MCP server. A search of the repo for modelcontextprotocol returns nothing, and the community attempts are 0–2★ or archived.
  • The Pulumi local OSS MCP server appears discontinued — the repo 404s and npm has been stale since September 2025; the documented path is now a hosted Pulumi Cloud endpoint.
  • There is no mature OSS "AI writes your Rego" tool. That gap is arguably this repo's thesis: the valuable artefact is the policy a human wrote, not the policy a model generated.

Related

  • Agentic DevOps Extravaganza — the sibling repo: K8sGPT and Robusta doing autonomous incident triage against a real Kubernetes API and a real LLM.

License

MIT — see LICENSE. The upstream tools retain their own licenses (Apache-2.0 throughout, except OpenTofu at MPL-2.0) and none of them are vendored here; they are enumerated in NOTICE.

Northwind Retail is fictional. The policy bundle, the provisioner set, the MCP server and every verdict in this README are not.

About

42→0: a real policy gate (OPA/Conftest) takes an AI agent's Kubernetes manifests from 42 violations to zero. Golden paths, Crossplane, an authz-gated MCP server.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages