Skip to content

Repository files navigation

nousergon-console

CI Coverage License

A read-only fleet index that persists nothing and reports what it cannot see.

Most monitoring surfaces are a pile of dashboards. A dashboard is a frozen answer to the questions its author had at authoring time — but monitoring is the business of being asked questions nobody anticipated: why is this stale · what else touched that artifact · what did this cycle cost · which of these has never run · who breaks if this one does. Each of those is a traversal, not a tile.

This is an index over a typed entity graph instead. Point it at the artifacts you already have; it builds a catalog you can search, link to, and walk.

Status: initial implementation. The contract is settled and normative — docs/contract.md is its single normative copy, and the sections below summarise it; the adapter layer, the entity index, search and the server are implemented. The implementation stack is Python + stdlib server — see docs/stack-decision.md and the Roadmap.

What makes it different

Service / metadata catalogs This
Target teams on cloud-native stacks (Kubernetes, Terraform, dbt, Airflow) one operator or a small team, over heterogeneous sources — object-store keys, cron jobs, systemd units, state machines, YAML registries, CI
State a database plus ingestion pipelines persists nothing — every figure is a projection of a fact durable somewhere else
Headline number catalog coverage the transparency-gap count — how much of your system the surface cannot see

That last row is the point. A surface that renders nine of fourteen services in depth is worse than a coarse one rendering all fourteen, because the five it omits are indistinguishable from five that are fine. So this one publishes its own blind spots as a first-class metric, and the objective for that number is zero.

If you already run Backstage, Port, OpenMetadata or DataHub over a standard cloud-native stack, use those — their connector libraries are the reason they exist. This is for the case where none of your sources has a stock connector.

The entity model

Everything rendered is a fact about exactly one of seven kinds:

Kind Is Identified by
Component anything that runs unattended and can fail with no human present component_id
Run one execution of a component run id
Cycle the business period runs belong to — a day, a weekly cadence, a deploy cycle id
Artifact a durable thing produced or consumed — an object key, a table, a report its key
Signal one named measurement over time, with its baseline metric name
Decision a ruling, a gate, a queued question, a policy clause tracker reference
Incident a failure record with its severity and class incident id

Cost is not a kind. It is a facet of Run, Cycle and Component, so "what did this cycle cost" is a traversal rather than a separate system with its own component list that will disagree with yours.

Three ways to reach everything

Every entity is reachable by name (search), by structure (navigation), and by relation (a link from anything adjacent). Three paths, independently sufficient — because a fact reachable only one way is reachable only by someone who already knows it exists, and that is a forensic tool rather than a measurement.

Every list is filterable by facet and by state — /component?state=UNREGISTERED — in both representations of the same URL. That is what makes a number's members navigable rather than merely enumerable: a count published without a way to reach the rows behind it reports a defect nobody can locate, which is how one newly-unregistered component hid inside a hundred-row exception table for two hours.

The relation direction that matters most is the reverse one. What did this produce is usually written down somewhere. Who breaks if this is stale exists nowhere unless the index derives it — and it is the only question an incident actually asks.

Getting a process or module onto it

Write one file. A descriptor, committed beside the thing it describes, naming the component and binding it to where its facts already are:

component_id: my-nightly-job
owner: brian
lifecycle: in-service

artifacts:
  - driver: object-store
    key: "s3://your-bucket/reports/latest.parquet"
    cadence_minutes: 1440

metrics:
  - driver: log-source                    # where the module already logs
    field: rows_written
    unit: rows
    baseline: 900
    render: count

consumes: ["s3://upstream/in.parquet"]

That is the whole integration. No adapter, no entry in the console's config, no edit to this repository — even when the data lives somewhere nothing else in your system uses.

The binding is in the descriptor rather than in the console's configuration for one reason, and it is the difference between free-by-design and free-by-coincidence: console config is per-source. Put the binding there and a component whose data lands somewhere new costs a console edit; put it here and it costs nothing, because the thing that knows where the data is, is the thing that put it there.

  • A driver reads a source shape — object store, state machine, log location, query, emitted envelope. Which bucket, group or table is yours is in your descriptor. A driver that could name an instance is an adapter written for one component wearing a driver's name, and a lint says so.
  • A descriptor naming a driver that does not exist fails the build, loudly. A typo and a genuinely-gone component must not look the same.
  • An artifact binding tells you about the artifact, not the component. A fresh output is not proof the job ran — inferring one from the other is how "it produced yesterday's file again" reads as healthy. Bind runs or metrics for the component's own state; doctor will tell you if you have not.
  • A metric is read from a structured record, addressed by field name — never a pattern matched against prose. A regex over log text couples you to a format nobody promised to keep, and it fails by matching nothing, silently, rendering that as absence.

Where a fact has nowhere natural to live, the module emits the versioned envelope and the descriptor points at that. It is one binding kind, not the path — most facts are already written down somewhere, and asking a module to write them twice is asking it to maintain two truths.

The emission envelope, for facts with nowhere else to live

A module emits, adds a registry row, and appears — with zero edits to this repository and zero edits to the console's configuration. That is the default path, and it is the whole path:

from console.emit import report, write_json

write_json(
    report(
        component_id="my-nightly-job",
        status="ok",                      # about this RUN: ok · attention · error
        cadence_minutes=1440,             # without this, staleness is not computable
        summary="processed 903 tickers, 0 rejected",
        deep_link="https://ci.example/run/1234",
        consumes=["s3://bucket/upstream.parquet"],
    ),
    "/var/lib/reports/my-nightly-job/latest.json",
)

Every field is optional with a declared default. A required field added later is a fleet-wide breaking change dressed as a schema improvement, and it lands on every emitter at once — most of which nobody is going to redeploy. The schema is published at console/schemas/component_report.schema.json and is part of the product contract.

A module's own numbers come along without any rendering code that knows about it. fields carries a descriptor per value:

fields={
    "tickers_scanned": {"value": 903, "unit": "tickers",
                        "baseline": 900, "render": "count"},
    "p99_latency":     {"value": 4.2, "unit": "s",
                        "baseline": None, "render": "duration"},
}
  • unit is required for a number. A measurement whose unit is inferred from context is the defect that emitted a normalized ratio, consumed it as raw share volume, and silently failed 901 of 903 tickers for months.
  • baseline: null is a declaration, and the number then renders as telemetry — plain, uncoloured. Green means better than the baseline; where there is no baseline there is no colour. An absent baseline and a null one are different facts and render differently.
  • An unrecognised field renders opaque and is counted, never dropped. A dropped field is a fact the emitter believes is on the surface and is not, and it fails silently on their side of a boundary they cannot see.
  • The render set is closed — an open one becomes a plugin API, and a plugin API is a per-module rendering path with a nicer name.

The cost of onboarding the next thing is a published number (target: zero edits). A surface whose coverage is bounded by how much adapter code somebody felt like writing will always render a subset while looking complete.

Writing an adapter is the exception, not the path — reserved for sources you do not control: a vendor API, a cloud control plane, someone else's registry. A source you write to is a source you can make emit.

Every view is also JSON, at the same URL

curl -H 'Accept: application/json' https://console.example/component/my-job
curl https://console.example/component/my-job.json          # same thing
curl https://console.example/.json                          # the landing view

Agents are a first-class reader of this surface, not an afterthought. One they cannot read forces every agent to re-derive system state from raw sources — and the agent's picture and the operator's picture then diverge exactly when something is wrong, because that is when the derivations differ.

It is not an API beside the UI: the resolver runs once and both representations render from its result. There is one router and one query, so the two cannot drift in coverage — a route that exists serves both, a route that does not 404s in both, and adding a route cannot add it to one and forget the other.

  • Accept: application/json is primary. */* deliberately does not count — a browser sends it, and defaulting that to JSON would make the human surface unreachable.
  • The .json suffix is a fallback, and only applies to a path that does not already resolve — because an artifact's identifier is its object key, so /artifact/ops/checks/x/latest.json names a real entity.
  • A 404 answers in the representation that was asked for. An agent that gets an HTML error page has to parse prose to learn what happened.
  • Every payload carries schema_version.

When something is not on the surface, the surface says why

$ console doctor my-nightly-job
my-nightly-job: adapter claim — nothing reported on it; 2 sources were asked (registry, checks)

  [ok  ] registry row: declared by registry
  [FAIL] adapter claim: nothing reported on it; 2 source(s) were asked (registry, checks)
         → check that the emission is landing where an enabled adapter reads, and that
           the adapter's key_pattern matches the path it is written to. An adapter that
           ran fine and found nothing looks exactly like one that was never enabled —
           compare its status in the index freshness block.

Onboarding fails silently by construction: nothing raises when a module is absent, and the absence looks exactly like "there is nothing to show". doctor walks the chain that would have put an identifier on the surface — registry row → adapter claim → merged entity → reachable by name, by structure, by relation — and names the first broken link with a next step. Reporting every failure at once buries the one that caused the others.

It is addressable too (/doctor/<id>, both representations), because a diagnosis nobody can link to has to be re-run by whoever gets asked about it. It exits non-zero, so it works as a check in a deploy script rather than only by eye.

Adapters

One adapter per source of truth. An adapter is a function from configuration to entities and edges, and it is the only thing that knows its source's shape.

  • Adapters do not know about each other. Cross-source relations are formed by the index over entity identifiers, never inside an adapter.
  • Every literal naming a bucket, ARN, host, port, path or component comes from configuration. None is compiled in — which is what makes this artifact usable by anyone but its author.
  • An adapter declares what it cannot supply. A source with no freshness stamp or no baseline says so, and its entities carry the corresponding state rather than a silent default.
  • An unreachable source renders its entities UNREPORTED and the adapter FAILED. It never empties the surface and never removes rows.
  • Adding a source is adding an adapter — no change to the model, the index, the router, or any view.

Configuration

config.example.yaml is the tracked shape; your real config.yaml is gitignored. It declares which adapters are enabled, what each points at, and which facets matter in your fleet. Nothing about your topology belongs anywhere else in this repository.

Rendering rules

Every rendered fact carries four fields — state · source · as-of · evidence link. A dot that cannot say how it knows is not yet trustworthy.

  • A row older than its declared cadence renders stale, not as its last value in normal styling.
  • Any roll-up states its denominator inline (12 / 14 reporting). One that cannot is not rendered.
  • A number with no baseline is rendered as telemetry, never coloured as a verdict.
  • Zero, null, empty, never-ran, never-triggered and never-observed are different facts and render as different things. No data is never drawn as green, and never drawn as nothing.
  • Component state comes from one closed, total vocabulary of twelve, and it has no fall-through. There is no UNKNOWN, no PENDING, no N/A: where the classifier cannot place a component the answer is UNREPORTED, which is loud and is a finding. What a component is deliberately not doing — DISABLED, DEPRECATED, RETIRED — is declared, never inferred, because a decision and a defect are indistinguishable from telemetry alone.
  • Not everything is a component. An artifact is fresh or stale; an issue is open or closed. Those rows carry the source's own value rather than being forced into a vocabulary that has no word for them.
  • Every addressable state has a URL built from entity identifiers. Paste it anywhere; it reproduces the view on a cold load.
  • The index has an as-of too, and it bounds every row's. Every page and every payload carries the index build time, the rebuild cadence, and each source's read. A surface built once at start and served all day renders every row frozen at boot while looking exactly like a live one — and that is invisible precisely because the rows stay internally consistent with each other. Past the shortest cadence its sources declared, the whole surface says so; a rebuild that fails keeps serving the previous index and marks it, rather than emptying the surface or pretending to be current.

Roadmap

  1. The adapter contract and reference adapters — done (filesystem/YAML registry, object store, checks-envelope, state-machine, Git host API, systemd units). See docs/adapters.md.
  2. The entity index, its URL scheme, and the relation graph — done.
  3. Generated navigation, global search, entity pages — done.
  4. The seven self-grading numbers, published on the surface itself.

The implementation stack is Python + stdlib server — decided in docs/stack-decision.md.

Development & testing

What it is. A local checkout of the same package the published nousergon-console entry point installs — no separate dev build, no compiled step.

Why it exists. The console is read-only and persists nothing (see above), so the test suite is the only place its correctness claims — the closed state vocabulary, the driver/adapter contract, the JSON/HTML parity, doctor's broken-link diagnosis — are actually checked; nothing about them is visible from running the server against real data once.

How to run it.

python3 -m venv .venv && .venv/bin/pip install -e ".[test]"
.venv/bin/python3 -m console               # serves locally against config.example.yaml

How to test it.

.venv/bin/python3 -m pytest tests/ -q --cov=console --cov-report=term-missing
.venv/bin/ruff check --select F821 .       # the undefined-name class that once merged green
.venv/bin/python3 -m console index --config config.example.yaml       # build-time namespace gate

pytest exits non-zero below the coverage floor in pyproject.toml [tool.coverage.report]; the badge above renders the figure CI last measured on main, from the same .coverage file the gate reads.

Where the deeper docs are. CONTRIBUTING.md for the rules every change is held to and how to propose one; docs/contract.md for the normative contracts (adapters, drivers, descriptors, claim merge, the emission envelope, the JSON API); docs/adapters.md for the per-adapter and per-driver reference; docs/stack-decision.md for why this is Python + stdlib with no framework.

Contributing

See CONTRIBUTING.md. One thing worth stating up front: this is a dogfooded tool, so its roadmap is driven by its authors' own use until there is a second real user. A feature request grounded in your actual use is welcome and is exactly what moves that line.

Licence

AGPL-3.0. See LICENSE and NOTICE.

About

Nous Ergon — a read-only fleet index that persists nothing and reports what it cannot see. Typed entity catalog (component · run · cycle · artifact · signal · decision · incident) over your existing durable artifacts, with one adapter per source.

Topics

Resources

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages