This companion Python service makes the router's opt-in hmm strategy
self-hostable. It serves immutable PCA/HMM/XGBoost artifacts through the
router's policy_router_v3 HTTP contract. Training, online learning,
WorkWeave registries, GCS, and managed-service credentials are intentionally
outside this package.
Add a Google API key to .env.local, then start the optional profile:
echo 'GOOGLE_API_KEY=...' >> .env.local
make up-hmmThe router still defaults to cluster. Select HMM through an installation
strategy or an authorized x-weave-router-strategy: hmm override. To stop the
sidecar and return to the normal stack, run make down-hmm && make up.
The Compose configuration downloads the public
hmm-model-v1
asset and verifies SHA-256 before extraction. Set HMM_PACKAGE_PATH instead
when running the sidecar directly with a local package.
HMM_PACKAGE_URL / HMM_PACKAGE_PATH (+ HMM_PACKAGE_SHA256) name the
default package. HMM_PACKAGE_REGISTRY optionally adds a bounded set of
further packages as a JSON list of {"sha256": ..., "url": ...} entries
(path may replace url for a pre-staged archive; at most 8 entries; https
only). Every entry is downloaded, digest-checked, extracted, manifest-verified
and loaded during startup — a failure in any entry fails readiness, and nothing
is ever fetched on a live request.
HMM_PACKAGE_REGISTRY='[{"sha256":"<sha256>","url":"https://.../hmm-model-v0.tar.gz"}]'/route, /preview (request field artifact_sha256) and /roster (query
?artifact_sha256=) select a loaded package by its outer digest; omitting it
serves the default. An unknown digest answers 404 rather than the default
package, so a router honouring an x-weave-policy-pin header fails closed.
/readyz lists every loaded digest under policy_artifact_sha256s.
The v1 artifact requires google/gemini-embedding-2 with 3,072 output
dimensions. The embedding is not a replaceable preprocessing detail: it is the
majority of the classifier feature vector and defines the HMM's emission space.
The sidecar therefore verifies a reference-vector probe at startup.
An OpenAI-compatible /embeddings endpoint can be used by setting:
HMM_EMBEDDING_PROVIDER=openai-compatible
HMM_EMBEDDING_BASE_URL=https://embedding-proxy.example/v1
HMM_EMBEDDING_API_KEY=...
HMM_EMBEDDING_MODEL=google/gemini-embedding-2It must expose the same underlying Google embedding model and pass the probe. Ollama, vLLM, or another local embedder requires an HMM/classifier package trained in that model's own vector space.
/route answers with a classification. The arm is chosen by the router, from
the declarative roster file (ROUTER_HMM_ROSTER_PATH, required alongside
ROUTER_HMM_SIDECAR_URL) using internal/router/hmm/selection:
predicted_label— the classifier's predicted complexity label;class_probabilities— the classifier's per-class probability map;ranked_fallback— the cluster ranking (probability, roster arms, eligible arms per group) the router walks;selected_roster_id,selected_provider,model— alwaysnull.
policy_router_v3 is a hard break from v1/v2 with no compatibility window:
this sidecar rejects a v1/v2 /route request, and a v3 router rejects a
response that names an arm or omits ranked_fallback — the turn fails with
HTTP 503 rather than serving a sidecar pick. Run the sidecar and the router
from matching releases. See
docs/HMM_GO_SELECTION.md.
The runtime accepts only hmm_router_frozen_package_v1 archives. It verifies:
- the outer package digest;
- safe archive paths, member count, and size limits;
- every file digest listed by the manifest;
- classifier/HMM dimensions and class order;
- the live embedding endpoint against the stored probe.
The public package contains data arrays and XGBoost JSON only. The legacy training pickle, training rows, prompts, embedding cache, local filesystem paths, and WorkWeave registry records are excluded.