Multi-tenancy, scale, and cost for enterprise AI. Separate shared platform capabilities from tenant-specific data, policy, and execution context. Every request carries authenticated tenant/domain identity — retrieval filters, memory, tools, quotas, and audit inherit it. Logical isolation for lower risk; physical isolation when required. Scale with stateless tiers, durable workflow state, async queues, backpressure, caching, and circuit-breaker fallbacks. Cost: route by complexity, smaller models for simple work, bounded context/iterations, token & tool budgets, and cost per successful business transaction.
Complements finops-platform-landing-zone
(cloud landing-zone FinOps) and enterprise-ai-platform-planes
(model gateway). This repo owns tenant identity + AI unit economics on the request path.
Shared platform (API · gateway · tools · obs)
│ authenticate tenant + principal
├─ Tenant A LOGICAL — shared index + filters
├─ Tenant B PHYSICAL — dedicated partition
└─ Budgets · RPM · cache · circuit · async queue
↓
FinOps: $ / successful business transaction
python -m venv .venv && source .venv/bin/activate
pip install pytest
pytest -q
./scripts/demo.shpip install fastapi uvicorn pydantic
PYTHONPATH=src uvicorn mt_ai.api:app --port 8093
# POST /v1/execute -H "X-Tenant-Id: tenant-retail-a" -H "X-Roles: analyst"
# GET /v1/finops/tenant-retail-aMulti-tenancy · Logical/Physical isolation · Rate limiting · Circuit breakers · Async queues ·
Embedding cache · Model routing · Token/tool budgets · FinOps unit economics · FastAPI