Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

multitenant-ai-scale-finops

CI

Multi-tenancy, scale, and cost for enterprise AI. Separate shared platform capabilities from tenant-specific data, policy, and execution context. Every request carries authenticated tenant/domain identity — retrieval filters, memory, tools, quotas, and audit inherit it. Logical isolation for lower risk; physical isolation when required. Scale with stateless tiers, durable workflow state, async queues, backpressure, caching, and circuit-breaker fallbacks. Cost: route by complexity, smaller models for simple work, bounded context/iterations, token & tool budgets, and cost per successful business transaction.

Complements finops-platform-landing-zone (cloud landing-zone FinOps) and enterprise-ai-platform-planes (model gateway). This repo owns tenant identity + AI unit economics on the request path.

Architecture

Shared platform (API · gateway · tools · obs)
        │  authenticate tenant + principal
        ├─ Tenant A LOGICAL  — shared index + filters
        ├─ Tenant B PHYSICAL — dedicated partition
        └─ Budgets · RPM · cache · circuit · async queue
                 ↓
        FinOps: $ / successful business transaction

Run

python -m venv .venv && source .venv/bin/activate
pip install pytest
pytest -q
./scripts/demo.sh
pip install fastapi uvicorn pydantic
PYTHONPATH=src uvicorn mt_ai.api:app --port 8093
# POST /v1/execute  -H "X-Tenant-Id: tenant-retail-a" -H "X-Roles: analyst"
# GET  /v1/finops/tenant-retail-a

Documentation

Toolbox

Multi-tenancy · Logical/Physical isolation · Rate limiting · Circuit breakers · Async queues · Embedding cache · Model routing · Token/tool budgets · FinOps unit economics · FastAPI

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages