Sanitized enterprise MLOps reference blueprint spanning governed features, model lifecycle, GPU serving, drift response, and serving economics.
-
Updated
Jul 28, 2026 - Python
Sanitized enterprise MLOps reference blueprint spanning governed features, model lifecycle, GPU serving, drift response, and serving economics.
ML model registry and serving — versioning, A/B testing, drift detection, model deployment, MinIO artifact storage
Distributed model inference engine with REST/gRPC serving, circuit breaker, load balancing and A/B testing
Scalable MNIST image classification API with FastAPI and ONNX, deployed on Kubernetes with HPA autoscaling
Production-style model inference platform — versioned registry, dynamic micro-batching (measured 8x throughput), canary + shadow deployments with instant rollback, load shedding (429+Retry-After), Prometheus metrics, async load-test harness
End-to-end MLOps: train, serve, Dockerize, and CI, with a model-quality gate on every push.
Production-grade ML inference service: multitask PyTorch served via FastAPI with dynamic batching, Prometheus metrics, 38 tests, Locust load tests, and multi-stage Docker (4GB → 800MB).
PyTorch model training and inference service exposed through a FastAPI REST API.
An OpenMOSS fork of SGLang with serving support for the MOSS-TTS family and MOSS-Audio
ML infrastructure on Kubernetes with model serving, Helm, Argo CD, MLflow flow, Prometheus metrics, and drift checks.
End-to-end MLOps pipeline for training, tracking, serving, and monitoring a Titanic survival prediction model.
Distributed fraud-scoring service with Go inference, model versioning, and a hand-written Raft consensus core.
Useful field notes on building AI systems that survive contact with reality.
Building Real-Time Inference Pipelines with Ray Serve
Demo ML repository, using Algorithmia Model Deployment Github Action, to auto deploy on an Algorithmia algorithm backed by Github
A Lazy, high throughput and blazing fast structured text generation backend.
Customize Nvidia Triton to use OpenShift Source to Image building
Text-to-Speech (TTS) service that fetches Wikipedia articles, cleans the text (removes citation brackets, edit tags, etc), splits content by section, and generates .mp3 files asynchronously.
To associate your repository with the model-serving topic, visit your repo's landing page and select "manage topics."