-
Red Hat
- Milan, Italy
Stars
A high-throughput and memory-efficient inference and serving engine for LLMs
Achieve state of the art inference performance with modern accelerators on Kubernetes
Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of …
Model Registry provides a single pane of glass for ML model developers to index and manage models, versions, and ML artifacts metadata. It fills a gap between model experimentation and production a…
TrustyAI Explainability Toolkit
Python bindings for TrustyAI's explainability library


