A high-throughput and memory-efficient inference and serving engine for LLMs
-
Updated
Oct 7, 2026 - Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Stop guessing GPU VRAM. CLI that sizes LLM weights + KV cache and picks the cheapest AWS GPU for vLLM, DeepSeek-R1, Llama 70B, and Qwen. JSON output for AI agents.
A fast CPU-based API for Qwen 2.5 using CTranslate2, hosted on Hugging Face Spaces.
⚡ Unified real-time forecasting platform for crypto, equities, energy & food commodities — probabilistic LightGBM (80% CI), seasonal-aware early warnings, live accuracy audit, Groq AI market analyst, Streamlit dashboard, FastAPI & Telegram alerts.
🤘 TT-NN operator library, and TT-Metalium low level kernel programming model.
Daily AI intelligence brief for founders. Python + Groq LLaMA 3.3-70B + GitHub Actions + cron-job.org. Delivers 90-second AI briefing to Telegram every morning.
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.8, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM5.3, Gemma4, Llava, Phi4, ...) (AAAI 2025).
Efficient Triton Kernels for LLM Training
Twice-daily AI-summarized news briefing to Telegram. Free forever. Python + Groq LLaMA 3.3-70B + GitHub Actions.
Build and manage modular AI agents with multi-model support, document RAG, memory, and tool integration using FastAPI and SQLite.
Access a curated list of over 1600 free and freemium public APIs across 51 categories for software development projects.
🚀 Pytorch Distributed native training library for LLMs/VLMs with OOTB Hugging Face support
Automate Google AI Studio with a thread-safe Python bridge using Playwright to programmatically control Gemini models and extract text from files.
Analyze code snippets with an AI assistant to find design flaws, suggest refactoring, generate tests, and audit SOLID principles efficiently.
Automate binary analysis by coordinating LLM agents with Ghidra, enabling scalable and precise reverse engineering workflows.
To associate your repository with the llama topic, visit your repo's landing page and select "manage topics."