Catch vendor lock-in before it catches you.
-
Updated
Jul 28, 2026 - Python
A large language model (LLM) is a type of machine learning model designed for understanding, generating, and interacting with human language. These models are trained on extensive datasets containing text from books, articles, websites, and other sources to learn patterns, context, and semantics in language. LLMs are widely used in applications like chatbots, code generation, translation, summarization, and more. They are often built using transformer architectures and are central to the field of generative AI.
Catch vendor lock-in before it catches you.
Hands-on labs for isolated AI workloads with Azure Container Apps Dynamic Sessions and ACA Sandboxes
RAG system for healthcare billing Q&A — RAGAS-evaluated (1.0 faithfulness, 0% hallucination), deployed live
Fine-tuned 3B model + RAG beats Claude Opus 4.8 raw at entity resolution (99% vs 78.5%) on unseen companies — Covent LLM Challenge submission.
Ask questions about your PDFs and get cited answers — or an honest "not found." Built from scratch in Python.
终端 AI Agent CLI — ReAct/Plan-and-Execute/Multi-Agent/MCP/RAG
AI Contract Intelligence Platform — RAG-powered natural-language Q&A, automated clause-risk flagging, and obligation-deadline tracking over enterprise contracts. FastAPI + React + PostgreSQL + ChromaDB + Ollama (Mistral 7B), fully local and free via Docker Compose.
Multi-provider LLM usage, cost & rate-limit meter with budget guards — Anthropic/OpenAI/OpenRouter/local, TUI + MCP + CI exit codes. Zero-dependency.
Boilerplate for building local LLM apps with Ollama
An autonomous, Multi-Agent AI system for Anti-Money Laundering (AML) that fuses deterministic structuring rules with unsupervised ML (PyOD) to dynamically detect financial crime and auto-generate explainable SARs.
PRISM: Pre/post-inference Runtime Inference Safety Monitor
Local-LLM RAG chat assistant for internal documentation — hybrid retrieval, citation-backed answers, zero cloud LLM calls. FastAPI + React, Docker Compose (Mac/GPU/AWS) + k8s stubs.
Distributed GPT-2 Large pretraining with PyTorch DDP across multi-node H100s — throughput, NCCL communication overhead, and scaling efficiency analysis on Nebius AI Cloud.