I build production-shaped AI and backend systems — RAG pipelines, LLM fine-tuning, voice agents, and the FastAPI / Postgres services that hold them up. I work end to end, with tests, migrations, and eval harnesses rather than notebooks, and I go deep on how the models actually behave.
- Fine-tuning Qwen2.5-1.5B for schema-conditioned text-to-SQL with an execution-accuracy eval harness
- Building AI automation workflows at Humai (Dubai) — LangGraph agent loops, a Gemma QLoRA fine-tune, a Pipecat voice agent, an MCP tool gateway
- Writing up a decoder-only GPT built from scratch in PyTorch
| Project | ||
|---|---|---|
| GraphRAG Hierarchical Chat | Graph-structured retrieval over long, cross-referenced document sets | repo · spec |
| PneumoScan AI | Explainable pneumonia detection from chest X-rays — 98.42% test accuracy, Grad-CAM overlays | repo · spec |
| YouSentimentAI | End-to-end MLOps pipeline for YouTube comment sentiment — DVC, MLflow, staged registry | repo · spec |
| LexiQE AI | Private, on-device legal-document Q&A — RAG over FAISS + FLAN-T5, cited answers | repo · spec |
| Decoder-only GPT, from scratch | Transformer, tokenizer, and training loop in raw PyTorch — to learn the internals | notes |
| Premium Business Planner | AI-assisted investor-ready plans — Gemini for prose, deterministic Python for the numbers | repo · spec |
Full case studies — problem, approach, outcome — at yadidiah.vercel.app/#work
llm / genai RAG (FAISS · pgvector · Pinecone) · LoRA / QLoRA (Unsloth)
LangChain · LangGraph · DSPy · prompt contracts + eval harnesses
Hugging Face · Ollama · PyTorch (from-scratch transformers)
backend Python · FastAPI · PostgreSQL · SQLAlchemy · Alembic · Redis / RQ
outbox pattern · idempotency keys · structured logging
NestJS · TypeScript
ml / data TensorFlow · Keras · scikit-learn · XGBoost · LightGBM
MLflow · DVC · Pandas · NumPy · Apache Spark
voice ai LiveKit · Pipecat · Deepgram · Whisper · Silero VAD · dual-STT validation
cloud / ops GCP (Cloud Run · GKE · Vertex AI) · AWS (SageMaker · Lambda)
Docker · Kubernetes · GitHub Actions
On-site notes on how LLMs, transformers, and retrieval actually behave — written from first principles, corrected against implementation:
- Deep learning, stated precisely
- The transformer forward pass, token by token
- What actually happens at inference — generation loop, KV cache, GQA/MQA, serving
- Building RAG from the boundaries in
- Backend patterns for AI work — durable jobs, outbox, idempotency
Longer posts on Medium → @yadidiah.k
Three peer-reviewed papers (IEEE · Springer, 2023–2025) on machine learning for IoT security — threat detection, intrusion detection, and systematized log analysis. Undergraduate thesis: implemented and benchmarked five volumetric segmentation architectures (U-Net, V-Net, SA-Net, E1D3, HDC) on 3D MRI brain-tumor data.
Full citations on LinkedIn.
LinkedIn ·
Medium ·
Kaggle ·
HackerRank ·
LeetCode ·
Stack Overflow ·
yadidiah.wrk@gmail.com


