I build LLM-backed systems where the interesting problem isn't the prompt — it's the boundary around it: what the model is allowed to touch, how a bad call fails safely, and how to prove the thing works before calling it done.
Recent work: a hybrid RAG + Postgres support agent with tool-scoped data access instead of text-to-SQL, a multi-agent GCP cost/reliability advisor on Vertex AI, and a config-driven document-verification engine.
Currently building: Work Cube — a general-purpose "recompute and compare" engine for auditing third-party calculations, config-driven so a new use case is a template, not a rewrite.
Started on OCR/scripting work in 2024–2025 (see InfoCraft). Since mid-2026, focused on LLM systems — agent architecture, evaluation, and the boundary between what a model decides and what it's allowed to touch.
Support Copilot — Hybrid RAG + Postgres support agent. Tool-scoped data access (no text-to-SQL), evaluated against a 26-question golden set scoring tool selection and answer correctness separately.
Cloud Cost Advisor — Multi-agent cost/reliability advisor on Google ADK + Vertex AI. A planner delegating to specialist agents, not one agent with a flat tool list.
minsurf — Discrete minimal-surface toolkit: soap-film form-finding and TPMS generation, validated against exact closed-form solutions. 74 tests passing.
