I design, build, and operate production AI systems end to end: the backend, the frontend, and the infrastructure they run on. I turn requirements into measurable targets, then move the numbers.
- Now — AI full-stack developer at Grinda AI, leading AI and full-stack development.
- Own — a sales email automation SaaS and a public-sector AI evaluation platform.
- Delivered — AI services to banks, public institutions, and an overseas securities firm.
- Comfortable with — closed networks, on-premise GPU clusters, and air-gapped deployment.
| 11,447 commits across 88 repositories | 2,318 pull requests in 5 main repositories |
| LLM p95 latency −70%, serving cost −38% | Sidebar 26×, list 125×, one query 411ms → 1.6ms |
| First-load bundle 21.7MB → 2.3MB | Icon webfont 3.74MB → 379KB |
| Three LLM/embedding models re-homed to on-prem DGX, zero downtime | A 7.83% bounce incident contained, then a sending governor built |
Context caching plus structured output cut p95 latency 70%, wall time 31%, and cost 38%; dropping HTML body generation removed another 49% of output tokens. A materialized summary table with pg_trgm GIN indexes made the app usable at scale. I owned deliverability: after a 7.83% bounce event I restored fail-closed validation and built a multi-domain sending governor with a database role pool, rate limiting, and an explicit fail-open policy. Billing runs on a hexagonal port with two providers behind it.
Five capabilities shipped: evaluation-criteria structuring, proposal and product-spec extraction, cross-compliance analysis, fact verification, and semantic search. I wrote the HWPX/HWP parser with magic-byte-first format detection, recovered text lost in shape sub-lists and table cells, and asserted word-level coverage against ground truth over a 36-file corpus. When the cloud GPU farm shut down I re-homed the language, vision, and embedding models to an on-premise DGX with zero downtime, which is what makes air-gapped public-sector bids possible. Passed a national AI verification technical review, on-site demonstration included.
| Tool | What it is | Built with |
|---|---|---|
| Nine Agents | VS Code fork with a 3×3 agent grid, preset rows, and broadcast input | TypeScript |
| zed2 | Zed fork tuned for multi-agent editing | Rust |
| rtk / ctx | Token-reduction toolchain that shrinks agent tool output | Rust |
| security-checker | Supply-chain scanner with OSV diffing, signed and notarized | Rust · Tauri 2 |
| Claude Code skills | ~90 skills covering deploys, databases, infrastructure, and reporting | TypeScript · Python |
| discord-claude-bridge | Agent chat bridge with a token ledger and concurrent request handling | Rust |
The page keeps one accent and one secondary so the badges read as a single system rather than a pile of logos.
| Token | Value | Where it is used |
|---|---|---|
--canvas |
#0B1220 |
Every badge label, banner background |
--accent |
#0F766E |
Languages, role badges, the rule under the title |
--accent-cyan |
#22D3EE |
AI and LLM group, the one highlight line in the banner |
--accent-2 |
#F59E0B / #B45309 |
Numbers that changed, infrastructure group |
--ink / --ink-2 |
#E6EDF3 / #9FB3C8 |
Banner title and supporting copy |
AI 제품을 기획 단계부터 운영까지 책임지고 개발합니다. 백엔드와 프론트엔드, 인프라를 모두 직접 다룹니다.
주식회사 그린다에이아이에서 AI 풀스택 개발자로 일합니다. 영업 이메일 자동화 SaaS와 공공 부문 AI 평가 시스템을 맡고 있습니다. 은행과 공공기관, 해외 증권사에 AI 서비스를 납품했습니다.
요구사항을 측정 가능한 지표로 바꾸어 놓고, 그 지표가 실제로 움직이는지 확인하면서 개선합니다. LLM 응답 지연 p95를 70% 단축하고 서빙 비용을 38% 절감했습니다. 폐쇄망 과제에서는 언어 모델과 임베딩 모델 세 종을 온프레미스 DGX로 무중단 이관했습니다. 업무에 쓰는 도구도 직접 만들어 씁니다.





