I build AI agents that reach production and stay there. Full-stack since 2019, AI-first since 2021. Model to interface, and the infrastructure holding it up.
When it is wrong, it can still show you exactly what it saw and why it concluded that. I build agents that are fast, accurate, and accountable when they are none of the above.
"Our agent is right, eventually, and nobody will wait that long." Speed in an agent is not a faster model, it is refusing to do the same work twice. I make the expensive parts of a run reusable: identical questions collapse into one computation instead of one per user, a step that already fetched something is never allowed to fetch it again, and the queries an agent writes are made deterministic so the same request produces the same result rather than a fresh guess. Users get an answer while they still care about the question.
"It gave us the wrong number and we cannot find out why." This is the failure that ends trust in an AI product, and it is an architecture problem, not a model problem. I build agents where the evidence outlives the answer: whatever a step concludes, the full data it actually saw is preserved alongside it rather than being summarised away. So when an answer is wrong, you can open it, see the real inputs, and say precisely where it went wrong. Every model call is traced end to end. Nothing has to be reproduced from memory.
"Our team is drowning in a manual process." Contract review, ad account management, campaign reporting, recruitment screening. I find the judgement-heavy work that people should keep, automate the mechanical work around it, and leave a human in the loop at the point where it actually matters.
"We need this built, and there is nobody to hand the other half to." I ship the API, the interface, the data pipeline, and the deployment. Small teams get a working product instead of a component that needs three more hires to become one.
Separately, in published research: model compression and alignment work, including a 91.25% reduction in KV-cache memory on an 8B model that improved long-context benchmark scores.
Founding AI Engineer at a stealth-stage startup, building an autonomous advertising platform. Unlaunched, so the details stay light.
Growth teams waste budget because nobody can watch every campaign every hour. The platform does. I built the system that ingests a brand's advertising, analytics, and creative data, monitors performance continuously, and says what changed and what to do about it in plain language, so a marketer without an analyst can act on their own numbers.
What that meant in practice:
- An orchestration layer running eleven independent background workers with resilient scheduling, so one slow integration never stalls the rest of the platform.
- A conversational analytics agent that turns a plain-English question into a warehouse query and answers over live campaign data, replacing dashboard archaeology.
- A caching layer that took multi-second page loads off the critical path. Concurrent requests for the same view collapse into a single computation, results stay servable while they refresh in the background, and any user action that changes the data invalidates it immediately. Fast and stale is a bug; this is fast and correct.
- Answers you can audit. Tool results are preserved in full rather than being retyped and truncated by each model that handles them. The agent works from a bounded view, the complete evidence travels to the user, and a wrong answer can always be traced back to the exact rows that produced it.
- A creative intelligence pipeline that reads the actual images and video in an ad account and connects creative choices to performance.
- A tested codebase, not a prototype: over 200 test files across the services, tracing on every model call, and monitoring that surfaces failures before customers report them.
Legal contract review platform. Lawyers reviewing agreements clause by clause. The hard part was not summarisation, it was making sure a retrieved clause arrived whole, because half an obligation reads as a different obligation entirely. Built the document pipeline, the clause-level risk scoring, and the review interface. → the technique, open-sourced
Voice and chat agent platform for service businesses. Missed calls are lost revenue for contractors. Built a configurable agent that answers, qualifies, and books appointments over real-time voice or chat, deployed for paying customers across multiple verticals from one codebase.
Marketing research and proposal platform. Trend ingestion through to a finished client-ready proposal, built with a team as a monorepo.
Model research, published openly. 15+ models on Hugging Face reproducing frontier techniques end to end, with evaluations rather than claims: 91.25% KV-cache compression with benchmark gains, 76.1% reward accuracy on preference alignment over a dataset I built myself, and 82.29% on TextVQA for a vision-language model.
| MCP-Server-AlphaVantage ⭐ 7 | Gives any assistant real market analysis tools. Independently security-audited by MseeP.ai. |
| RepoMind | Ask a codebase questions. Reviews its own answer and retries rather than guessing. |
| HF2Reasoning ⭐ 2 | Turns any public dataset into a reasoning dataset, so small teams can build training data without a labelling budget. |
| iRoPE Implementation | Long-context attention rebuilt from a paper description, with an honest account of the limits. |
| SafeTensors Converter | Converts model checkpoints out of a format that executes code on load. Verifies every tensor. Also a hosted app. |
| AI Recruitment Synapse | Matches candidates to roles and explains why, instead of returning an unexplained score. |
| SynGen | Builds legal question-answer datasets for a jurisdiction that has almost none. |
Alongside these: agent orchestration and evaluation frameworks, vector and graph databases, hybrid retrieval, distributed task queues, LLM observability tooling, and load testing.
| 2025 – Present | Founding AI Engineer | Stealth startup · autonomous advertising |
| 2024 – 2025 | Full-Stack Engineer | Yeild AI · marketing research and proposals |
| 2024 | Senior AI Engineer | BotsCrew · conversational agent platforms |
| 2022 – 2024 | Founding AI Engineer | 108 AI · legal AI agents |
| 2021 – 2022 | AI Research Engineer | Queryloop AI · applied LLM research |
| 2019 – 2021 | Software Engineer, Full-Stack | Devsinc · started as an intern, stayed two years |
B.E. Software Engineering · National University of Sciences and Technology, Islamabad · GPA 3.81
Most of my work lives in private repositories, so the graph here is quiet. This page is the public trace of it.
Open to remote AI engineering and AI-first full-stack roles.

