Computer Science at the University of Southern California · Los Angeles
I work where machine learning meets systems engineering: smaller models, structured outputs, reproducible evaluations, and the infrastructure required to make all of that useful. I care about the measurements after accuracy too—calibration, latency, throughput, and cost.
Right now, I am exploring how non-autoregressive models can make fast, typed decisions over text, and how those systems compare with much larger language models on real workloads.
| Project | What I am investigating |
|---|---|
| Laya | A 421M-parameter, non-autoregressive decision engine for typed choice, score, and yes/no questions over text in 100+ languages. |
| Jev · Laya · Qwen classification study | A reproducible comparison on 12,000 U.S. federal IT solicitations, evaluated against what a reseller actually quoted. |
| Laya CUDA benchmark | Throughput, p99 latency, and cost across laptop, workstation, datacenter, and MIG GPUs—with optimized runtimes and LLM/API baselines. |
| GLiNER2 | Unified, schema-driven information extraction from unstructured text. |
research question → reproducible run → honest metric → engineering decision
- Evaluate on operational data, not just convenient datasets.
- Measure confidence and calibration, not only top-line accuracy.
- Treat latency, throughput, and cost as model-quality dimensions.
- Publish the method, limitations, and artifacts needed to reproduce a result.
Python · PyTorch · CUDA · Transformers · vLLM · TensorRT · ONNX Runtime · FastAPI · Docker · PostgreSQL

