Skip to content
View bhushankinge's full-sized avatar

Block or report bhushankinge

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
bhushankinge/README.md
Bhushan Kinge — AI systems builder. A visual pipeline turns unstructured text into typed, measured decisions.
AI systems · model evaluation · efficient inference
Computer Science at the University of Southern California · Los Angeles

I build decision systems for messy, real-world data.

I work where machine learning meets systems engineering: smaller models, structured outputs, reproducible evaluations, and the infrastructure required to make all of that useful. I care about the measurements after accuracy too—calibration, latency, throughput, and cost.

Right now, I am exploring how non-autoregressive models can make fast, typed decisions over text, and how those systems compare with much larger language models on real workloads.

Selected work

Project What I am investigating
Laya A 421M-parameter, non-autoregressive decision engine for typed choice, score, and yes/no questions over text in 100+ languages.
Jev · Laya · Qwen classification study A reproducible comparison on 12,000 U.S. federal IT solicitations, evaluated against what a reseller actually quoted.
Laya CUDA benchmark Throughput, p99 latency, and cost across laptop, workstation, datacenter, and MIG GPUs—with optimized runtimes and LLM/API baselines.
GLiNER2 Unified, schema-driven information extraction from unstructured text.

How I approach the work

research question → reproducible run → honest metric → engineering decision
  • Evaluate on operational data, not just convenient datasets.
  • Measure confidence and calibration, not only top-line accuracy.
  • Treat latency, throughput, and cost as model-quality dimensions.
  • Publish the method, limitations, and artifacts needed to reproduce a result.

Working with

Python · PyTorch · CUDA · Transformers · vLLM · TensorRT · ONNX Runtime · FastAPI · Docker · PostgreSQL


Interested in efficient AI systems, evaluation, or information extraction? Explore the projects above or browse my repositories.

Pinned Loading

  1. laya-cuda-bench laya-cuda-bench Public

    How many decisions/s can one NVIDIA GPU serve under a p99 SLO, and what does a million cost? Reproducible benchmark of the Laya 421M decision models on RTX 2000 Ada, RTX PRO 5000/6000 Blackwell and…

    Jupyter Notebook

  2. jev-laya-classification-bench jev-laya-classification-bench Public

    Typed-decision models (Jev API, Laya 421M) vs Qwen3.5-35B on 12,000 real U.S. federal IT solicitations, graded against actual reseller quotes: accuracy, calibration, Wilson-bounded auto-accept cuto…

    Python