I'm interested in building trainable and reliable AI agents, with a focus on LLM post-training, Agentic RL, agent infrastructure, and industrial AI systems.
-
LLM Post-Training
- SFT, Preference Optimization, RLHF / RLVR
- PPO, GRPO and reinforcement learning for tool-using agents
- Data synthesis, trajectory construction and bad-case mining
-
Agentic RL & Agent Training
- Multi-turn and long-horizon agent training
- Tool-use and environment interaction
- Trajectory collection and credit assignment
- Verifiers, reward design and executable evaluation environments
-
Agent Infrastructure
- Agent Harness / Runtime
- Tool execution, sandboxing and state management
- Rollout systems and training-serving integration
- Evals, observability and failure recovery
-
Industrial AI
- AI agents for CAD / CAE / engineering software
- GUI + API hybrid agents
- Verifiable workflows for engineering tasks
- Turning industrial software into trainable agent environments
Languages
Python · C++ · Go
LLM / Training
PyTorch · Transformers · verl · DeepSpeed · Megatron-LM
Inference
vLLM · SGLang
Systems
CUDA · Docker · Kubernetes · Linux
Agentic RL · Post-Training · RL Infrastructure · Agent Harness
Long-Horizon Agents · Verifiable Rewards · Industrial Agents
🚀 My current goal is to understand and build the complete agent learning loop:
Environment → Rollout → Trajectory → Verifier → Reward → Post-Training → Evaluation → Deployment
and explore how this loop can be applied to real-world industrial software and engineering workflows.

