Skip to content
View Schofi's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report Schofi

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Schofi/README.md

👋 Hi, I'm @Schofi

I'm interested in building trainable and reliable AI agents, with a focus on LLM post-training, Agentic RL, agent infrastructure, and industrial AI systems.

🔬 Currently exploring

  • LLM Post-Training

    • SFT, Preference Optimization, RLHF / RLVR
    • PPO, GRPO and reinforcement learning for tool-using agents
    • Data synthesis, trajectory construction and bad-case mining
  • Agentic RL & Agent Training

    • Multi-turn and long-horizon agent training
    • Tool-use and environment interaction
    • Trajectory collection and credit assignment
    • Verifiers, reward design and executable evaluation environments
  • Agent Infrastructure

    • Agent Harness / Runtime
    • Tool execution, sandboxing and state management
    • Rollout systems and training-serving integration
    • Evals, observability and failure recovery
  • Industrial AI

    • AI agents for CAD / CAE / engineering software
    • GUI + API hybrid agents
    • Verifiable workflows for engineering tasks
    • Turning industrial software into trainable agent environments

🛠 Tech Stack

Languages Python · C++ · Go

LLM / Training PyTorch · Transformers · verl · DeepSpeed · Megatron-LM

Inference vLLM · SGLang

Systems CUDA · Docker · Kubernetes · Linux

🧠 Research Interests

Agentic RL · Post-Training · RL Infrastructure · Agent Harness
Long-Horizon Agents · Verifiable Rewards · Industrial Agents


🚀 My current goal is to understand and build the complete agent learning loop:

Environment → Rollout → Trajectory → Verifier → Reward → Post-Training → Evaluation → Deployment

and explore how this loop can be applied to real-world industrial software and engineering workflows.

Popular repositories Loading

  1. PageIndex PageIndex Public

    📄 页索引:基于推理的 RAG 文档索引系统

    Python 4 2

  2. gex gex Public

    Forked from ikun2021/gex

    go微服务实践-数字货币交易平台

    Go 1

  3. play-with-llm-application play-with-llm-application Public

    Jupyter Notebook 1

  4. Happy-With-Data-Structure Happy-With-Data-Structure Public

    学习数据结构的一些代码,包括数组,链表,二分搜索树,集合,并查集,线段树,AVL树,红黑树,哈希表,如有错误,多多指教。

    Java 1

  5. minikeyvalue minikeyvalue Public

    Forked from geohot/minikeyvalue

    A distributed key value store in under 1000 lines. Used in production at comma.ai

    Go

  6. go-fastdfs go-fastdfs Public

    Forked from sjqzhang/go-fastdfs

    go-fastdfs 是一个简单的分布式文件系统(私有云存储),具有无中心、高性能,高可靠,免维护等优点,支持断点续传,分块上传,小文件合并,自动同步,自动修复。Go-fastdfs is a simple distributed file system (private cloud storage), with no center, high performance, high relia…

    Go