Kuaishou Technology — Research Intern
Aug 2025 – Jul 2026
Agent post-training for tool use and decision making. Led the AgentBrew training pipeline and the reasoning module of LBM, working with Dr. Qingpeng Cai and Dr. Yewen Li.
I am a Ph.D. student at the College of Computing and Data Science, Nanyang Technological University, advised by Prof. Bo An. I received my B.S. in Computer Science from Xidian University.
I work on post-training language-model agents. Most of my recent work starts from the same practical constraint: real agent logs are plentiful, but online rollouts, reward models, and task verifiers usually are not. Concretely, I am interested in
I am always happy to talk about potential collaborations — feel free to reach out by email.
* denotes equal contribution. See also my Google Scholar.
AgentBrew post-trains tool-use agents from raw interaction logs with no online rollouts, external rewards, or task verifiers. Retrospective task inference plus PMI-based step-level credit assignment turns noisy trajectories into weighted SFT supervision.
Separating high-level reasoning from precise numerical action lets an LLM agent operate in continuous decision spaces; the reasoning module is trained with offline RL, avoiding simulator and real-world rollouts.
Code generation as anytime local search over programs. A revision reward model ranks candidates by revision distance, guiding both hill climbing and genetic search on LiveCodeBench and TACO. In collaboration with Alibaba Tongyi Lab.
CoSo measures each token's causal influence on the executable action and concentrates exploration on action-critical tokens, with convergence and policy-improvement guarantees.
Q* casts multi-step reasoning as heuristic search and learns a plug-and-play Q-value model to score the next reasoning step, guiding decoding without fine-tuning the LLM for the task. Evaluated on GSM8K, MATH, and MBPP. With Skywork AI.
A variational structured memory that stores semantic concepts alongside sample-specific episodic memories, with structured recall for rapid adaptation to unseen generation tasks.
Aug 2025 – Jul 2026
Agent post-training for tool use and decision making. Led the AgentBrew training pipeline and the reasoning module of LBM, working with Dr. Qingpeng Cai and Dr. Yewen Li.
Singapore
College of Computing and Data Science. Advisor: Prof. Bo An.
2020 – 2024
GPA 3.9 / 4.0.