2025.12 — NOW
北京大学前沿计算研究中心
本科生科研助理 · 导师:董豪
参与具身模型的 post-training 与评测。手上的工作比较具体:多源机器人数据清洗和对齐、SFT / CoT 样本构造、RL 候选筛选、评测脚本维护,以及错误案例回流。
这段工作让我真正意识到,模型“失败”本身并不自动成为有用数据。只有把输入、轨迹、判分和目标都记录清楚,失败才有可能进入下一轮训练。
Center for Frontier Computing Research, PKU
Undergraduate Research Assistant · Advisor: Hao Dong
I work on post-training and evaluation for embodied models: cleaning and aligning robot data, building SFT and CoT samples, filtering RL candidates, maintaining evaluation scripts, and feeding failure cases into the next round.
The work made one thing clear: a model failure does not automatically become useful data. Inputs, trajectories, scoring, and targets all need to be recorded well enough for the failure to teach the next run anything.