CritPT-RL
面向 scientific coding tasks 的完整 RL post-training 流程,覆盖数据生成、reward 设计、GRPO 训练、checkpoint evaluation 和 official-style evaluation。
A complete RL post-training pipeline for scientific coding tasks, covering data generation, reward design, GRPO training, checkpoint evaluation, and official-style evaluation.
查看项目 View project
