reward-model 2 MLRL004: RLHF / Agent Training Jul 1, 2026 HKUDS011: DeepInnovator 作为 Scientific Idea Foundation Model 与 Research Innovation Training Layer Jun 25, 2026