《高级机器学习》:面向生成式人工智能时代的机器学习教材,涵盖强化学习、大模型训练与对齐、生成建模和多模态学习等内容。
-
Updated
Sep 28, 2026 - TeX
《高级机器学习》:面向生成式人工智能时代的机器学习教材,涵盖强化学习、大模型训练与对齐、生成建模和多模态学习等内容。
The implementation for our paper: TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning.
Based on the virtual world built with IsaacSim, accomplish the training and deployment of real-world reinforcement learning for Hil-Serl.
基于ms-swift的Qwen3_8B 金融推理两阶段后训练(LoRA SFT->GRPO)。
Interactive Autonomous Navigation Research Lab for UGVs featuring localization, path planning, DWA, MPC, Adaptive MPC, reinforcement learning hooks, safety supervision, live visualization, benchmarking, replay, and automated report generation.
To associate your repository with the reforcement-learning topic, visit your repo's landing page and select "manage topics."