header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

# Hugging Face让Claude Code、Codex直接参与RL训练 Hugging Face 最近推出了一项新功能,让 Claude Code、Codex 等 AI 编程助手能够直接参与到强化学习(RL)训练流程中。这一举措旨在打通"AI 写代码"与"AI 训练 AI"之间的壁垒,让编程智能体成为 RL 训练闭环中的一环。 ## 核心思路 传统 RL 训练流程中,环境搭建、奖励函数设计、训练脚本调试等环节往往需要人类工程师反复介入。Hugging Face 的新方案允许 Claude Code、Codex 这类具备代码生成与执行能力的智能体直接接入训练环境,承担以下角色: - **自动编写和修改训练代码**:根据任务目标生成 RL 训练脚本 - **调试训练过程**:读取日志、定位报错、迭代修复 - **设计奖励函数**:根据环境反馈调整 reward shaping - **参与策略迭代**:在训练循环中动态调整超参数或环境配置 ## 技术实现 这一功能依托于 Hugging Face 现有的生态组件,可能涉及: - **TRL(Transformer R

动察 Beating AI News Flash: Hugging Face has open-sourced a Multi-harness RL solution that can directly plug existing Coding Agents such as Claude Code, Codex, and OpenCode into the reinforcement learning process to train open-source models connected to them, without needing to build a separate training environment for each Agent.


Large model companies such as OpenAI have long been training models in their own Agent environments. This time, Hugging Face has turned similar capabilities into a general open-source tool, allowing other teams to directly use existing Agents for RL.


Moreover, the model does not have to be trained in only one type of Agent. In a single training run, tasks can be run in turn across Claude Code, Codex, OpenCode, and Mini-SWE-Agent, enabling the same model to adapt to different prompts, tools, and execution logic, reducing the "imbalance" of only being good at one kind of Harness.


They tested with a 2.6-billion-parameter small model from Liquid AI. With the same weights and only a change of Agent, the task pass rate could drop from about 62% to 33%. After training simultaneously in four Agent environments, the average pass rate increased from 42.2% to 54.2%, and the improvement across different Agents was also more balanced than training in only a single environment.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish