OpenRLHF
View on GitHubAn Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
OpenRLHF is a Ray + vLLM distributed RLHF framework with an agent-based execution pipeline, supporting PPO, GRPO, REINFORCE++, DAPO, async RL, and VLM/multi-turn agent training for LLMs up to 70B+ parameters.
Use Cases
Train LLMs with RLHF/PPO/GRPO at scaleReinforcement fine-tuning of reasoning modelsMulti-turn agent RL in external environmentsVision-language model RLHF with image inputsCustom reward function training for single-turn agentsAsync RL training pipelinesDistributed RLHF on multi-GPU clusters up to 70B+ paramsReproducing DeepSeek-R1-style reasoning trainingSupervised fine-tuning and reward model trainingLoRA-based parameter-efficient RLHF
Built With
- Language
- Python
- Frameworks
- Ray · vLLM · DeepSpeed · HuggingFace Transformers · PyTorch · Accelerate · FlashAttention · bitsandbytes · PEFT · NCCL
Tags
RLHF · reinforcement-learning · PPO · GRPO · REINFORCE++ · agentic-RL · distributed-training · Ray · vLLM · DeepSpeed · multi-turn-agents · VLM · async-RL · LoRA · reward-model · fine-tuning