Vibe Coding Discover

AI Frameworks

Online-RLHF

View on GitHub

A recipe for online RLHF and online iterative DPO.

★ 54648 forksPython—RLHFlow

Recipe and scripts for online iterative RLHF: SFT, reward modeling, vLLM response generation, reward annotation, and iterative DPO training loops. Reproduces LLaMA3-8B alignment comparable to Llama3-8B-Instruct using only open-source data.

Use Cases

Reproducing online iterative RLHF alignment from open dataRunning an automated iterative DPO training loopAnnotating generated responses with reward/preference modelsBulk response generation with vLLM for preference datasetsTraining SFT checkpoints on instruction dataAligning Llama3-8B to match or beat Llama3-8B-InstructResearch on RLHF vs offline alignment methods

Built With

Language
Python
Frameworks
vLLM · axolotl · DeepSpeed · HuggingFace Transformers · FastChat · alignment-handbook · accelerate · wandb

Tags

rlhf · dpo · llm-alignment · online-rlhf · iterative-dpo · reward-modeling · sft · preference-optimization · llama3 · vllm · deepspeed · axolotl · training-pipeline · fine-tuning