Online-RLHF
View on GitHubA recipe for online RLHF and online iterative DPO.
Recipe and scripts for online iterative RLHF: SFT, reward modeling, vLLM response generation, reward annotation, and iterative DPO training loops. Reproduces LLaMA3-8B alignment comparable to Llama3-8B-Instruct using only open-source data.
Use Cases
Reproducing online iterative RLHF alignment from open dataRunning an automated iterative DPO training loopAnnotating generated responses with reward/preference modelsBulk response generation with vLLM for preference datasetsTraining SFT checkpoints on instruction dataAligning Llama3-8B to match or beat Llama3-8B-InstructResearch on RLHF vs offline alignment methods
Built With
- Language
- Python
- Frameworks
- vLLM · axolotl · DeepSpeed · HuggingFace Transformers · FastChat · alignment-handbook · accelerate · wandb
Tags
rlhf · dpo · llm-alignment · online-rlhf · iterative-dpo · reward-modeling · sft · preference-optimization · llama3 · vllm · deepspeed · axolotl · training-pipeline · fine-tuning