Vibe Coding Discover

Use Cases

Annotating Generated Responses With Reward/preference Models

Published projects tagged with this use case.

1 project

Online-RLHF

★ 546

Recipe and scripts for online iterative RLHF: SFT, reward modeling, vLLM response generation, reward annotation, and iterative DPO training loops. Reproduces LLaMA3-8B alignment comparable to Llama3-8B-Instruct using only open-source data.

AI Frameworks | Python · rlhf · dpo

View Project →