Preference datasets trl-lib/hh-rlhf-helpful-base Viewer • Updated Jan 8, 2025 • 46.2k • 183 • 3 trl-lib/lm-human-preferences-descriptiveness Viewer • Updated Jan 8, 2025 • 6.26k • 40 • 2 trl-lib/lm-human-preferences-sentiment Viewer • Updated Jan 8, 2025 • 6.26k • 27 trl-lib/rlaif-v Viewer • Updated Jan 8, 2025 • 83.1k • 165 • 4
Prompt-completion datasets trl-lib/tldr Viewer • Updated Jan 8, 2025 • 130k • 4.3k • 32 trl-lib/OpenMathReasoning Viewer • Updated Apr 26, 2025 • 3.2M • 191
Unpaired preference datasets trl-lib/ultrafeedback-gpt-3.5-turbo-helpfulness Viewer • Updated Jan 8, 2025 • 16.6k • 44 • 5 trl-lib/kto-mix-14k Viewer • Updated Mar 25, 2024 • 15k • 211 • 9
trl-lib/ultrafeedback-gpt-3.5-turbo-helpfulness Viewer • Updated Jan 8, 2025 • 16.6k • 44 • 5
Online-DPO trl-lib/pythia-1b-deduped-tldr-sft 1B • Updated Aug 2, 2024 • 1.19k trl-lib/pythia-2.8b-deduped-tldr-sft Updated Aug 2, 2024 • 65 trl-lib/pythia-2.8b-deduped-tldr-rm Updated Aug 2, 2024 • 7 trl-lib/pythia-6.9b-deduped-tldr-sft Updated Aug 2, 2024 • 11
Stepwise supervision datasets trl-lib/math_shepherd Viewer • Updated Jan 8, 2025 • 445k • 2.63k • 12 trl-lib/prm800k Viewer • Updated Jan 8, 2025 • 41.2k • 158 • 7
Prompt-only datasets trl-lib/ultrafeedback-prompt Viewer • Updated Jan 8, 2025 • 39.8k • 605 • 10 trl-lib/DeepMath-103K Viewer • Updated Nov 14, 2025 • 103k • 1.78k • 15
Comparing DPO with IPO and KTO A collection of chat models to explore the differences between three alignment techniques: DPO, IPO, and KTO. teknium/OpenHermes-2.5-Mistral-7B Text Generation • 7B • Updated Feb 19, 2024 • 5.76k • • 909 Intel/orca_dpo_pairs Viewer • Updated Nov 29, 2023 • 12.9k • 2.26k • 324
teknium/OpenHermes-2.5-Mistral-7B Text Generation • 7B • Updated Feb 19, 2024 • 5.76k • • 909
Preference datasets trl-lib/hh-rlhf-helpful-base Viewer • Updated Jan 8, 2025 • 46.2k • 183 • 3 trl-lib/lm-human-preferences-descriptiveness Viewer • Updated Jan 8, 2025 • 6.26k • 40 • 2 trl-lib/lm-human-preferences-sentiment Viewer • Updated Jan 8, 2025 • 6.26k • 27 trl-lib/rlaif-v Viewer • Updated Jan 8, 2025 • 83.1k • 165 • 4
Stepwise supervision datasets trl-lib/math_shepherd Viewer • Updated Jan 8, 2025 • 445k • 2.63k • 12 trl-lib/prm800k Viewer • Updated Jan 8, 2025 • 41.2k • 158 • 7
Prompt-completion datasets trl-lib/tldr Viewer • Updated Jan 8, 2025 • 130k • 4.3k • 32 trl-lib/OpenMathReasoning Viewer • Updated Apr 26, 2025 • 3.2M • 191
Prompt-only datasets trl-lib/ultrafeedback-prompt Viewer • Updated Jan 8, 2025 • 39.8k • 605 • 10 trl-lib/DeepMath-103K Viewer • Updated Nov 14, 2025 • 103k • 1.78k • 15
Unpaired preference datasets trl-lib/ultrafeedback-gpt-3.5-turbo-helpfulness Viewer • Updated Jan 8, 2025 • 16.6k • 44 • 5 trl-lib/kto-mix-14k Viewer • Updated Mar 25, 2024 • 15k • 211 • 9
trl-lib/ultrafeedback-gpt-3.5-turbo-helpfulness Viewer • Updated Jan 8, 2025 • 16.6k • 44 • 5
Comparing DPO with IPO and KTO A collection of chat models to explore the differences between three alignment techniques: DPO, IPO, and KTO. teknium/OpenHermes-2.5-Mistral-7B Text Generation • 7B • Updated Feb 19, 2024 • 5.76k • • 909 Intel/orca_dpo_pairs Viewer • Updated Nov 29, 2023 • 12.9k • 2.26k • 324
teknium/OpenHermes-2.5-Mistral-7B Text Generation • 7B • Updated Feb 19, 2024 • 5.76k • • 909
Online-DPO trl-lib/pythia-1b-deduped-tldr-sft 1B • Updated Aug 2, 2024 • 1.19k trl-lib/pythia-2.8b-deduped-tldr-sft Updated Aug 2, 2024 • 65 trl-lib/pythia-2.8b-deduped-tldr-rm Updated Aug 2, 2024 • 7 trl-lib/pythia-6.9b-deduped-tldr-sft Updated Aug 2, 2024 • 11