Human evaluation infrastructure for RLHF, model eval, safety testing, and alignment research.
Leaderboard of LLMs based on detailed human feedback