AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl 0.4B • Updated Apr 30 • 21
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_propsft_propprefix_nokl 0.4B • Updated Apr 30 • 20
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-184_eval-dataset Viewer • Updated May 1 • 6.45k • 13
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-26_eval-dataset Viewer • Updated May 1 • 6.45k • 12
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-78_eval-dataset Viewer • Updated May 1 • 6.45k • 9
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-52_eval-dataset Viewer • Updated May 1 • 6.45k • 13
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-104_eval-dataset Viewer • Updated May 1 • 6.45k • 11
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft0.3_prefix_nokl_checkpoint-255_eval-dataset Viewer • Updated May 1 • 6.45k • 13
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_mergedsft_prefix_nokl_checkpoint-255_eval-dataset Viewer • Updated Apr 30 • 6.45k • 9
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-255_eval-dataset Viewer • Updated Apr 30 • 6.45k • 19
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-104_eval-dataset Viewer • Updated Apr 30 • 6.45k • 22
AdversarialRLHF/rloo_pythia410m_tldr6.9b_rm410mdata_checkpoint-52_eval-dataset Viewer • Updated Apr 30 • 6.45k • 14