• Precios
  • Iniciar sesión
Preference Optimization

Preference Optimization

por SaiKrishna Rallabandi
Educación
Temporada 1
ASFT: Aligned Supervised Fine-Tuning through Absolute Likelihood
This paper proposes a new method for fine-tuning large language models (LLMs) called Aligned Supervised Fine-Tuning (ASFT). ASFT addresses limitations of existing Direct Preference Optimization (DPO) methods by optimizing the absolute likelihood of generating human-preferred responses rather than relying on relative likelihoods. Unlike DPO, ASFT does not require a reference model and is less sensitive to the initial state of the model, leading to more efficient and robust training. The authors demonstrate the effectiveness of ASFT through extensive experiments on various benchmark datasets, showing significant performance improvements compared to existing methods.
T1 · E1
8 oct 2024
11:17
© RSS America LLC.
Podcast Standards CertifiedIAB Tech Lab
Recursos
  • Novedades ✨
  • Afiliados
  • Blog
  • Prensa
  • Kit de prensa
  • Media Kit
  • Ayuda
Más en RSS.com
  • Partners
  • Reseñas
  • Herramientas
  • Audio a Video
  • Encuentra mi feed
  • Cambia a RSS.com
Legal
  • Politica de Cookies
  • Política de Privacidad y Cookies
  • Términos del servicio
Comunidad
Los mejores shows disponibles en RSS.com.
Escuchalos ahora
Elige idioma