Skip to results
MLSift
← Feed
routineNLP & Language ModelsDPO2606.12881

Direct Preference Optimization for Chatbot Fine-Tuning: An Empirical Study

Yvonne Qiu, Dezhi Yu, ShuoJia Fu

cs.CL cs.LG

Abstract

We present an approach to fine-tuning large language models using Direct Preference Optimization (DPO), a reinforcement learning technique. Our experimental results demonstrate that DPO simplifies the training pipeline, improves computational efficiency, and achieves competitive performance. The evaluation using BLEU, ROUGE, and cosine similarity metrics indicates effective learning and convergence, though further investigation is needed to address observed training instability.

Topics

Classified with taxonomy v2 on Wed, 2 Sept 2026.

The PDF is 1–3 MB. Open it in your browser's viewer, or load it here.

Open PDF