Incantory
Sign in
SkillMITNot scanned

Preference Optimization

Align a fine-tuned model with preference data using DPO, ORPO, KTO, or SimPO. Use when preference pairs or thumbs-up/down feedback exist, when choosing between preference-optimization methods, or when a DPO run needs hyperparameters or debugging.

wshobson/agentsv10 stars · 0 forks · 0 makes≈1.9K tokens

Family tree

1 prompt · 1 version shown · 0 makes · Full screen

No forks or makes yet. When someone forks this prompt or shares something they made with it, it grows here.

The tree as a list