This is a heavily interactive web application, and JavaScript is required. Simple HTML interfaces are possible, but that is not what this is.
Post
Computational Linguistics Journal
complingjournal.bsky.social
did:plc:isbkpkr7jinlcksktfeeyp7g
Have you noticed that preference datasets in reinforcement learning/DPO are not fully utilised? This article proposes a framework called Vote-based Preference Optimisation that leverages user voting data to better align language models with diverse subjective preferences: https://doi.org/10.1162/COLI.a.579
2026-08-10T08:44:46.273Z