Other
Apple researchers report that grounded, multi-dimensional rubric rewards improve open-domain question-answering performance by 6.5% over an instruction-tuned baseline and 4% over flat rubric variants.
Verified
Emerging
1 source
- First seen
- Last updated
- Event ID
4ddb6ca8-0c68-4f6a-9d3e-b89699164bed
Source trail
-
machinelearning.apple.comFrom Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge AnswersAnnouncement PublishedOpen source ↗
Observation timeline
- Announcement Publishedmachinelearning.apple.com · ORIGINAL
Research & reporting
Editorial work is a separate, human-published layer. It never defines this Event.
No editorial Article is attached. The Event remains public because editorial publication is optional.