arXiv cs.CL
7/21/2026

RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
Short summary
RIMS is a three-stage preference optimization framework for small-scale language models in RAG settings, using synthetic chain-of-thought data generation, differentiable soft aggregation of multiple preference pairs, and smoothed objective optimization. It provably tightens gradient alignment to the oracle objective compared to hard selection and outperforms state-of-the-art baselines on four multi-hop QA benchmarks under noisy retrieval.
- •Three-stage framework: synthetic CoT preference data, soft multi-pair aggregation, and smoothed preference optimization
- •Theoretically proves tighter gradient alignment than hard selection with controllable error bounds
- •Outperforms SOTA baselines on four multi-hop QA benchmarks across multiple SLM backbones
Generated with AI, which can make mistakes.
Is this a good recommendation for you?