Back to feed
arXiv cs.CL
arXiv cs.CL
7/21/2026
RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation

RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation

Short summary

RIMS is a three-stage preference optimization framework for small-scale language models in RAG settings, using synthetic chain-of-thought data generation, differentiable soft aggregation of multiple preference pairs, and smoothed objective optimization. It provably tightens gradient alignment to the oracle objective compared to hard selection and outperforms state-of-the-art baselines on four multi-hop QA benchmarks under noisy retrieval.

  • Three-stage framework: synthetic CoT preference data, soft multi-pair aggregation, and smoothed preference optimization
  • Theoretically proves tighter gradient alignment than hard selection with controllable error bounds
  • Outperforms SOTA baselines on four multi-hop QA benchmarks across multiple SLM backbones

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more