arXiv cs.LG
8/3/2026

LARA: Lightweight Adapters in the Residual Stream for Composable Adaptation and Alignment
Short summary
LARA introduces lightweight residual-stream adapters that correct hidden states at selected layers without modifying base weights, matching LoRA at equal parameter counts. Its key advantage is a scale parameter enabling smooth interpolation between base and adapted behavior at inference. Multiple behaviors can coexist on one frozen model—seven behaviors fit on a 1.5B model with ~33 MB overhead—and are routed per token, making it suitable for hosting many behaviors on-device.
- •LARA adapts frozen models via residual-stream corrections, matching LoRA performance
- •Scale parameter enables graded interpolation between base and adapted behavior
- •Seven behaviors coexist on one 1.5B model with ~33 MB total overhead
Generated with AI, which can make mistakes.
Is this a good recommendation for you?

