AR
arXiv CS.AI
7/14/2026

Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking
Short summary
This paper introduces the Format Sensitivity Index (FSI) and Parseability Sensitivity Index (PSI) to measure how prompt-wrapper formatting affects LLM benchmark scores. Across 140,000 OpenRouter generations spanning 7 QA tasks, 5 wrapper families, and 4 models (7B–72B), mean FSI varies by over 30x and is largely driven by compliance failures. The authors argue that reporting accuracy without wrapper variance is statistically fragile and offer practical recommendations for benchmarking and structured-output deployments.
- •Introduces FSI and PSI metrics to quantify prompt-wrapper-induced score variance in LLM benchmarks
- •140k generations across 7 tasks, 5 wrappers, 4 models show FSI varies 30x+, driven by compliance failures
- •Recommends reporting wrapper variance and parseability alongside accuracy for robust benchmarking
Generated with AI, which can make mistakes.
Is this a good recommendation for you?
