Back to feed
arXiv cs.LG
arXiv cs.LG
7/16/2026
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

Short summary

This paper formalizes when to invoke expensive LLMs in streaming inference pipelines by casting it as a risk-based sequential stopping problem. It proves six theoretical results including sublinear regret bounds, convergence guarantees for adaptive thresholds, and optimality of threshold policies. Empirical validation on turbofan degradation data with real LLM calls confirms sublinear regret and shows anomaly-score-driven risk functions dominate baselines by an order of magnitude on Pareto AUC.

  • Formalizes LLM invocation timing as a risk-based sequential stopping problem with six proven theoretical results
  • Achieves O(sqrt(T log T)) regret on stationary streams, extending to changepoint settings
  • Empirically validates on CMAPSS turbofan data with real LLM calls, outperforming RouteLLM-style routers and contextual bandits

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more