AR
arXiv CS.AI
7/15/2026

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
Short summary
GenAI Evaluation is a governed, configuration-driven pipeline for large-scale evaluation of retail conversational agents using LLM-as-a-judge scoring. It processes ~50,000 records daily and has evaluated over two million interactions, assessing helpfulness, truthfulness, clarity, tone, and translation quality. Validated against 12,980 human-labeled records, the pipeline achieved a macro F1 of 0.93 and 89% human-acceptability accuracy for translation.
- •Governed pipeline evaluates retail chatbot interactions at scale using LLM-as-a-judge with schema-constrained scoring
- •Processes ~50,000 records daily; over 2M interactions evaluated with selective re-evaluation for cost efficiency
- •Achieved macro F1 of 0.93 validated against 12,980 stratified-random human-labeled records across 14 intents and 129 sub-domains
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



