Back to feed
AR
arXiv CS.AI
7/15/2026
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

Short summary

GenAI Evaluation is a governed, configuration-driven pipeline for large-scale evaluation of retail conversational agents using LLM-as-a-judge scoring. It processes ~50,000 records daily and has evaluated over two million interactions, assessing helpfulness, truthfulness, clarity, tone, and translation quality. Validated against 12,980 human-labeled records, the pipeline achieved a macro F1 of 0.93 and 89% human-acceptability accuracy for translation.

  • Governed pipeline evaluates retail chatbot interactions at scale using LLM-as-a-judge with schema-constrained scoring
  • Processes ~50,000 records daily; over 2M interactions evaluated with selective re-evaluation for cost efficiency
  • Achieved macro F1 of 0.93 validated against 12,980 stratified-random human-labeled records across 14 intents and 129 sub-domains

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more