
RAGthoven at SemEval-2026: Multi-Stage Humor Generation Pipeline Ties Frontier Model Baseline
Original: RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar
Short summary
RAGthoven is a multi-stage LLM pipeline for multilingual humor generation (Planner, Best-of-N Writer, Reflector, Judge) grounded in computational humor theory, evaluated at SemEval-2026 Task 1. It ties for Rank 1 with the Gemini 2.5 Flash baseline across English, Spanish, and Chinese, leading in Spanish by 42 Elo points but trailing in English and Chinese within statistical ties. Agentic variants (ReAct-style and autonomous orchestration) did not outperform the simpler non-agentic pipeline despite higher tool-call budgets, suggesting diminishing returns from elaborate scaffolding with strong frontier models.
- •RAGthoven ties Gemini 2.5 Flash baseline for Rank 1 in multilingual humor generation across three languages
- •Agentic variants with higher tool-call budgets did not outperform the simpler non-agentic pipeline
- •Results suggest language-dependent diminishing returns from multi-stage prompt engineering with strong frontier models
Generated with AI, which can make mistakes.
Is this a good recommendation for you?