Back to feed
AR
arXiv CS.AI
7/14/2026
Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction

Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction

Short summary

This work demonstrates that a compact 1.5B parameter language model (Qwen2.5-1.5B) aligned via GRPO can serve as an effective edge-deployed control agent when paired with a digital-twin validator and reprompting loop. In thermal-control simulations it achieved 91.5% action-alignment accuracy at 3.84s mean latency, with a 95% in-range rate under symbolic re-mapping. The architecture offers a practical path toward reconfigurable autonomous control without large cloud-based models.

  • Qwen2.5-1.5B aligned via GRPO achieves 91.5% action-alignment in thermal-control simulations
  • Multi-agent framework pairs action agent with digital-twin validator and reprompting agent for iterative correction
  • 3.84s mean inference latency enables practical edge deployment for closed-loop industrial control

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more