Back to feed
r/MachineLearning
r/MachineLearning
7/14/2026
[P] RL-training Qwen3.6 to RL-train tool using AI models [P]

[P] RL-training Qwen3.6 to RL-train tool using AI models [P]

Short summary

A developer built and open-sourced a recursive RL system where a Qwen3.6-35B agent learns to write complete training jobs for smaller Qwen models, with the inner model's improvement as the reward signal. Over 54 outer-loop steps (~1,750 GPU jobs, ~$1.3k total), the agent learned to pick better base models, use hyperparameters, and generalize to held-out tasks. The full harness, task families, and code are on GitHub.

  • RL-trained Qwen3.6 agent that writes and dispatches RL training jobs for smaller models
  • Achieved 0.0→0.63 reward over 54 steps with generalization to held-out tasks at ~$1.3k total cost
  • Fully open-sourced: harness, reward code, GPU orchestration, and Tinker RL scripts on GitHub

Generated with AI, which can make mistakes.

Is this a good recommendation for you?

Comments

Failed to load comments. Please try again.

Explore more