r/MachineLearning
7/14/2026
![[P] RL-training Qwen3.6 to RL-train tool using AI models [P]](https://preview.redd.it/hg7ww6ute8dh1.png?width=140&height=75&auto=webp&s=d9c4aa6843cd8b9f2a480a97e0f8469ec80b4d41)
[P] RL-training Qwen3.6 to RL-train tool using AI models [P]
Short summary
A developer built and open-sourced a recursive RL system where a Qwen3.6-35B agent learns to write complete training jobs for smaller Qwen models, with the inner model's improvement as the reward signal. Over 54 outer-loop steps (~1,750 GPU jobs, ~$1.3k total), the agent learned to pick better base models, use hyperparameters, and generalize to held-out tasks. The full harness, task families, and code are on GitHub.
- •RL-trained Qwen3.6 agent that writes and dispatches RL training jobs for smaller models
- •Achieved 0.0→0.63 reward over 54 steps with generalization to held-out tasks at ~$1.3k total cost
- •Fully open-sourced: harness, reward code, GPU orchestration, and Tinker RL scripts on GitHub
Generated with AI, which can make mistakes.
Is this a good recommendation for you?


