Fri 2 Oct 2026

UTC Live

Super Intelligence
News and research
Every item sourced

SUPENCE

Superintelligence latest news and info

Tips and corrections
Write to Supence
contact@supence.info

Source Microsoft Research · Published · Open the original ↗

Lab notes

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them.

AI agents have evolved from single models to complex full-stack systems built from models, tools, and execution environments. Their capabilities increasingly depend on the agent harness that coordinates them from outside the model. Reinforcement learning (RL) is an approach where AI systems learn through trial and error, guided by rewards and penalties for their actions. RL can make those agents better, but most agent RL systems require developers to reimplement the agent inside the training framework. That is costly, and it means the agent being trained is not quite the agent that gets deployed.

To address this, researchers at Microsoft Research Asia have introduced the Harnessed Agentic RL training paradigm and open-sourced a fully rebuilt Agent Lightning v1.0 (opens in new tab) . Compared with the original, Agent Lightning, v1.0 puts more emphasis on staying lightweight, on integrating with real harnesses, and on a complete, reproducible agent RL training pipeline.

The opening of the post, quoted unchanged from Microsoft Research. The full text continues at the source.

Continue reading at Microsoft Research ↗

More on Supence Newest first

  1. The problem with your AI therapist

    Capability · Financial Times

  2. What if AI prefers CVs written by AI?

    Capability · Financial Times