RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System
General agentic RL · corresponding-author work
Latent Collaboration in Multi-Agent Systems (LatentMAS)
Efficient multi-agent reasoning in latent space · corresponding-author work
OpenClaw-RL: Train Any Agent Simply by Talking
Fully asynchronous RL for interactive agents
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models (TraceRL)
Trajectory-aware RL for diffusion LLMs · corresponding-author work
MMaDA: Multimodal Large Diffusion Language Models
Unified multimodal generative modeling
Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning (ReasonFlux-Coder)
Coder / unit-tester co-evolution via RL · corresponding-author work
ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs
Trajectory-aware process reward models · corresponding-author work
ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
Structured reasoning and post-training
Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models
Structured reasoning with thought templates