Selected publications

Work across models, agents, RL, and discovery.

This page highlights representative work connected to the research trajectory behind DIG. For the complete publication record, see the PI's Google Scholar profile.

ICML 2026

RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System

General agentic RL · corresponding-author work

ICML 2026Spotlight · Top 3%

Latent Collaboration in Multi-Agent Systems (LatentMAS)

Efficient multi-agent reasoning in latent space · corresponding-author work

TECH REPORT

OpenClaw-RL: Train Any Agent Simply by Talking

Fully asynchronous RL for interactive agents

ICLR 2026

Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models (TraceRL)

Trajectory-aware RL for diffusion LLMs · corresponding-author work

NEURIPS 2025

MMaDA: Multimodal Large Diffusion Language Models

Unified multimodal generative modeling

NEURIPS 2025Spotlight · Top 3%

Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning (ReasonFlux-Coder)

Coder / unit-tester co-evolution via RL · corresponding-author work

NEURIPS 2025

ReasonFlux-PRM: Trajectory-Aware PRMs for Long Chain-of-Thought Reasoning in LLMs

Trajectory-aware process reward models · corresponding-author work

PREPRINT

ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates

Structured reasoning and post-training

NEURIPS 2024Spotlight · Top 3%

Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models

Structured reasoning with thought templates

ICML 2024

Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Generative modeling foundation (RPG)

Research record

70+ papers across machine learning and AI.

Representative venues include NeurIPS, ICML, ICLR, CVPR, ACL and related conferences, with multiple Spotlight, Oral, and Best Paper recognitions.

Complete bibliographyUse Google Scholar for the full and most up-to-date publication record.