CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes Paper • 2608.27455 • Published 9 days ago • 11
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses Paper • 2608.24876 • Published 11 days ago • 28
FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory Paper • 2608.04530 • Published about 1 month ago • 14
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published about 1 month ago • 27
OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration Paper • 2605.28805 • Published May 27 • 11
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Paper • 2606.19534 • Published Jun 17 • 65
LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing Paper • 2606.06042 • Published Jun 4 • 25
Benchmarking and Evolving Reason-Reflect-Rectify for Reflective Visual Generation Paper • 2605.19639 • Published May 19
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Paper • 2606.19534 • Published Jun 17 • 65
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision Paper • 2604.12002 • Published Apr 13 • 12
GenEnv: Difficulty-Aligned Co-Evolution Between LLM Agents and Environment Simulators Paper • 2512.19682 • Published Dec 22, 2025 • 19
Beyond Binary Preference: Aligning Diffusion Models to Fine-grained Criteria by Decoupling Attributes Paper • 2601.04300 • Published Jan 7 • 3
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning Paper • 2512.13278 • Published Dec 15, 2025
RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System Paper • 2602.02488 • Published Feb 2 • 36