Verifier-Induced Support Reshaping in On-Policy Optimization
Paper • 2608.00220 • Published • 4
nlu
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
UniWorld-Design: From Pixel Generation to Layer-Native Design