Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows Paper • 2605.27922 • Published May 27
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Paper • 2605.10912 • Published May 11 • 46
UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks Paper • 2607.08768 • Published 26 days ago • 34
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild Paper • 2605.27882 • Published May 27 • 17
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows Paper • 2605.14678 • Published May 19 • 108
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop Paper • 2606.22557 • Published Jun 21
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World Paper • 2605.26086 • Published May 25 • 25
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training Paper • 2607.05804 • Published 28 days ago • 19
MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents Paper • 2606.16748 • Published Jun 15 • 7
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments Paper • 2605.10779 • Published May 11
On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflows Paper • 2605.06110 • Published May 7 • 17