StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 6 days ago • 9
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published 21 days ago • 24
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published 6 days ago • 9
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Paper • 2606.12191 • Published Jun 10 • 71
SWE-Touch: Benchmarking Coding Agents When Users Touch the Code Paper • 2608.02499 • Published 21 days ago • 24
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published 27 days ago • 158
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity Paper • 2607.00248 • Published Jun 30 • 32
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 42 • 5
GATE: Graph-based Adaptive Tool Evolution Across Diverse Tasks Paper • 2502.14848 • Published Feb 20, 2025 • 1
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 42
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 42 • 5
DV-World: Benchmarking Data Visualization Agents in Real-World Scenarios Paper • 2604.25914 • Published Apr 28 • 42