Agentic Transaction: Towards ACID-Compliant Agent Systems Paper • 2608.13900 • Published 6 days ago • 25
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs Paper • 2608.16391 • Published 3 days ago • 12
view article Article The Heterogeneous Feature of RoPE-based Attention in Long-Context LLMs SII-xrliu • Nov 15, 2025 • 16
Parameter Exploration for RLVR via Variational Learning Paper • 2608.09805 • Published 10 days ago • 6
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation Paper • 2608.11616 • Published 8 days ago • 6
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models Paper • 2608.06729 • Published 13 days ago • 2
NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs Paper • 2608.08107 • Published 12 days ago • 3
Persistent Recursive Worlds Enable Autonomous Software Evolution Paper • 2608.10450 • Published 8 days ago • 6
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research Paper • 2608.11216 • Published Jul 20 • 13
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents Paper • 2608.11878 • Published 8 days ago • 10
Can Test-Time Scaling Improve World Foundation Model? Paper • 2503.24320 • Published Mar 31, 2025 • 1
MedTsLLM: Leveraging LLMs for Multimodal Medical Time Series Analysis Paper • 2408.07773 • Published Aug 14, 2024 • 1
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Paper • 2505.11875 • Published May 17, 2025 • 2
The Art of Scaling Test-Time Compute for Large Language Models Paper • 2512.02008 • Published Dec 1, 2025 • 6
Faster and Better LLMs via Latency-Aware Test-Time Scaling Paper • 2505.19634 • Published May 26, 2025 • 1