view article Article What Is MiniMax H3 (Hailuo 3.0)? The Open-Weight Multimodal Video Model, Explained ResterChed • 24 days ago • 2
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published 21 days ago • 61
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 22 days ago • 102
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 88
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 217
DPWriter: Reinforcement Learning with Diverse Planning Branching for Creative Writing Paper • 2601.09609 • Published Jan 14 • 4
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE Paper • 2507.21802 • Published Jul 29, 2025 • 20
V-JEPA 2 Collection A frontier video understanding model developed by FAIR, Meta, which extends the pretraining objectives of https://ai.meta.com/blog/v-jepa-yann • 8 items • Updated Jun 13, 2025 • 229
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Paper • 2511.04570 • Published Nov 6, 2025 • 242
Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation Paper • 2510.01284 • Published Sep 30, 2025 • 37
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Paper • 2510.03117 • Published Oct 3, 2025 • 12
Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding? Paper • 2505.14321 • Published May 20, 2025 • 11
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization Paper • 2503.23377 • Published Mar 30, 2025 • 57
BSharedRAG: Backbone Shared Retrieval-Augmented Generation for the E-commerce Domain Paper • 2409.20075 • Published Sep 30, 2024 • 2
ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering Paper • 2503.16867 • Published Mar 21, 2025 • 12