arxiv:2602.02493
ZehongMa
zehongma
AI & ML interests
MLLMs, Image/Video Generation, Multi-modal Representation Learning
Recent Activity
upvoted a paper 1 day ago
MiniWorld: Democratizing the Training of Video World Models from Scratch authored a paper 18 days ago
PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss upvoted a paper about 2 months ago
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion TransformerOrganizations
None yet