Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Docs
  • Enterprise
  • Pricing

  • Log In
  • Sign Up
Mu Cai's picture
6 9 3

Mu Cai

mucai
tolgacangoz's profile picture thomas-yanxin's profile picture variante's profile picture
·
https://pages.cs.wisc.edu/~mucai/
  • MuCai7
  • mu-cai

AI & ML interests

Computer Vision, Deep Learning, 3D Vision, Vision and Language,

Recent Activity

upvoted a paper 4 days ago
When Tokens Talk Too Much: A Survey of Multimodal Long-Context Token Compression across Images, Videos, and Audios
upvoted a paper 18 days ago
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
upvoted a paper 6 months ago
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
View all activity

Organizations

vgbench's profile picture CounterCurate's profile picture

authored 2 papers 10 months ago

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Paper • 2410.10818 • Published Oct 14, 2024 • 17

Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos

Paper • 2410.02763 • Published Oct 3, 2024 • 7
authored 2 papers about 1 year ago

LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

Paper • 2406.20095 • Published Jun 28, 2024 • 18

Matryoshka Multimodal Models

Paper • 2405.17430 • Published May 27, 2024 • 35
Company
TOS Privacy About Jobs
Website
Models Datasets Spaces Pricing Docs