Learning Video LLM with Streaming Speech Transcription at Scale (CVPR 2025)
Joya Chen
chenjoya
AI & ML interests
Video LLM
Recent Activity
updated a dataset about 3 hours ago
chenjoya/Live-WhisperX-526K upvoted a paper about 1 month ago
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning upvoted a paper about 1 month ago
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity