Podcast
The podcast where we breakdown the recent AI papers and explain them in simple terms for you to understand.
QSVideo: Human-Style Search that Tames Long Videos – AI Breakdown
This episode breaks down QSVideo, a human-inspired framework that finds the few critical frames in long videos by rewriting questions, weighting object/action/location importance, and scoring frames with a lightweight semantic ranker.
QSVideo uses diversity based metric, an anchor-and-expand (or recency-first for streams) traversal, and parallel ranking to deliver notable accuracy gains on LVBench and StreamingBench, run much faster than prior methods, and plug into existing video-language models without retraining.
News
- QSVideo: Human-Style Search that Tames Long VideosThis episode breaks down QSVideo, a human-inspired framework that finds the few critical frames in long videos by rewriting questions,… Read more: QSVideo: Human-Style Search that Tames Long Videos
- Beyond Language Modeling: An Exploration of Multimodal PretrainingIn this episode, we discuss Beyond Language Modeling: An Exploration of Multimodal Pretraining by Shengbang Tong, David Fan, John Nguyen,… Read more: Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Mode Seeking meets Mean Seeking for Fast Long Video GenerationIn this episode, we discuss Mode Seeking meets Mean Seeking for Fast Long Video Generation by Shengqu Cai, Weili Nie,… Read more: Mode Seeking meets Mean Seeking for Fast Long Video Generation
- Recursive Language ModelsIn this episode, we discuss Recursive Language Models by Alex L. Zhang, Tim Kraska, Omar Khattab. The paper introduces Recursive… Read more: Recursive Language Models
- PaperBanana: Automating Academic Illustration for AI ScientistsIn this episode, we discuss PaperBanana: Automating Academic Illustration for AI Scientists by Dawei Zhu, Rui Meng, Yale Song, Xiyu… Read more: PaperBanana: Automating Academic Illustration for AI Scientists