Fetching the paper…
Reading the bibliography…
Generating engaging, accurate short-form videos from scientific papers is challenging due to content complexity and the gap between expert authors and readers.
Moviepy: a python library for video editing
Zulko · 2014
Earlier work this paper cites.
Contextually customized video summaries via natural language
Jinsoo Choi, Tae-Hyun Oh, and In So Kweon · 2018
Earlier work this paper cites.
A dataset of peer reviews (PeerRead): Collection, insights and NLP applications
Dongyeop Kang, Waleed Ammar, Bhavana Dalvi, Madeleine van Zuylen, Sebastian Kohlmeier, Eduard Hovy, and Roy Schwartz · 2018
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
“it took me almost 30 minutes to practice this.” performance and production practices in dance challenge videos on tiktok
Daniel Klug · 2020
Earlier work this paper cites.
Doc2ppt: Automatic presentation slides generation from scientific documents
Tsu-Jui Fu, William Yang Wang, Daniel J. McDuff, and Yale Song · 2021
Earlier work this paper cites.
Augmenting scientific papers with just-in-time, position-sensitive definitions of terms and symbols
Andrew Head, Kyle Lo, Dongyeop Kang, Raymond Fok, Sam Skjonsberg, Daniel S Weld, and Marti A Hearst · 2021
Earlier work this paper cites.
Video diffusion models
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet · 2022
Earlier work this paper cites.
Automated video editing based on learned styles using lstm-gan
Hsin-I Huang, Chi-Sheng Shih, and Zi-Lin Yang · 2022
Earlier work this paper cites.
Mcvd: Masked conditional video diffusion for prediction, generation, and interpolation
Vikram Voleti, Alexia Jolicoeur-Martineau, and Christopher Pal · 2022
Earlier work this paper cites.
Generating videos with dynamics-aware implicit generative adversarial networks
Sihyun Yu, Jihoon Tack, Sangwoo Mo, Hyunsu Kim, Junho Kim, Jung-Woo Ha, and Jinwoo Shin · 2022
Earlier work this paper cites.
A suite of generative tasks for multi-level multimodal webpage understanding
Andrea Burns, Krishna Srinivasan, Joshua Ainslie, Geoff Brown, Bryan A. Plummer, Kate Saenko, Jianmo Ni, and Mandy Guo · 2023
Earlier work this paper cites.
Speed is all you need: On-device acceleration of large diffusion models via gpu-aware optimizations
Yu-Hui Chen, Raman Sarokin, Juhyun Lee, Jiuqiang Tang, Chuo-Ling Chang, Andrei Kulik, and Matthias Grundmann · 2023
Earlier work this paper cites.
Creator-friendly algorithms: Behaviors, challenges, and design opportunities in algorithmic platforms
Yoonseo Choi, Eun Jeong Kang, Min Kyung Lee, and Juho Kim · 2023
Earlier work this paper cites.
Establishing tiktok as a platform for informal learning: Evidence from mixed-methods analysis of creators and viewers
Sourojit Ghosh and Andrea Figueroa · 2023
Earlier work this paper cites.
For who page? tiktok creators’ algorithmic dependencies
Laura Herman · 2023
Cited alongside, same era.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman · 2023
Cited alongside, same era.
Supporting video authoring for communication of research results
Katharina Wünsche, Laura Koesten, Torsten Möller, and Jian Chen · 2023
Cited alongside, same era.
Tuning large multimodal models for videos using reinforcement learning from AI feedback
Daechul Ahn, Yura Choi, Youngjae Yu, Dongyeop Kang, and Jonghyun Choi · 2024
Cited alongside, same era.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh · 2024
Cited alongside, same era.
Consisti2v: Enhancing visual consistency for image-to-video generation
Weiming Ren, Huan Yang, Ge Zhang, Cong Wei, Xinrun Du, Wenhao Huang, and Wenhu Chen · 2024
Later among the works it cites.
A rockslide-generated tsunami in a greenland fjord rang earth for 9 days
Kristian Svennevig, Stephen P. Hicks, Thomas Forbriger, Thomas Lecocq, Rudolf Widmer-Schnidrig, Anne Mangeney, Clément Hibert, Niels J. Korsgaard, Antoine Lucas, Claudio Satriano, Robert E. Anthony, Aurélien Mordret, Sven Schippkus, Søren Rysgaard, Wieter Boone, Steven J. Gibbons, Kristen L. Cook, Sylfest Glimsdal, Finn Løvholt, Koen Van Noten, Jelle D. Assink, Alexis Marboeuf, Anthony Lomax, Kris Vanneste, Taka’aki Taira, Matteo Spagnolo, Raphael De Plaen, Paula Koelemeijer, Carl Ebeling, Andrea Cannata, William D. Harcourt, David G. Cornwell, Corentin Caudron, Piero Poli, Pascal Bernard, Eric Larose, Eleonore Stutzmann, Peter H. Voss, Bjorn Lund, Flavio Cannavo, Manuel J. Castro-Díaz, Esteban Chaves, Trine Dahl-Jensen, Nicolas De Pinho Dias, Aline Déprez, Roeland Develter, Douglas Dreger, Läslo G. Evers, Enrique D. Fernández-Nieto, Ana M. G. Ferreira, Gareth Funning, Alice-Agnes Gabriel, Marc Hendrickx, Alan L. Kafka, Marie Keiding, Jeffrey Kerby, Shfaqat A. Khan, Andreas Kjær Dideriksen, Oliver D. Lamb, Tine B. Larsen, Bradley Lipovsky, Ikha Magdalena, Jean-Philippe Malet, Mikkel Myrup, Luis Rivera, Eugenio Ruiz-Castillo, Selina Wetter, and Bastien Wirtz · 2024
Later among the works it cites.
Rong-Cheng Tu, Wenhao Sun, Zhao Jin, Jingyi Liao, Jiaxing Huang, and Dacheng Tao · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karin De Langis, Ryan Koo, and Dongyeop Kang · 2024
Cited alongside, same era.
Kubrick: Multimodal agent collaborations for synthetic video generation, 2024
Liu He, Yizhi Song, Hejun Huang, Daniel Aliaga, and Xin Zhou · 2024
Cited alongside, same era.
Storyagent: Customized storytelling video generation via multi-agent collaboration, 2024
Panwen Hu, Jin Jiang, Jianqi Chen, Mingfei Han, Shengcai Liao, Xiaojun Chang, and Xiaodan Liang · 2024
Cited alongside, same era.
Genmac: Compositional text-to-video generation with multi-agent collaboration
Kaiyi Huang, Yukun Huang, Xuefei Ning, Zinan Lin, Yu Wang, and Xihui Liu · 2024
Cited alongside, same era.
Threads of subtlety: Detecting machine-generated texts through discourse motifs
Zae Myung Kim, Kwang Lee, Preston Zhu, Vipul Raheja, and Dongyeop Kang · 2024
Cited alongside, same era.
Longform multimodal lay summarization of scientific papers: Towards automatically generating science blogs from research articles
Sandeep Kumar, Guneet Singh Kohli, Tirthankar Ghosal, and Asif Ekbal · 2024
Cited alongside, same era.
Multimodal ArXiv: A dataset for improving scientific comprehension of large vision-language models
Lei Li, Yuqi Wang, Runxin Xu, Peiyi Wang, Xiachong Feng, Lingpeng Kong, and Qi Liu · 2024
Cited alongside, same era.
Jiuniu Wang, Zehua Du, Yuyuan Zhao, Bo Yuan, Kexiang Wang, Jian Liang, Yaxi Zhao, Yihen Lu, Gengliang Li, Junlong Gao, Xin Tu, and Zhenyu Guo · 2024
Later among the works it cites.
KNowNEt:Guided Health Information Seeking from LLMs via Knowledge Graph Integration
Youfu Yan, Yu Hou, Yongkang Xiao, Rui Zhang, and Qianwen Wang · 2024
Later among the works it cites.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu · 2024
Later among the works it cites.
Mora: Enabling generalist video generation via a multi-agent framework, 2024
Zhengqing Yuan, Yixin Liu, Yihan Cao, Weixiang Sun, Haolong Jia, Ruoxi Chen, Zhaoxu Li, Bin Lin, Li Yuan, Lifang He, Chi Wang, Yanfang Ye, and Lichao Sun · 2024
Later among the works it cites.
Llava-next: A strong zero-shot video understanding model, April 2024
Yuanhan Zhang, Bo Li, haotian Liu, Yong jae Lee, Liangke Gui, Di Fu, Jiashi Feng, Ziwei Liu, and Chunyuan Li · 2024
Later among the works it cites.
Contextual document embeddings
John Xavier Morris and Alexander M Rush · 2025
Closest in time.
Flash talks competition, September 2020
Research!America · 2025
Closest in time.
Querying databases with function calling, 2025
Connor Shorten, Charles Pierse, Thomas Benjamin Smith, Karel D’Oosterlinck, Tuana Celik, Erika Cardenas, Leonie Monigatti, Mohd Shukri Hasan, Edward Schmuhl, Daniel Williams, Aravind Kesiraju, and Bob van Luijt · 2025
Closest in time.
Videoagent: Self-improving video generation, 2025
Achint Soni, Sreyas Venkataraman, Abhranil Chandra, Sebastian Fischmeister, Percy Liang, Bo Dai, and Sherry Yang · 2025
Closest in time.
Scholawrite: A dataset of end-to-end scholarly writing process
Linghe Wang, Minhwa Lee, Ross Volkov, Luan Tuyen Chau, and Dongyeop Kang · 2025
Closest in time.
Filmagent: A multi-agent framework for end-to-end film automation in virtual 3d spaces, 2025
Zhenran Xu, Longyue Wang, Jifang Wang, Zhouyi Li, Senbao Shi, Xue Yang, Yiyu Wang, Baotian Hu, Jun Yu, and Min Zhang · 2025
Closest in time.