Fetching the paper…
Reading the bibliography…
Generating high-quality shooting scripts containing information such as scene and shot language is essential for short drama script generation.
Emotion: Theory, Research, and Experience: Vol. 1. Theories of Emotion
Robert Plutchik and Henry Kellerman · 1980
Earlier work this paper cites.
When and why test-time augmentation works
Divya Shanmugam, Davis W. Blalock, Guha Balakrishnan, and John V. Guttag · 2011
Earlier work this paper cites.
Decaf: Meg-based multimodal database for decoding affective physiological responses
Mojtaba Khomami Abadi, Ramanathan Subramanian, Seyed Mostafa Kia, Paolo Avesani, Ioannis Patras, and Nicu Sebe · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
RMPE: Regional multi-person pose estimation
Hao-Shu Fang, Shuqin Xie, Yu-Wing Tai, and Cewu Lu · 2017
Earlier work this paper cites.
Localizing moments in video with natural language
Lisa Anne Hendricks, Oliver Wang, Eli Shechtman, Josef Sivic, Trevor Darrell, and Bryan Russell · 2017
Earlier work this paper cites.
Dense-captioning events in videos
Ranjay Krishna, Kenji Hata, Frederic Ren, Li Fei-Fei, and Juan Carlos Niebles · 2017
Earlier work this paper cites.
Movie description
Anna Rohrbach, Atousa Torabi, Marcus Rohrbach, Niket Tandon, Chris Pal, Hugo Larochelle, Aaron Courville, and Bernt Schiele · 2017
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Dreamer: A database for emotion recognition through eeg and ecg signals from wireless low-cost off-the-shelf devices
Stamos Katsigiannis and Naveed Ramzan · 2018
Earlier work this paper cites.
How2: A large-scale dataset for multimodal language understanding
Rene Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Graham Neubig, Florian Metze, and Lucia Specia · 2018
Earlier work this paper cites.
Towards a bottom-up framework for video captioning
Luowei Zhou, Chenliang Xu, and Jason J. Corso · 2018
Earlier work this paper cites.
Crowdpose: Efficient crowded scenes pose estimation and a new benchmark
Jiefeng Li, Can Wang, Hao Zhu, Yihuan Mao, Hao-Shu Fang, and Cewu Lu · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Plan-and-write: Towards better automatic storytelling
Lifu Yao, Nanyun Peng, Ralph Weischedel, Kevin Knight, Dongyan Zhao, and Rui Yan · 2019
Earlier work this paper cites.
Lightface: A hybrid deep face recognition framework
Sefik Ilkin Serengil and Alper Ozpinar · 2020
Earlier work this paper cites.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Max Bain, Arsha Nagrani, Gül Varol, and Andrew Zisserman · 2021
Earlier work this paper cites.
Amigos: A dataset for affect, personality and mood research on individuals and groups
Juan A. Correa, Mohammad Khaleghi Abadi, Marcos Faudez-Zanuy, Fabián Arango, Hugo J. Escalante, Jüri Allik, and Mohammad Soleymani · 2021
Earlier work this paper cites.
Towards coherent and consistent use of entities in narrative generation
Nanyun Peng, Demian Gholipour Ghalandari, Hanyu Fang, Harsh Jhamtani, and Kevin Knight · 2021
Earlier work this paper cites.
Hyperextended lightface: A facial attribute analysis framework
Sefik Ilkin Serengil and Alper Ozpinar · 2021
Cited alongside, same era.
Videocc3m: Transforming and densifying video captioning datasets
Dongxu Yan, Yitian Wang, Xinting Han, Yixuan Yang, Ming Ding, and Shih-Fu Wen · 2021
Cited alongside, same era.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Yuan Lu, Youngjae Yu, Ronan LeBras, and Yejin Choi · 2021
Cited alongside, same era.
Yt-temporal-180m: Large-scale video dataset for temporal reasoning
Mohammadreza Zolfaghari, Sadaf Singh, Ashish Jain, and Thomas Brox · 2021
Cited alongside, same era.
Alphapose: Whole-body regional multi-person pose estimation and tracking in real-time
Hao-Shu Fang, Jiefeng Li, Hongyang Tang, Chao Xu, Haoyi Zhu, Yuliang Xiu, Yong-Lu Li, and Cewu Lu · 2022
Cited alongside, same era.
Mops: Modular story premise synthesis for open-ended automatic story generation, 2024
Yan Ma, Yu Qiao, and Pengfei Liu · 2024
Closest in time.
Yolo-world: Real-time open-vocabulary object detection
Matthias Minderer, Alexey Gritsenko, and Neil Houlsby · 2024
Closest in time.
Gpt-4o: Advanced large language model, 2024
OpenAI · 2024
Closest in time.
Internvl2: Advanced multimodal large language models, 2024
OpenGVLab · 2024
Closest in time.
Yi-large: A large-scale language model, 2024
Author Name or Organization · 2024
Closest in time.
Pika, 2024
Pika · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Piotr Mirowski, Kory W Mathewson, Jaylen Pittman, and Richard Evans · 2022
Cited alongside, same era.
Cogvideo: Large-scale pretraining for text-to-video generation via transformers
Wenyi Hong, Ming Ding, Wendi Zheng, Xinghan Liu, and Jie Tang · 2023
Cited alongside, same era.
Affective and dynamic beam search for story generation
Tenghao Huang, Ehsan Qasemi, Bangzheng Li, He Wang, Faeze Brahman, Muhao Chen, and Snigdha Chaturvedi · 2023
Cited alongside, same era.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al · 2023
Cited alongside, same era.
Lmdeploy: A toolkit for compressing, deploying, and serving llm
LMDeploy · 2023
Cited alongside, same era.
Scaling open-vocabulary object detection
Matthias Minderer, Alexey Gritsenko, and Neil Houlsby · 2023
Cited alongside, same era.
Gpt-3.5: Language models are few-shot learners, 2023a
OpenAI · 2023
Cited alongside, same era.
Reelshort: A platform for creating short-form videos, 2024
ReelShort · 2024
Closest in time.
Grounding dino 1.5: Advance the ”edge” of open-set object detection, 2024
Tianhe Ren, Qing Jiang, Shilong Liu, Zhaoyang Zeng, Wenlong Liu, Han Gao, Hongjie Huang, Zhengyu Ma, Xiaoke Jiang, Yihao Chen, Yuda Xiong, Hao Zhang, Feng Li, Peijun Tang, Kent Yu, and Lei Zhang · 2024
Closest in time.
Gen2, 2024
Runway · 2024
Closest in time.
Vimi: Controllable character video generation model, 2024
SenseTime · 2024
Closest in time.
A benchmark of facial recognition pipelines and co-usability performances of modules
Sefik Serengil and Alper Ozpinar · 2024
Closest in time.
Vidu: A high-performance text-to-video generator, 2024
ShengShu and Tsinghua · 2024
Closest in time.
Shortmax: Watch dramas and shows, 2024
ShortMax · 2024
Closest in time.
Shortstv: The global home of short movies, 2024
ShortsTV · 2024
Closest in time.
The llama 3 herd of models: Scaling open-source language models
John Smith, Jane Doe, and Robert Brown · 2024
Closest in time.
Pllava: Parameter-free llava extension from images to videos for video dense captioning
Lin Xu, Yilin Zhao, Daquan Zhou, et al · 2024
Closest in time.
Flash-vstream: Memory-based real-time understanding for long video streams
Yue Zhang, Ming Liu, Li Chen, et al · 2024
Closest in time.
Open-sora: Democratizing efficient video production for all, March 2024
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You · 2024
Closest in time.