Fetching the paper…
Reading the bibliography…
We explore a novel video creation experience, namely Video Creation by Demonstration.
Adam: A method for stochastic optimization
D. P. Kingma · 2015
Earlier work this paper cites.
Quo vadis, action recognition? A new model and the Kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
The “something something” video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
Scaling egocentric vision: The EPIC-KITCHENS dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Earlier work this paper cites.
Towards accurate generative models of video: A new metric & challenges
T. Unterthiner, S. Van Steenkiste, K. Kurach, R. Marinier, M. Michalski, and S. Gelly · 2018
Earlier work this paper cites.
Learning to forecast and refine residual motion for image-to-video generation
L. Zhao, X. Peng, Y. Tian, M. Kapadia, and D. Metaxas · 2018
Earlier work this paper cites.
First order motion model for image animation
A. Siarohin, S. Lathuilière, S. Tulyakov, E. Ricci, and N. Sebe · 2019
Earlier work this paper cites.
Talking face generation by conditional recurrent adversarial network
Y. Song, J. Zhu, D. Li, X. Wang, and H. Qi · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
ViViT: A video vision transformer
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Zero-shot text-to-image generation
A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever · 2021
Earlier work this paper cites.
Flow guided transformable bottleneck networks for motion retargeting
J. Ren, M. Chai, O. J. Woodford, K. Olszewski, and S. Tulyakov · 2021
Earlier work this paper cites.
Motion representations for articulated animation
A. Siarohin, O. J. Woodford, J. Ren, M. Chai, and S. Tulyakov · 2021
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole · 2021
Earlier work this paper cites.
RT-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Controllable video generation through global and local motion dynamics
A. Davtyan and P. Favaro · 2022
Cited alongside, same era.
Ego4D: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Cited alongside, same era.
Show me what and tell me how: Video synthesis via multimodal conditioning
L. Han, J. Ren, H.-Y. Lee, F. Barbieri, K. Olszewski, S. Minaee, D. Metaxas, and S. Tulyakov · 2022
Cited alongside, same era.
Flexible diffusion modeling of long videos
W. Harvey, S. Naderiparizi, V. Masrani, C. Weilbach, and F. Wood · 2022
Cited alongside, same era.
Classifier-free diffusion guidance
J. Ho and T. Salimans · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Video generation models as world simulators
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh · 2024
Closest in time.
Genie: Generative interactive environments
J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al · 2024
Closest in time.
MagicPose: Realistic human poses and facial expressions retargeting with identity-aware diffusion
D. Chang, Y. Shi, Q. Gao, H. Xu, J. Fu, G. Song, Q. Yan, Y. Zhu, X. Yang, and M. Soleymani · 2024
Closest in time.
Photorealistic video generation with diffusion models
A. Gupta, L. Yu, K. Sohn, X. Gu, M. Hahn, F.-F. Li, I. Essa, L. Jiang, and J. Lezama · 2024
Closest in time.
Animate Anyone: Consistent and controllable image-to-video synthesis for character animation
L. Hu · 2024
Closest in time.
Image Conductor: Precision control for interactive video synthesis
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
Layered controllable video generation
J. Huang, Y. Jin, K. M. Yi, and L. Sigal · 2022
Cited alongside, same era.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Cited alongside, same era.
Progressive distillation for fast sampling of diffusion models
T. Salimans and J. Ho · 2022
Cited alongside, same era.
InternVideo: General video foundation models via generative and discriminative learning
Y. Wang, K. Li, Y. Li, Y. He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wang, et al · 2022
Cited alongside, same era.
Align your latents: High-resolution video synthesis with latent diffusion models
A. Blattmann, R. Rombach, H. Ling, T. Dockhorn, S. W. Kim, S. Fidler, and K. Kreis · 2023
Cited alongside, same era.
GAIA-1: A generative world model for autonomous driving
A. Hu, L. Russell, H. Yeo, Z. Murez, G. Fedoseev, A. Kendall, J. Shotton, and G. Corrado · 2023
Cited alongside, same era.
Y. Li, X. Wang, Z. Zhang, Z. Wang, Z. Yuan, L. Xie, Y. Zou, and Y. Shan · 2024
Closest in time.
Towards world simulator: Crafting physical commonsense-based benchmark for video generation
F. Meng, J. Liao, X. Tan, W. Shao, Q. Lu, K. Zhang, Y. Cheng, D. Li, Y. Qiao, and P. Luo · 2024
Closest in time.
Spectral motion alignment for video motion transfer using diffusion models
G. Y. Park, H. Jeong, S. W. Lee, and J. C. Ye · 2024
Closest in time.
Movie Gen: A cast of media foundation models
A. Polyak, A. Zohar, A. Brown, A. Tjandra, A. Sinha, A. Lee, A. Vyas, B. Shi, C.-Y. Ma, C.-Y. Chuang, et al · 2024
Closest in time.
Diffusion models are real-time game engines
D. Valevski, Y. Leviathan, M. Arar, and S. Fruchter · 2024
Closest in time.
DragAnything: Motion control for anything using entity representation
W. Wu, Z. Li, Y. Gu, R. Zhao, Y. He, D. J. Zhang, M. Z. Shou, Y. Li, T. Gao, and D. Zhang · 2024
Closest in time.
Pandora: Towards general world model with natural language actions and video states
J. Xiang, G. Liu, Y. Gu, Q. Gao, Y. Ning, Y. Zha, Z. Feng, T. Tao, S. Hao, Y. Shi, et al · 2024
Closest in time.
Video diffusion models are training-free motion interpreter and controller
Z. Xiao, Y. Zhou, S. Yang, and X. Pan · 2024
Closest in time.
Learning interactive real-world simulators
M. Yang, Y. Du, K. Ghasemipour, J. Tompson, D. Schuurmans, and P. Abbeel · 2024
Closest in time.
VideoGLUE: Video general understanding evaluation of foundation models
L. Yuan, N. B. Gundavarapu, L. Zhao, H. Zhou, Y. Cui, L. Jiang, X. Yang, M. Jia, T. Weyand, L. Friedman, et al · 2024
Closest in time.