Fetching the paper…
Reading the bibliography…
We present a method for text-driven perpetual view generation -- synthesizing long-term videos of various scenes solely, given an input text prompt describing the scene and camera poses.
An image inpainting technique based on the fast marching method
Alexandru Cristian Telea · 2004
Earlier work this paper cites.
Infinite images: Creating and exploring a large photorealistic virtual space
Biliana Kaneva, Josef Sivic, Antonio Torralba, Shai Avidan, and William T. Freeman · 2010
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Stereo magnification: Learning view synthesis using multiplane images
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely · 2018
Earlier work this paper cites.
Soft rasterizer: A differentiable renderer for image-based 3d reasoning
Shichen Liu, Tianye Li, Weikai Chen, and Hao Li · 2019
Earlier work this paper cites.
Consistent video depth estimation
Xuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen, and Johannes Kopf · 2020
Earlier work this paper cites.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng · 2020
Earlier work this paper cites.
Giraffe: Representing scenes as compositional generative neural feature fields
Michael Niemeyer and Andreas Geiger · 2020
Earlier work this paper cites.
Softmax splatting for video frame interpolation
Simon Niklaus and Feng Liu · 2020
Earlier work this paper cites.
Accelerating 3d deep learning with pytorch3d
Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari · 2020
Earlier work this paper cites.
Graf: Generative radiance fields for 3d-aware image synthesis
Katja Schwarz, Yiyi Liao, Michael Niemeyer, and Andreas Geiger · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Earlier work this paper cites.
Single-view view synthesis with multiplane images
Richard Tucker and Noah Snavely · 2020
Earlier work this paper cites.
Synsin: End-to-end view synthesis from a single image
Olivia Wiles, Georgia Gkioxari, Richard Szeliski, and Justin Johnson · 2020
Earlier work this paper cites.
Multiview neural surface reconstruction by disentangling geometry and appearance
Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman · 2020
Earlier work this paper cites.
pi-gan: Periodic implicit generative adversarial networks for 3d-aware image synthesis
Eric Chan, Marco Monteiro, Petr Kellnhofer, Jiajun Wu, and Gordon Wetzstein · 2021
Earlier work this paper cites.
Unconstrained scene generation with locally conditioned radiance fields
Terrance DeVries, Miguel Angel Bautista, Nitish Srivastava, Graham W Taylor, and Joshua M Susskind · 2021
Earlier work this paper cites.
Stylenerf: A style-based 3d-aware generator for high-resolution image synthesis
Jiatao Gu, Lingjie Liu, Peng Wang, and Christian Theobalt · 2021
Earlier work this paper cites.
Putting nerf on a diet: Semantically consistent few-shot view synthesis
Ajay Jain, Matthew Tancik, and Pieter Abbeel · 2021
Earlier work this paper cites.
Pathdreamer: A world model for indoor navigation
Jing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge, and Peter Anderson · 2021
Cited alongside, same era.
Robust consistent video depth estimation
Johannes Kopf, Xuejian Rong, and Jia-Bin Huang · 2021
Cited alongside, same era.
Infinite nature: Perpetual view generation of natural scenes from a single image
Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, and Angjoo Kanazawa · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Cited alongside, same era.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Cited alongside, same era.
Pixelsynth: Generating a 3d-consistent experience from a single image
Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
René Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler, and Vladlen Koltun · 2022
Later among the works it cites.
Look outside the room: Synthesizing a consistent long-term 3d scene video from a single image
Xuanchi Ren and Xiaolong Wang · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L. Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mahdavi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David Fleet, and Mohammad Norouzi · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xiaoyue Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chris Rockwell, David F. Fouhey, and Justin Johnson · 2021
Cited alongside, same era.
Geometry-free view synthesis: Transformers and no 3d priors, 2021
Robin Rombach, Patrick Esser, and Björn Ommer · 2021
Cited alongside, same era.
pixelNeRF: Neural radiance fields from one or few images
Alex Yu, Vickie Ye, Matthew Tancik, and Angjoo Kanazawa · 2021
Cited alongside, same era.
Consistent depth of moving objects in video
Zhoutong Zhang, Forrester Cole, Richard Tucker, William T Freeman, and Tali Dekel · 2021
Cited alongside, same era.
Omri Avrahami, Ohad Fried, and Dani Lischinski · 2022
Cited alongside, same era.
Text2live: Text-driven layered image and video editing
Omer Bar-Tal, Dolev Ofri-Amar, Rafail Fridman, Yoni Kasten, and Tali Dekel · 2022
Cited alongside, same era.
Gaudi: A neural architect for immersive 3d scene generation
Miguel Angel Bautista, Pengsheng Guo, Samira Abnar, Walter Talbott, Alexander Toshev, Zhuoyuan Chen, Laurent Dinh, Shuangfei Zhai, Hanlin Goh, Daniel Ulbricht, et al · 2022
Cited alongside, same era.
Later among the works it cites.
Stable-dreamfusion: Text-to-3d with stable-diffusion, 2022
Jiaxiang Tang · 2022
Later among the works it cites.
Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation
Jay Zhangjie Wu, Yixiao Ge, Xintao Wang, Weixian Lei, Yuchao Gu, Wynne Hsu, Ying Shan, Xiaohu Qie, and Mike Zheng Shou · 2022
Later among the works it cites.
Sinnerf: Training neural radiance fields on complex scenes from a single image
Dejia Xu, Yifan Jiang, Peihao Wang, Zhiwen Fan, Humphrey Shi, and Zhangyang Wang · 2022
Later among the works it cites.
Clip aesthetic score predictor
LAION AI · 2023
Closest in time.
Pix2video: Video editing using image diffusion
Duygu Ceylan, Chun-Hao Huang, and Niloy J. Mitra · 2023
Closest in time.
Persistent nature: A generative model of unbounded 3d worlds
Lucy Chai, Richard Tucker, Zhengqi Li, Phillip Isola, and Noah Snavely · 2023
Closest in time.
Scenedreamer: Unbounded 3d scene generation from 2d image collections
Zhaoxi Chen, Guangcong Wang, and Ziwei Liu · 2023
Closest in time.
Structure and content-guided video synthesis with diffusion models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis · 2023
Closest in time.
Text2room: Extracting textured 3d meshes from 2d text-to-image models
Lukas Höllein, Ang Cao, Andrew Owens, Justin Johnson, and Matthias Nießner · 2023
Closest in time.
Videofusion: Decomposed diffusion models for high-quality video generation
Zhengxiong Luo, Dayou Chen, Yingya Zhang, Yan Huang, Liang Wang, Yujun Shen, Deli Zhao, Jingren Zhou, and Tieniu Tan · 2023
Closest in time.
Fatezero: Fusing attentions for zero-shot text-based video editing
Chenyang Qi, Xiaodong Cun, Yong Zhang, Chenyang Lei, Xintao Wang, Ying Shan, and Qifeng Chen · 2023
Closest in time.
Text-to-4d dynamic scene generation
Uriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual, Iurii Makarov, Filippos Kokkinos, Naman Goyal, Andrea Vedaldi, Devi Parikh, Justin Johnson, and Yaniv Taigman · 2023
Closest in time.
Consistent view synthesis with pose-guided diffusion models
Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia-Bin Huang, and Johannes Kopf · 2023
Closest in time.
Painting 3d nature in 2d: View synthesis of natural scenes from a single semantic mask
Shangzhan Zhang, Sida Peng, Tianrun Chen, Linzhan Mou, Haotong Lin, Kaicheng Yu, Yiyi Liao, and Xiaowei Zhou · 2023
Closest in time.
Make-a-protagonist: Generic video editing with an ensemble of experts
Yuyang Zhao, Enze Xie, Lanqing Hong, Zhenguo Li, and Gim Hee Lee · 2023
Closest in time.