Fetching the paper…
Reading the bibliography…
We present a novel approach for generating 360-degree high-quality, spatio-temporally coherent human videos from a single image.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero. 2014 · 2014
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation. In
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Deformable gans for pose-based human image generation. In
Aliaksandr Siarohin, Enver Sangineto, Stéphane Lathuiliere, and Nicu Sebe. 2018 · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Appearance and pose-conditioned human image generation using deformable gans
Aliaksandr Siarohin, Stéphane Lathuilière, Enver Sangineto, and Nicu Sebe. 2019a · 2019
Earlier work this paper cites.
First order motion model for image animation
Aliaksandr Siarohin, Stéphane Lathuilière, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019b · 2019
Earlier work this paper cites.
AIST Dance Video Database: Multi-genre, Multi-dancer, and Multi-camera Database for Dance Information Processing. In
Shuhei Tsuchida, Satoru Fukayama, Masahiro Hamasaki, and Masataka Goto. 2019 · 2019
Earlier work this paper cites.
G3AN: Disentangling appearance and motion for video generation. In
Yaohui Wang, Piotr Bilinski, Francois Bremond, and Antitza Dantcheva. 2020 · 2020
Earlier work this paper cites.
Learning high fidelity depths of dressed humans by watching social media dance videos. In
Yasamin Jafarian and Hyun Soo Park. 2021 · 2021
Earlier work this paper cites.
Segmenter: Transformer for semantic segmentation. In
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. 2021 · 2021
Earlier work this paper cites.
A good image generator is what you need for high-resolution video synthesis
Yu Tian, Jian Ren, Menglei Chai, Kyle Olszewski, Xi Peng, Dimitris N Metaxas, and Sergey Tulyakov. 2021 · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention. In
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. 2021 · 2021
Earlier work this paper cites.
One-shot free-view neural talking-head synthesis for video conferencing. In
Ting-Chun Wang, Arun Mallya, and Ming-Yu Liu. 2021 · 2021
Earlier work this paper cites.
SegFormer: Simple and efficient design for semantic segmentation with transformers
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. 2021 · 2021
Earlier work this paper cites.
Function4D: Real-time Human Volumetric Capture from Very Sparse Consumer RGBD Sensors. In
Tao Yu, Zerong Zheng, Kaiwen Guo, Pengpeng Liu, Qionghai Dai, and Yebin Liu. 2021 · 2021
Cited alongside, same era.
Tokens-to-token vit: Training vision transformers from scratch on imagenet. In
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zi-Hang Jiang, Francis EH Tay, Jiashi Feng, and Shuicheng Yan. 2021 · 2021
Cited alongside, same era.
Light-weight Multi-person Total Capture Using Sparse Multi-view Cameras. In
Yuxiang Zhang, Zhe Li, Liang An, Mengcheng Li, Tao Yu, and Yebin Liu. 2021 · 2021
Cited alongside, same era.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al · 2021
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models. In
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Cited alongside, same era.
Wonder3d: Single image to 3d using cross-domain diffusion
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al · 2023
Later among the works it cites.
SMPL: A skinned multi-person linear model
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2023 · 2023
Later among the works it cites.
Vdt: General-purpose video diffusion transformers via mask modeling. In
Haoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo, Zhiwu Lu, Ping Luo, and Mingyu Ding. 2023 · 2023
Later among the works it cites.
Scalable diffusion models with transformers. In
William Peebles and Saining Xie. 2023 · 2023
Later among the works it cites.
Zero123++: a single image to consistent multi-view diffusion base model
Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. 2023a · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Twindom 3D Avatar Dataset
Twindom. 2022 · 2022
Cited alongside, same era.
All are worth words: A vit backbone for diffusion models. In
Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu. 2023 · 2023
Cited alongside, same era.
Person image synthesis via denoising diffusion model. In
Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, and Fahad Shahbaz Khan. 2023 · 2023
Cited alongside, same era.
Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion. In
Michael J Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang. 2023 · 2023
Cited alongside, same era.
Dna-rendering: A diverse neural actor repository for high-fidelity human-centric rendering. In
Wei Cheng, Ruixiang Chen, Siming Fan, Wanqi Yin, Keyu Chen, Zhongang Cai, Jingbo Wang, Yang Gao, Zhengming Yu, Zhengyu Lin, et al · 2023
Cited alongside, same era.
Humans in 4D: Reconstructing and Tracking Humans with Transformers. In
Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa*, and Jitendra Malik*. 2023 · 2023
Cited alongside, same era.
Animate anyone: Consistent and controllable image-to-video synthesis for character animation
Li Hu, Xin Gao, Peng Zhang, Ke Sun, Bang Zhang, and Liefeng Bo. 2023 · 2023
Cited alongside, same era.
Mvdream: Multi-view diffusion for 3d generation
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. 2023b · 2023
Later among the works it cites.
Disco: Disentangled control for referring human dance generation in real world
Tan Wang, Linjie Li, Kevin Lin, Chung-Ching Lin, Zhengyuan Yang, Hanwang Zhang, Zicheng Liu, and Lijuan Wang. 2023 · 2023
Later among the works it cites.
Magicanimate: Temporally consistent human image animation using diffusion model
Zhongcong Xu, Jianfeng Zhang, Jun Hao Liew, Hanshu Yan, Jia-Wei Liu, Chenxu Zhang, Jiashi Feng, and Mike Zheng Shou. 2023 · 2023
Later among the works it cites.
Generating Holistic 3D Human Motion from Speech. In
Hongwei Yi, Hualin Liang, Yifei Liu, Qiong Cao, Yandong Wen, Timo Bolkart, Dacheng Tao, and Michael J Black. 2023 · 2023
Later among the works it cites.
Fast training of diffusion models with masked transformers
Hongkai Zheng, Weili Nie, Arash Vahdat, and Anima Anandkumar. 2023 · 2023
Later among the works it cites.
CameraCtrl: Enabling Camera Control for Text-to-Video Generation
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. 2024 · 2024
Closest in time.
Motion-x: A large-scale 3d expressive whole-body human motion dataset
Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. 2024 · 2024
Closest in time.
Latte: Latent diffusion transformer for video generation
Xin Ma, Yaohui Wang, Gengyun Jia, Xinyuan Chen, Ziwei Liu, Yuan-Fang Li, Cunjian Chen, and Yu Qiao. 2024 · 2024
Closest in time.
Video generation models as world simulators
OpenAI. 2024 · 2024
Closest in time.
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion
Shiyuan Yang, Liang Hou, Haibin Huang, Chongyang Ma, Pengfei Wan, Di Zhang, Xiaodong Chen, and Jing Liao. 2024 · 2024
Closest in time.
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu. 2024 · 2024
Closest in time.