Fetching the paper…
Reading the bibliography…
We address the task of generating temporally consistent and physically plausible images of actions and object state transformations.
Error-tolerant image compositing
Michael W Tao, Micah K Johnson, and Sylvain Paris · 2013
Earlier work this paper cites.
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero · 2014
Earlier work this paper cites.
Deep multi-scale video prediction beyond mean square error
Michael Mathieu, Camille Couprie, and Yann LeCun · 2015
Earlier work this paper cites.
Convolutional lstm network: A machine learning approach for precipitation nowcasting
Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros · 2017
Earlier work this paper cites.
Dual motion gan for future-flow embedded video prediction
Xiaodan Liang, Lisa Lee, Wei Dai, and Eric P Xing · 2017
Earlier work this paper cites.
Deep predictive coding networks for video prediction and unsupervised learning
William Lotter, Gabriel Kreiman, and David Cox · 2017
Earlier work this paper cites.
Conditional image synthesis with auxiliary classifier gans
Augustus Odena, Christopher Olah, and Jonathon Shlens · 2017
Earlier work this paper cites.
Unpaired image-to-image translation using cycle-consistent adversarial networks
Jun-Yan Zhu, Taesung Park, Phillip Isola, and Alexei A Efros · 2017
Earlier work this paper cites.
Toward multimodal image-to-image translation
Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A Efros, Oliver Wang, and Eli Shechtman · 2017
Earlier work this paper cites.
Stargan: Unified generative adversarial networks for multi-domain image-to-image translation
Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo · 2018
Earlier work this paper cites.
Visual reinforcement learning with imagined goals
Ashvin V Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Earlier work this paper cites.
Zero-shot visual imitation
Deepak Pathak, Parsa Mahmoudieh, Guanghao Luo, Pulkit Agrawal, Dian Chen, Yide Shentu, Evan Shelhamer, Jitendra Malik, Alexei A. Efros, and Trevor Darrell · 2018
Earlier work this paper cites.
High-resolution image synthesis and semantic manipulation with conditional gans
Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Andrew Tao, Jan Kautz, and Bryan Catanzaro · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He · 2018
Earlier work this paper cites.
Toward realistic image compositing with adversarial learning
Bor-Chun Chen and Andrew Kae · 2019
Earlier work this paper cites.
Predicting future frames using retrospective cycle gan
Yong-Hoon Kwon and Min-Gyu Park · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Planning with goal-conditioned policies
Soroush Nasiriany, Vitchyr Pong, Steven Lin, and Sergey Levine · 2019
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu · 2019
Earlier work this paper cites.
Coin: A large-scale dataset for comprehensive instructional video analysis
Yansong Tang, Dajun Ding, Yongming Rao, Yu Zheng, Danyang Zhang, Lili Zhao, Jiwen Lu, and Jie Zhou · 2019
Earlier work this paper cites.
Action modifiers: Learning from adverbs in instructional videos
Hazel Doughty, Ivan Laptev, Walterio Mayol-Cuevas, and Dima Damen · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Contextual imagined goals for self-supervised robotic learning
Ashvin Nair, Shikhar Bahl, Alexander Khazatsky, Vitchyr Pong, Glen Berseth, and Sergey Levine · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Earlier work this paper cites.
Cross-domain correspondence learning for exemplar-based image translation
Pan Zhang, Bo Zhang, Dong Chen, Lu Yuan, and Fang Wen · 2020
Cited alongside, same era.
Learning generalizable robotic reward functions from “in-the-wild” human videos
Annie S Chen, Suraj Nair, and Chelsea Finn · 2021
Cited alongside, same era.
Pre-trained image processing transformer
Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao · 2021
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Learning temporal dynamics from cycles in narrated video
Dave Epstein, Jiajun Wu, Cordelia Schmid, and Chen Sun · 2021
Cited alongside, same era.
Taming transformers for high-resolution image synthesis
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Later among the works it cites.
Palette: Image-to-image diffusion models
Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al · 2022
Later among the works it cites.
Look for the change: Learning object states and state-modifying actions from untrimmed web videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev, and Josef Sivic · 2022
Later among the works it cites.
Pretraining is all you need for image-to-image translation
Tengfei Wang, Ting Zhang, Bo Zhang, Hao Ouyang, Dong Chen, Qifeng Chen, and Fang Wen · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Cited alongside, same era.
Intrinsic image harmonization
Zonghui Guo, Haiyong Zheng, Yufeng Jiang, Zhaorui Gu, and Bing Zheng · 2021
Cited alongside, same era.
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans · 2021
Cited alongside, same era.
Variational diffusion models
Diederik Kingma, Tim Salimans, Ben Poole, and Jonathan Ho · 2021
Cited alongside, same era.
Future frame prediction network for video anomaly detection
Weixin Luo, Wen Liu, Dongze Lian, and Shenghua Gao · 2021
Cited alongside, same era.
Sdedit: Guided image synthesis and editing with stochastic differential equations
Chenlin Meng, Yutong He, Yang Song, Jiaming Song, Jiajun Wu, Jun-Yan Zhu, and Stefano Ermon · 2021
Cited alongside, same era.
Learning graph embeddings for compositional zero-shot learning
Muhammad Ferjad Naeem, Yongqin Xian, Federico Tombari, and Zeynep Akata · 2021
Cited alongside, same era.
P3iv: Probabilistic procedure planning from instructional videos with weak supervision
He Zhao, Isma Hadji, Nikita Dvornik, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson · 2022
Later among the works it cites.
Universal guidance for diffusion models
Arpit Bansal, Hong-Min Chu, Avi Schwarzschild, Soumyadip Sengupta, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Closest in time.
Instructpix2pix: Learning to follow image editing instructions
Tim Brooks, Aleksander Holynski, and Alexei A. Efros · 2023
Closest in time.
Learning video-conditioned policies for unseen manipulation tasks
Elliot Chane-Sane, Cordelia Schmid, and Ivan Laptev · 2023
Closest in time.
Learning universal policies via text-guided video generation
Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Joshua B Tenenbaum, Dale Schuurmans, and Pieter Abbeel · 2023
Closest in time.
Stepformer: Self-supervised step discovery and localization in instructional videos
Nikita Dvornik, Isma Hadji, Ran Zhang, Konstantinos G Derpanis, Richard P Wildes, and Allan D Jepson · 2023
Closest in time.
An edit friendly ddpm noise space: Inversion and manipulations
Inbar Huberman-Spiegelglas, Vladimir Kulikov, and Tomer Michaeli · 2023
Closest in time.
Imagic: Text-based real image editing with diffusion models
Bahjat Kawar, Shiran Zada, Oran Lang, Omer Tov, Huiwen Chang, Tali Dekel, Inbar Mosseri, and Michal Irani · 2023
Closest in time.
Putting people in their place: Affordance-aware human insertion into scenes
Sumith Kulal, Tim Brooks, Alex Aiken, Jiajun Wu, Jimei Yang, Jingwan Lu, Alexei A Efros, and Krishna Kumar Singh · 2023
Closest in time.
Zero-shot image-to-image translation
Gaurav Parmar, Krishna Kumar Singh, Richard Zhang, Yijun Li, Jingwan Lu, and Jun-Yan Zhu · 2023
Closest in time.
Localizing object-level shape variations with text-to-image diffusion models
Or Patashnik, Daniel Garibi, Idan Azuri, Hadar Averbuch-Elor, and Daniel Cohen-Or · 2023
Closest in time.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman · 2023
Closest in time.
Chop & learn: Recognizing and generating object-state compositions
Nirat Saini, Hanyu Wang, Archana Swaminathan, Vinoj Jayasundara, Bo He, Kamal Gupta, and Abhinav Shrivastava · 2023
Closest in time.
Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models
George Stein et al · 2023
Closest in time.
Plug-and-play diffusion features for text-driven image-to-image translation
Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel · 2023
Closest in time.
Manipulate by seeing: Creating manipulation controllers from pre-trained representations
Jianren Wang, Sudeep Dasari, Mohan Kumar Srirama, Shubham Tulsiani, and Abhinav Gupta · 2023
Closest in time.
Feature prediction diffusion model for video anomaly detection
Cheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang, and Wenjun Wang · 2023
Closest in time.
Paint by example: Exemplar-based image editing with diffusion models
Binxin Yang, Shuyang Gu, Bo Zhang, Ting Zhang, Xuejin Chen, Xiaoyan Sun, Dong Chen, and Fang Wen · 2023
Closest in time.
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala · 2023
Closest in time.
Learning procedure-aware video representation from instructional videos and their narrations
Yiwu Zhong, Licheng Yu, Yang Bai, Shangwen Li, Xueting Yan, and Yin Li · 2023
Closest in time.
Multi-task learning of object states and state-modifying actions from web videos
Tomáš Souček, Jean-Baptiste Alayrac, Antoine Miech, Ivan Laptev, and Josef Sivic · 2024
Closest in time.