Fetching the paper…
Reading the bibliography…
We introduce Cosmos-Transfer, a conditional world generation model that can generate world simulations based on multiple spatial control inputs of various modalities such as segmentation, depth, and edge.
Information Retrieval
C.J. van Rijsbergen · 1979
Earlier work this paper cites.
A computational approach to edge detection
John Canny · 1986
Earlier work this paper cites.
Bilateral filtering for gray and color images
Carlo Tomasi and Roberto Manduchi · 1998
Earlier work this paper cites.
Image quality assessment: From error visibility to structural similarity
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli · 2004
Earlier work this paper cites.
Depth map prediction from a single image using a multi-scale deep network
David Eigen, Christian Puhrsch, and Rob Fergus · 2014
Earlier work this paper cites.
Unsupervised learning of depth and ego-motion from video
Tinghui Zhou, Matthew Brown, Noah Snavely, and David G Lowe · 2017
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Semantic image synthesis with spatially-adaptive normalization
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu · 2019
Earlier work this paper cites.
Few-shot video-to-video synthesis
Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Jan Kautz, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Domain stylization: A fast covariance matching framework towards domain adaptation
Aysegul Dundar, Ming-Yu Liu, Zhiding Yu, Ting-Chun Wang, John Zedlewski, , and Jan Kautz · 2020
Earlier work this paper cites.
Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning
Caelan Reed Garrett, Tomás Lozano-Pérez, and Leslie Pack Kaelbling · 2020
Earlier work this paper cites.
World-consistent video-to-video synthesis
Arun Mallya, Ting-Chun Wang, Karan Sapra, and Ming-Yu Liu · 2020
Earlier work this paper cites.
Rl-cyclegan: Reinforcement learning aware simulation-to-real
Kanishka Rao, Chris Harris, Alex Irpan, Sergey Levine, Julian Ibarz, and Mohi Khansari · 2020
Earlier work this paper cites.
You only need adversarial supervision for semantic image synthesis
Vadim Sushko, Edgar Schönfeld, Dan Zhang, Juergen Gall, Bernt Schiele, and Anna Khoreva · 2020
Earlier work this paper cites.
Retinagan: An object-aware approach to sim-to-real transfer
Daniel Ho, Kanishka Rao, Zhuo Xu, Eric Jang, Mohi Khansari, and Yunfei Bai · 2021
Earlier work this paper cites.
Multimodal conditional image synthesis with product-of-experts GANs
Xun Huang, Arun Mallya, Ting-Chun Wang, and Ming-Yu Liu · 2022
Earlier work this paper cites.
Semantic-shape adaptive feature modulation for semantic image synthesis
Zhengyao Lv, Xiaoming Li, Zhenxing Niu, Bing Cao, and Wangmeng Zuo · 2022
Earlier work this paper cites.
Robot learning from randomized simulations: A review
Fabio Muratore, Fabio Ramos, Greg Turk, Wenhao Yu, Michael Gienger, and Jan Peters · 2022
Earlier work this paper cites.
Retrieval-based spatially adaptive normalization for semantic image synthesis
Yupeng Shi, Xiao Liu, Yuxiang Wei, Zhongqin Wu, and Wangmeng Zuo · 2022
Earlier work this paper cites.
Semantic image synthesis via diffusion models
Weilun Wang, Jianmin Bao, Wengang Zhou, Dongdong Chen, Dong Chen, Lu Yuan, and Houqiang Li · 2022
Earlier work this paper cites.
Fast-vid2vid: Spatial-temporal compression for video-to-video synthesis
Long Zhuo, Guangcong Wang, Shikai Li, Wayne Wu, and Ziwei Liu · 2022
Earlier work this paper cites.
Universal guidance for diffusion models
Arpit Bansal, Hong-Min Chu, Avi Schwarzschild, Soumyadip Sengupta, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Cited alongside, same era.
Training-free layout control with cross-attention guidance
Minghao Chen, Iro Laina, and Andrea Vedaldi · 2023
Cited alongside, same era.
Shortcut-v2v: Compression framework for video-to-video translation based on temporal redundancy reduction
Chaeyeon Chung, Yeojeong Park, Seunghwan Choi, Munkhsoyol Ganbat, and Jaegul Choo · 2023
Cited alongside, same era.
Structure and content-guided video synthesis with diffusion models
Patrick Esser, Johnathan Chiu, Parmida Atighehchian, Jonathan Granskog, and Anastasis Germanidis · 2023
Cited alongside, same era.
Composer: Creative and controllable image synthesis with composable conditions
Lianghua Huang, Di Chen, Yu Liu, Yujun Shen, Deli Zhao, and Jingren Zhou · 2023
Junsong Chen, Yue Wu, Simian Luo, Enze Xie, Sayak Paul, Ping Luo, Hang Zhao, and Zhenguo Li · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team · 2024
Later among the works it cites.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al · 2024
Later among the works it cites.
Ego-Exo4D: Understanding skilled human activity from first- and third-person perspectives, 2024
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Humansd: A native skeleton-guided diffusion model for human image generation
Xuan Ju, Ailing Zeng, Chenchen Zhao, Jianan Wang, Lei Zhang, and Qiang Xu · 2023
Cited alongside, same era.
Gligen: Open-set grounded text-to-image generation
Yuheng Li, Haotian Liu, Qingyang Wu, Fangzhou Mu, Jianwei Yang, Jianfeng Gao, Chunyuan Li, and Yong Jae Lee · 2023
Cited alongside, same era.
Hyperhuman: Hyper-realistic human generation with latent structural diffusion
Xian Liu, Jian Ren, Aliaksandr Siarohin, Ivan Skorokhodov, Yanyu Li, Dahua Lin, Xihui Liu, Ziwei Liu, and Sergey Tulyakov · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Cited alongside, same era.
Scenario diffusion: Controllable driving scenario generation with diffusion
Ethan Pronovost, Meghana Reddy Ganesina, Noureldin Hendy, Zeyu Wang, Andres Morales, Kai Wang, and Nick Roy · 2023
Cited alongside, same era.
Unicontrol: A unified diffusion model for controllable visual generation in the wild
Can Qin, Shu Zhang, Ning Yu, Yihao Feng, Xinyi Yang, Yingbo Zhou, Huan Wang, Juan Carlos Niebles, Caiming Xiong, Silvio Savarese, et al · 2023
Cited alongside, same era.
Cutlass, 2023
Vijay Thakkar, Pradeep Ramani, Cris Cecka, Aniket Shivam, Honghao Lu, Ethan Yan, Jack Kosaian, Mark Hoemmen, Haicheng Wu, Andrew Kerr, et al · 2023
Cited alongside, same era.
Peekaboo: Interactive video generation via masked-diffusion
Yash Jain, Anshul Nasery, Vibhav Vineet, and Harkirat Behl · 2024
Later among the works it cites.
Hydra-mdp: End-to-end multimodal planning with multi-target hydra-distillation
Zhenxin Li, Kailin Li, Shihao Wang, Shiyi Lan, Zhiding Yu, Yishen Ji, Zhiqi Li, Ziyue Zhu, Jan Kautz, Zuxuan Wu, et al · 2024
Later among the works it cites.
Han Lin, Jaemin Cho, Abhay Zala, and Mohit Bansal · 2024
Later among the works it cites.
Yifan Lu, Xuanchi Ren, Jiawei Yang, Tianchang Shen, Zhangjie Wu, Jun Gao, Yue Wang, Siheng Chen, Mike Chen, Sanja Fidler, and Jiahui Huang · 2024
Later among the works it cites.
Place: Adaptive layout-semantic fusion for semantic image synthesis
Zhengyao Lv, Yuxiang Wei, Wangmeng Zuo, and Kwan-Yee K Wong · 2024
Later among the works it cites.
T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models
Chong Mou, Xintao Wang, Liangbin Xie, Yanze Wu, Jian Zhang, Zhongang Qi, and Ying Shan · 2024
Later among the works it cites.
Genie 2: A large-scale foundation world model
Jack Parker-Holder, Philip Ball, Jake Bruce, Vibhavari Dasagi, Kristian Holsheimer, Christos Kaplanis, Alexandre Moufarek, Guy Scully, Jeremy Shar, Jimmy Shi, et al · 2024
Later among the works it cites.
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, et al · 2024
Later among the works it cites.
Anycontrol: create your artwork with versatile control on text-to-image generation
Yanan Sun, Yanchen Liu, Yinhao Tang, Wenjie Pei, and Kai Chen · 2024
Later among the works it cites.
Exploring generative ai for sim2real in driving data synthesis
Haonan Zhao, Yiting Wang, Thomas Bashford-Rogers, Valentina Donzella, and Kurt Debattista · 2024
Later among the works it cites.
Fast-vid2vid++: Spatial-temporal distillation for real-time video-to-video synthesis
Long Zhuo, Guangcong Wang, Shikai Li, Wayne Wu, and Ziwei Liu · 2024
Later among the works it cites.
AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems, 2025
AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xu Huang, Shu Jiang, et al · 2025
Closest in time.
Semantic image synthesis via class-adaptive cross-attention
Tomaso Fontanini, Claudio Ferrari, Giuseppe Lisanti, Massimo Bertozzi, and Andrea Prati · 2025
Closest in time.
Data scaling laws in imitation learning for robotic manipulation, 2025
Fanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen, Jiacheng You, and Yang Gao · 2025
Closest in time.
Cosmos world foundation model platform for physical ai
NVIDIA · 2025
Closest in time.
Gen3c: 3d-informed world-consistent video generation with precise camera control
Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier-David, Thomas Müller, Alexander Keller, Sanja Fidler, and Jun Gao · 2025
Closest in time.