Fetching the paper…
Reading the bibliography…
The success of large language models (LLMs) has fostered a new research trend of multi-modality large language models (MLLMs), which changes the paradigm of various fields in computer vision.
Deep Unfolding Network for Image Super-Resolution, March 2020
Kai Zhang, Luc Van Gool, and Radu Timofte · 2003
Earlier work this paper cites.
Taming Transformers for High-Resolution Image Synthesis, June 2021
Patrick Esser, Robin Rombach, and Björn Ommer · 2012
Earlier work this paper cites.
Adaptive noise reduction scheme for salt and pepper
Tina Gebreyohannes and Dong-Yoon Kim · 2012
Earlier work this paper cites.
Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network, May 2017
Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, and Wenzhe Shi · 2017
Earlier work this paper cites.
Deep joint rain detection and removal from a single image
Wenhan Yang, Robby T Tan, Jiashi Feng, Jiaying Liu, Zongming Guo, and Shuicheng Yan · 2017
Earlier work this paper cites.
Esrgan: Enhanced super-resolution generative adversarial networks
Xintao Wang, Ke Yu, Shixiang Wu, Jinjin Gu, Yihao Liu, Chao Dong, Yu Qiao, and Chen Change Loy · 2018
Earlier work this paper cites.
The Unreasonable Effectiveness of Deep Features as a Perceptual Metric, April 2018
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
A salt and pepper noise image denoising method based on the generative classification
Bo Fu, Xiaoyang Zhao, Chuanming Song, Ximing Li, and Xianghai Wang · 2019
Earlier work this paper cites.
Defocus deblurring using dual-pixel data
Abdullah Abuolaim and Michael S Brown · 2020
Earlier work this paper cites.
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever · 2020
Earlier work this paper cites.
A Random CNN Sees Objects: One Inductive Bias of CNN and Its Applications, December 2021
Yun-Hao Cao and Jianxin Wu · 2021
Earlier work this paper cites.
Masked Autoencoders Are Scalable Vision Learners, December 2021
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2021
Earlier work this paper cites.
Generating images with sparse representations
Charlie Nash, Jacob Menick, Sander Dieleman, and Peter W Battaglia · 2021
Earlier work this paper cites.
Learning Transferable Visual Models From Natural Language Supervision, February 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Earlier work this paper cites.
Zero-Shot Text-to-Image Generation, February 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Cited alongside, same era.
Learning from Randomly Initialized Neural Network Features, February 2022
Ehsan Amid, Rohan Anil, Wojciech Kotłowski, and Manfred K. Warmuth · 2022
Cited alongside, same era.
BEiT: BERT Pre-Training of Image Transformers, September 2022
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2022
Cited alongside, same era.
MAGMA – Multimodal Augmentation of Generative Models through Adapter-based Finetuning, October 2022
Constantin Eichenberg, Sidney Black, Samuel Weinbach, Letitia Parcalabescu, and Anette Frank · 2022
Cited alongside, same era.
PixelLM: Pixel Reasoning with Large Multimodal Model, December 2023
Zhongwei Ren, Zhicheng Huang, Yunchao Wei, Yao Zhao, Dongmei Fu, Jiashi Feng, and Xiaojie Jin · 2023
Later among the works it cites.
RoFormer: Enhanced Transformer with Rotary Position Embedding, November 2023
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Later among the works it cites.
Jinguo Zhu, Xiaohan Ding, Yixiao Ge, Yuying Ge, Sijie Zhao, Hengshuang Zhao, Xiaohua Wang, and Ying Shan · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2022
Cited alongside, same era.
Autoregressive Image Generation using Residual Quantization, March 2022
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han · 2022
Cited alongside, same era.
Tape: Task-agnostic prior embedding for image restoration
Lin Liu, Lingxi Xie, Xiaopeng Zhang, Shanxin Yuan, Xiangyu Chen, Wengang Zhou, Houqiang Li, and Qi Tian · 2022
Cited alongside, same era.
High-Resolution Image Synthesis with Latent Diffusion Models, April 2022
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Cited alongside, same era.
Planting a SEED of Vision in Large Language Model, August 2023
Yuying Ge, Yixiao Ge, Ziyun Zeng, Xintao Wang, and Ying Shan · 2023
Cited alongside, same era.
Jiasen Lu, Christopher Clark, Sangho Lee, Zichen Zhang, Savya Khosla, Ryan Marten, Derek Hoiem, and Aniruddha Kembhavi · 2023
Cited alongside, same era.
Linearly Mapping from Image to Text Space, March 2023
Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick · 2023
Cited alongside, same era.
Scalable Diffusion Models with Transformers, March 2023
William Peebles and Saining Xie · 2023
Cited alongside, same era.
Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, Xiangwen Kong, Xiangyu Zhang, Kaisheng Ma, and Li Yi · 2024
Closest in time.
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis, March 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach · 2024
Closest in time.
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation, April 2024
Yuying Ge, Sijie Zhao, Jinguo Zhu, Yixiao Ge, Kun Yi, Lin Song, Chen Li, Xiaohan Ding, and Ying Shan · 2024
Closest in time.
The Platonic Representation Hypothesis, May 2024
Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola · 2024
Closest in time.
Unified Language-Vision Pretraining in LLM with Dynamic Discrete Visual Tokenization, March 2024
Yang Jin, Kun Xu, Kun Xu, Liwei Chen, Chao Liao, Jianchao Tan, Quzhe Huang, Bin Chen, Chenyi Lei, An Liu, Chengru Song, Xiaoqiang Lei, Di Zhang, Wenwu Ou, Kun Gai, and Yadong Mu · 2024
Closest in time.
DiffBIR: Towards Blind Image Restoration with Generative Diffusion Prior, April 2024
Xinqi Lin, Jingwen He, Ziyan Chen, Zhaoyang Lyu, Bo Dai, Fanghua Yu, Wanli Ouyang, Yu Qiao, and Chao Dong · 2024
Closest in time.
Kosmos-G: Generating Images in Context with Multimodal Large Language Models, March 2024
Xichen Pan, Li Dong, Shaohan Huang, Zhiliang Peng, Wenhu Chen, and Furu Wei · 2024
Closest in time.
Frozen Transformers in Language Models Are Effective Visual Encoder Layers, May 2024
Ziqi Pang, Ziyang Xie, Yunze Man, and Yu-Xiong Wang · 2024
Closest in time.
Chameleon: Mixed-Modal Early-Fusion Foundation Models, May 2024
Chameleon Team · 2024
Closest in time.