Fetching the paper…
Reading the bibliography…
Current attention algorithms (e.g., self-attention) are stimulus-driven and highlight all the salient objects in an image.
Orienting of attention
Michael I Posner · 1980
Earlier work this paper cites.
The helmholtz machine
Peter Dayan, Geoffrey E Hinton, Radford M Neal, and Richard S Zemel · 1995
Earlier work this paper cites.
Neural mechanisms of selective visual attention
Robert Desimone, John Duncan, et al · 1995
Earlier work this paper cites.
Perception as Bayesian inference
David C Knill and Whitman Richards · 1996
Earlier work this paper cites.
Effects of attention on orientation-tuning functions of single neurons in macaque cortical area v4
Carrie J McAdams and John HR Maunsell · 1999
Earlier work this paper cites.
Visual spatial characterization of macaque v1 neurons
Michael P Sceniak, Michael J Hawken, and Robert Shapley · 2001
Earlier work this paper cites.
Circuits for local and global signal integration in primary visual cortex
Alessandra Angelucci, Jonathan B Levitt, Emma JS Walton, Jean-Michel Hupe, Jean Bullier, and Jennifer S Lund · 2002
Earlier work this paper cites.
Nature and interaction of signals from the receptive field center and surround in macaque v1 neurons
James R Cavanaugh, Wyeth Bair, and J Anthony Movshon · 2002
Earlier work this paper cites.
Top-down influence in early visual processing: a bayesian perspective
Tai Sing Lee · 2002
Earlier work this paper cites.
Attentional modulation strength in cortical area mt depends on stimulus contrast
Julio C Martınez-Trujillo and Stefan Treue · 2002
Earlier work this paper cites.
Analysis and synthesis of visual images in the brain: evidence for pattern theory
T Sing Lee · 2003
Earlier work this paper cites.
Hierarchical bayesian inference in the visual cortex
Tai Sing Lee and David Mumford · 2003
Earlier work this paper cites.
Top-down control of visual attention in object detection
Aude Oliva, Antonio Torralba, Monica S Castelhano, and John M Henderson · 2003
Earlier work this paper cites.
Inference, attention, and decision in a bayesian neural architecture
Angela J Yu and Peter Dayan · 2004
Earlier work this paper cites.
Bayesian inference and attentional modulation in the visual cortex
Rajesh PN Rao · 2005
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
Geoffrey E Hinton and Ruslan R Salakhutdinov · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Theory and dynamics of perceptual bistability
Paul R Schrater and Rashmi Sundareswara · 2006
Earlier work this paper cites.
Vision as bayesian inference: analysis by synthesis?
Alan Yuille and Daniel Kersten · 2006
Earlier work this paper cites.
Sparse coding via thresholding and local competition in neural circuits
Christopher J Rozell, Don H Johnson, Richard G Baraniuk, and Bruno A Olshausen · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
The normalization model of attention
John H Reynolds and David J Heeger · 2009
Earlier work this paper cites.
What and where: A bayesian inference theory of attention
Sharat Chikkerur, Thomas Serre, Cheston Tan, and Tomaso Poggio · 2010
Earlier work this paper cites.
Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion
Pascal Vincent, Hugo Larochelle, Isabelle Lajoie, Yoshua Bengio, Pierre-Antoine Manzagol, and Léon Bottou · 2010
Earlier work this paper cites.
Visual attention: The past 25 years
Marisa Carrasco · 2011
Earlier work this paper cites.
An object-based bayesian framework for top-down visual attention
Ali Borji, Dicky Sihite, and Laurent Itti · 2012
Cited alongside, same era.
The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2012
Cited alongside, same era.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Cited alongside, same era.
Understanding vision: theory, models, and data
Li Zhaoping · 2014
Cited alongside, same era.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Neural networks with recurrent generative feedback
Yujia Huang, James Gornet, Sihui Dai, Zhiding Yu, Tan Nguyen, Doris Tsao, and Anima Anandkumar · 2020
Later among the works it cites.
Meta-amortized variational inference and learning
Mike Wu, Kristy Choi, Noah Goodman, and Stefano Ermon · 2020
Later among the works it cites.
Synthesize then compare: Detecting failures and anomalies for semantic segmentation
Yingda Xia, Yi Zhang, Fengze Liu, Wei Shen, and Alan L Yuille · 2020
Later among the works it cites.
Beit: Bert pre-training of image transformers
Hangbo Bao, Li Dong, and Furu Wei · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Huijuan Xu and Kate Saenko · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Cited alongside, same era.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Gang Chen · 2021
Later among the works it cites.
Convit: Improving vision transformers with soft convolutional inductive biases
Stéphane d’Ascoli, Hugo Touvron, Matthew L Leavitt, Ari S Morcos, Giulio Biroli, and Levent Sagun · 2021
Later among the works it cites.
Rethinking spatial dimensions of vision transformers
Byeongho Heo, Sangdoo Yun, Dongyoon Han, Sanghyuk Chun, Junsuk Choe, and Seong Joon Oh · 2021
Later among the works it cites.
Perceiver io: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Tdaf: Top-down attention framework for vision tasks
Bo Pang, Yizhuo Li, Jiefeng Li, Muchen Li, Hanwen Cao, and Cewu Lu · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Global filter networks for image classification
Yongming Rao, Wenliang Zhao, Zheng Zhu, Jiwen Lu, and Jie Zhou · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Later among the works it cites.
Region similarity representation learning
Tete Xiao, Colorado J Reed, Xiaolong Wang, Kurt Keutzer, and Trevor Darrell · 2021
Later among the works it cites.
An empirical study of training end-to-end vision-and-language transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Pengchuan Zhang, Lu Yuan, Nanyun Peng, et al · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Object discovery and representation networks
Olivier J Hénaff, Skanda Koppula, Evan Shelhamer, Daniel Zoran, Andrew Jaegle, Andrew Zisserman, João Carreira, and Relja Arandjelović · 2022
Later among the works it cites.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
Towards robust vision transformer
Xiaofeng Mao, Gege Qi, Yuefeng Chen, Xiaodan Li, Ranjie Duan, Shaokai Ye, Yuan He, and Hui Xue · 2022
Later among the works it cites.
Open vocabulary semantic segmentation with patch aligned contrastive learning
Jishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang, Ashish Shah, Philip HS Torr, and Ser-Nam Lim · 2022
Later among the works it cites.
Visual attention emerges from recurrent sparse reconstruction
Baifeng Shi, Yale Song, Neel Joshi, Trevor Darrell, and Xin Wang · 2022
Later among the works it cites.
Unsupervised learning of structured representations via closed-loop transcription
Shengbang Tong, Xili Dai, Yubei Chen, Mingyang Li, Zengyi Li, Brent Yi, Yann LeCun, and Yi Ma · 2022
Later among the works it cites.
Groupvit: Semantic segmentation emerges from text supervision
Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, and Xiaolong Wang · 2022
Later among the works it cites.
Understanding the robustness in vision transformers
Daquan Zhou, Zhiding Yu, Enze Xie, Chaowei Xiao, Animashree Anandkumar, Jiashi Feng, and Jose M Alvarez · 2022
Later among the works it cites.