Fetching the paper…
Reading the bibliography…
As transformers are equivariant to the permutation of input tokens, encoding the positional information of tokens is necessary for many tasks.
Engineering applications of noncommutative harmonic analysis: with emphasis on rotation and motion groups
Gregory S Chirikjian · 2000
Earlier work this paper cites.
Transforming auto-encoders
Geoffrey E Hinton, Alex Krizhevsky, and Sida D Wang · 2011
Earlier work this paper cites.
Transformation properties of learned visual representations
Taco S Cohen and Max Welling · 2014
Earlier work this paper cites.
ShapeNet: An Information-Rich 3D Model Repository
Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Structure-from-motion revisited
Johannes Lutz Schönberger and Jan-Michael Frahm · 2016
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Feature pyramid networks for object detection
Tsung-Yi Lin, Piotr Dollár, Ross B. Girshick, Kaiming He, Bharath Hariharan, and Serge J. Belongie · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Interpretable transformations with encoder-decoder networks
Daniel E Worrall, Stephan J Garbin, Daniyar Turmukhambetov, and Gabriel J Brostow · 2017
Earlier work this paper cites.
Explorations in homeomorphic variational auto-encoding
Luca Falorsi, Pim De Haan, Tim R Davidson, Nicola De Cao, Maurice Weiler, Patrick Forré, and Taco S Cohen · 2018
Earlier work this paper cites.
Matrix capsules with em routing
Geoffrey E Hinton, Sara Sabour, and Nicholas Frosst · 2018
Earlier work this paper cites.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm De Vries, Vincent Dumoulin, and Aaron Courville · 2018
Earlier work this paper cites.
Unsupervised geometry-aware representation for 3d human pose estimation
Helge Rhodin, Mathieu Salzmann, and Pascal Fua · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Non-local neural networks
Xiaolong Wang, Ross B. Girshick, Abhinav Gupta, and Kaiming He · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Stereo magnification
Tinghui Zhou, Richard Tucker, John Flynn, Graham Fyffe, and Noah Snavely · 2018
Earlier work this paper cites.
Monocular neural image based rendering with continuous view control
Xu Chen, Jie Song, and Otmar Hilliges · 2019
Earlier work this paper cites.
Gauge equivariant convolutional networks and the icosahedral cnn
Taco Cohen, Maurice Weiler, Berkay Kicanaoglu, and Max Welling · 2019
Earlier work this paper cites.
Stand-alone self-attention in vision models
Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens · 2019
Earlier work this paper cites.
Root mean square layer normalization
Biao Zhang and Rico Sennrich · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Equivariant neural rendering
Emilien Dupont, Miguel Bautista Martin, Alex Colburn, Aditya Sankar, Josh Susskind, and Qi Shan · 2020
Cited alongside, same era.
Epipolar transformers
Yihui He, Rui Yan, Katerina Fragkiadaki, and Shoou-I Yu · 2020
Cited alongside, same era.
Attentive group equivariant convolutional networks
David Romero, Erik Bekkers, Jakub Tomczak, and Mark Hoogendoorn · 2020
Cited alongside, same era.
Stereo radiance fields (SRF): learning view synthesis for sparse views of novel scenes
Julian Chibane, Aayush Bansal, Verica Lazova, and Gerard Pons-Moll · 2021
Cited alongside, same era.
Unsupervised learning of equivariant structure from sequences
Takeru Miyato, Masanori Koyama, and Kenji Fukumizu · 2022
Later among the works it cites.
Learning symmetric embeddings for equivariant world models
Jung Yeon Park, Ondrej Biza, Linfeng Zhao, Jan Willem van de Meent, and Robin Walters · 2022
Later among the works it cites.
Geometric transformer for fast and robust point cloud registration
Zheng Qin, Hao Yu, Changjian Wang, Yulan Guo, Yuxing Peng, and Kai Xu · 2022
Later among the works it cites.
Translating Images into Maps
Avishkar Saha, Oscar Mendez Maldonado, Chris Russell, and Richard Bowden · 2022
Later among the works it cites.
Generalizable patch-based neural rendering
Mohammed Suhail, Carlos Esteves, Leonid Sigal, and Ameesh Makadia · 2022
Later among the works it cites.
A length-extrapolatable transformer
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gauge equivariant mesh cnns: Anisotropic convolutions on geometric graphs
Pim De Haan, Maurice Weiler, Taco Cohen, and Max Welling · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Gauge equivariant transformer
Lingshen He, Yiming Dong, Yisen Wang, Dacheng Tao, and Zhouchen Lin · 2021
Cited alongside, same era.
Infinite nature: Perpetual view generation of natural scenes from a single image
Andrew Liu, Richard Tucker, Varun Jampani, Ameesh Makadia, Noah Snavely, and Angjoo Kanazawa · 2021
Cited alongside, same era.
Vision transformers for dense prediction
René Ranftl, Alexey Bochkovskiy, and Vladlen Koltun · 2021
Cited alongside, same era.
Categorical depth distribution network for monocular 3d object detection
Cody Reading, Ali Harakeh, Julia Chae, and Steven L. Waslander · 2021
Cited alongside, same era.
Later among the works it cites.
Cross-view Transformers for real-time Map-view Semantic Segmentation
Brady Zhou and Philipp Krähenbühl · 2022
Later among the works it cites.
Geometric algebra transformers
Johann Brehmer, Pim De Haan, Sönke Behrends, and Taco Cohen · 2023
Closest in time.
Explicit correspondence matching for generalizable neural radiance fields
Yuedong Chen, Haofei Xu, Qianyi Wu, Chuanxia Zheng, Tat-Jen Cham, and Jianfei Cai · 2023
Closest in time.
Learning to render novel views from wide-baseline stereo pairs
Yilun Du, Cameron Smith, Ayush Tewari, and Vincent Sitzmann · 2023
Closest in time.
Lrm: Large reconstruction model for single image to 3d
Yicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi, Yang Zhou, Difan Liu, Feng Liu, Kalyan Sunkavalli, Trung Bui, and Hao Tan · 2023
Closest in time.
Neural fourier transform: A general approach to equivariant representation learning
Masanori Koyama, Kenji Fukumizu, Kohei Hayashi, and Takeru Miyato · 2023
Closest in time.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Closest in time.
Bevsegformer: Bird’s eye view semantic segmentation from arbitrary camera rigs
Lang Peng, Zhirong Chen, Zhangjie Fu, Pengpeng Liang, and Erkang Cheng · 2023
Closest in time.
RePAST: Relative pose attention scene representation transformer
Aleksandr Safin, Daniel Durckworth, and Mehdi SM Sajjadi · 2023
Closest in time.
Safety-enhanced autonomous driving using interpretable sensor fusion transformer
Hao Shao, Letian Wang, Ruobing Chen, Hongsheng Li, and Yu Liu · 2023
Closest in time.
3dppe: 3d point positional encoding for transformer-based multi-camera 3d object detection
Changyong Shu, Jiajun Deng, Fisher Yu, and Yifan Liu · 2023
Closest in time.
Is attention all that nerf needs?
Mukund Varma, Peihao Wang, Xuxi Chen, Tianlong Chen, Subhashini Venugopalan, and Zhangyang Wang · 2023
Closest in time.
Geometry-biased transformers for novel view synthesis
Naveen Venkat, Mayank Agarwal, Maneesh Singh, and Shubham Tulsiani · 2023
Closest in time.
Exploring object-centric temporal modeling for efficient multi-view 3d object detection
Shihao Wang, Yingfei Liu, Tiancai Wang, Ying Li, and Xiangyu Zhang · 2023
Closest in time.
Novel view synthesis with diffusion models
Daniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho, Andrea Tagliasacchi, and Mohammad Norouzi · 2023
Closest in time.
Unifying flow, stereo and depth estimation
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, Fisher Yu, Dacheng Tao, and Andreas Geiger · 2023
Closest in time.
Triplane meets gaussian splatting: Fast and generalizable single-view 3d reconstruction with transformers
Zi-Xin Zou, Zhipeng Yu, Yuan-Chen Guo, Yangguang Li, Ding Liang, Yan-Pei Cao, and Song-Hai Zhang · 2023
Closest in time.