The perceptron, a perceiving and recognizing automaton Project Para
F. Rosenblatt · 1957
Earlier work this paper cites.
Principles of neurodynamics. perceptrons and the theory of brain mechanisms
F. ROSENBLATT · 1961
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart et al · 1985
Earlier work this paper cites.
Autoencoders, minimum description length, and helmholtz free energy
G. E. Hinton and R. S. Zemel · 1994
Earlier work this paper cites.
High-level vision: Object recognition and visual cognition
S. Ullman et al · 1996
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun et al · 1998
Earlier work this paper cites.
The average distances in random graphs with given expected degrees
F. Chung and L. Lu · 2002
Earlier work this paper cites.
Perceptual organization in vision: Behavioral and neural perspectives
R. Kimchi et al · 2003
Earlier work this paper cites.
A non-local algorithm for image denoising
A. Buades et al · 2005
Earlier work this paper cites.
Model compression
C. Buciluǎ et al · 2006
Earlier work this paper cites.
Cvonline: The evolving, distributed, non-proprietary, on-line compendium of computer vision
R. B. Fisher · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent et al · 2008
Earlier work this paper cites.
What are they doing?: Collective activity classification using spatio-temporal relationship among people
W. Choi et al · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Improving the speed of neural networks on cpus
V. Vanhoucke et al · 2011
Earlier work this paper cites.
Spectral sparsification of graphs
D. A. Spielman and S.-H. Teng · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky et al · 2012
Earlier work this paper cites.
TransTrack: Multiple Object Tracking with Transformer
Original
P. Sun et al · 2012
Earlier work this paper cites.
Top-down saliency detection via contextual pooling
J. Zhu et al · 2014
Earlier work this paper cites.
Do deep nets really need to be deep?
J. Ba and R. Caruana · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
R. Livni et al · 2014
Earlier work this paper cites.
Diannao: a small-footprint high-throughput accelerator for ubiquitous machine-learning
T. Chen et al · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Original
J. Chung et al · 2014
Earlier work this paper cites.
Multiple object recognition with visual attention
J. Ba et al · 2014
Earlier work this paper cites.
Recurrent models of visual attention
V. Mnih et al · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin et al · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau et al · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren et al · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
J. Long et al · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Original
G. Hinton et al · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu et al · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu et al · 2015
Earlier work this paper cites.
A decomposable attention model for natural language inference
A. Parikh et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He et al · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Original
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Layer normalization
Original
J. L. Ba et al · 2016
Earlier work this paper cites.
Conditional image generation with pixelcnn decoders
Original
A. v. d. Oord et al · 2016
Earlier work this paper cites.
Context encoders: Feature learning by inpainting
D. Pathak et al · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani et al · 2017
Earlier work this paper cites.
Convolutional sequence to sequence learning
J. Gehring et al · 2017
Earlier work this paper cites.
Focal loss for dense object detection
T.-Y. Lin et al · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
C. R. Qi et al · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi et al · 2017
Earlier work this paper cites.
Neural discrete representation learning
A. v. d. Oord et al · 2017
Earlier work this paper cites.
Two-stream transformer networks for video-based face alignment
H. Liu et al · 2017
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming
Z. Liu et al · 2017
Earlier work this paper cites.
Residual attention network for image classification
F. Wang et al · 2017
Earlier work this paper cites.
End-to-end dense video captioning with masked transformer
L. Zhou et al · 2018
Earlier work this paper cites.
Image transformer
N. Parmar et al · 2018
Earlier work this paper cites.
Self-attention with relative position representations
P. Shaw et al · 2018
Earlier work this paper cites.
Cascade r-cnn: Delving into high quality object detection
Z. Cai and N. Vasconcelos · 2018
Earlier work this paper cites.
Graph r-cnn for scene graph generation
J. Yang et al · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
J. Frankle and M. Carbin · 2018
Earlier work this paper cites.
Non-local neural networks
X. Wang et al · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
A. Radford et al · 2018
Earlier work this paper cites.
Learn to pay attention
S. Jetley et al · 2018
Earlier work this paper cites.
Attribute-aware attention model for fine-grained representation learning
K. Han et al · 2018
Earlier work this paper cites.
Diagnose like a radiologist: Attention guided convolutional neural network for thorax disease classification
Original
Q. Guan et al · 2018
Earlier work this paper cites.
Squeeze-and-excitation networks
J. Hu et al · 2018
Earlier work this paper cites.
Psanet: Point-wise spatial attention network for scene parsing
H. Zhao et al · 2018
Earlier work this paper cites.
Attention u-net: Learning where to look for the pancreas
O. Oktay et al · 2018
Earlier work this paper cites.
Compact generalized non-local network
K. Yue et al · 2018
Earlier work this paper cites.
Beyond grids: Learning graph representations for visual recognition
Y. Li and A. Gupta · 2018
Earlier work this paper cites.
Symbolic graph reasoning meets convolutions
X. Liang et al · 2018
Earlier work this paper cites.
Relation networks for object detection
H. Hu et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin et al · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
P. Michel et al · 2019
Earlier work this paper cites.
Adaptive input representations for neural language modeling
A. Baevski and M. Auli · 2019
Earlier work this paper cites.
Learning deep transformer models for machine translation
Q. Wang et al · 2019
Earlier work this paper cites.
Understanding and improving layer normalization
J. Xu et al · 2019
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford et al · 2019
Earlier work this paper cites.
Fcos: Fully convolutional one-stage object detection
Z. Tian et al · 2019
Earlier work this paper cites.
Video action transformer network
R. Girdhar et al · 2019
Earlier work this paper cites.
Temporal transformer networks: Joint learning of invariant and discriminative time warping
S. Lohit et al · 2019
Earlier work this paper cites.
Video multitask transformer network
H. Seong et al · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
C. Sun et al · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Original
L. H. Li et al · 2019
Earlier work this paper cites.
Q8bert: Quantized 8bit bert
Original
O. Zafrir et al · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Original
V. Sanh et al · 2019
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
S. Sun et al · 2019
Earlier work this paper cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Original
I. Turc et al · 2019
Earlier work this paper cites.
Proxquant: Quantized neural networks via proximal operators
Y. Bai et al · 2019
Earlier work this paper cites.
Efficient 8-bit quantization of transformer neural machine language translation model
Original
A. Bhandare et al · 2019
Earlier work this paper cites.
Quantized transformer
C. Fan · 2019
Earlier work this paper cites.
transformers. zip: Compressing transformers with pruning and quantization
R. Cheong and R. Daniel · 2019
Earlier work this paper cites.
Nat: Neural architecture transformer for accurate and compact architectures
Y. Guo et al · 2019
Earlier work this paper cites.
The evolved transformer
D. So et al · 2019
Earlier work this paper cites.
A large-scale study of representation learning with the visual task adaptation benchmark
Original
X. Zhai et al · 2019
Earlier work this paper cites.
Robust neural machine translation with doubly adversarial inputs
Y. Cheng et al · 2019
Earlier work this paper cites.
Is attention interpretable?
S. Serrano and N. A. Smith · 2019
Earlier work this paper cites.
Attention is not not explanation
S. Wiegreffe and Y. Pinter · 2019
Earlier work this paper cites.
Towards understanding the role of over-parametrization in generalization of neural networks
B. Neyshabur et al · 2019
Earlier work this paper cites.
Davinci: A scalable architecture for neural network computing
H. Liao et al · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Original
Y. Liu et al · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang et al · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
L. Dong et al · 2019
Earlier work this paper cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Original
M. Lewis et al · 2019
Earlier work this paper cites.
Ernie: Enhanced language representation with informative entities
Original
Z. Zhang et al · 2019
Earlier work this paper cites.
Knowledge enhanced contextual word representations
Original
M. E. Peters et al · 2019
Earlier work this paper cites.