Fetching the paper…
Reading the bibliography…
We introduce semi-parametric inducing point networks (SPIN), a general-purpose architecture that can query the training set at inference time in a compute-efficient manner.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy · 1907
Earlier work this paper cites.
Connectionist learning procedures, 1989
Geoffrey E. Hinton · 1989
Earlier work this paper cites.
An introduction to kernel and nearest-neighbor nonparametric regression
N. S. Altman · 1992
Earlier work this paper cites.
Support vector machines
Marti A. Hearst, Susan T Dumais, Edgar Osuna, John Platt, and Bernhard Scholkopf · 1998
Earlier work this paper cites.
Sampling techniques for kernel methods
Dimitris Achlioptas, Frank McSherry, and Bernhard Schölkopf · 2001
Earlier work this paper cites.
Greedy function approximation: A gradient boosting machine
Jerome H. Friedman · 2001
Earlier work this paper cites.
Modeling linkage disequilibrium and identifying recombination hotspots using single-nucleotide polymorphism data
Na Li and Matthew Stephens · 2003
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2003
Earlier work this paper cites.
Efficient non-parametric function induction in semi-supervised learning
Olivier Delalleau, Yoshua Bengio, and Nicolas Le Roux · 2005
Earlier work this paper cites.
Sparse gaussian processes using pseudo-inputs
Edward Snelson and Zoubin Ghahramani · 2005
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2007
Earlier work this paper cites.
The Encylopedia of Molecular Biology
John Kendrew · 2009
Earlier work this paper cites.
Genotype imputation
Yun Li, Cristen Willer, Serena Sanna, and Gonçalo Abecasis · 2009
Earlier work this paper cites.
Variational learning of inducing variables in sparse gaussian processes
Michalis Titsias · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Remarks on some nonparametric estimates of a density function
Richard A Davis, Keh-Shin Lii, and Dimitris N Politis · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Sharp analysis of low-rank kernel matrix approximations
Francis Bach · 2013
Earlier work this paper cites.
Efficient optimization for sparse gaussian process regression
Yanshuai Cao, Marcus A Brubaker, David J Fleet, and Aaron Hertzmann · 2013
Earlier work this paper cites.
Deep gaussian processes
Andreas Damianou and Neil D Lawrence · 2013
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Mcmc for variationally sparse gaussian processes
James Hensman, Alexander G Matthews, Maurizio Filippone, and Zoubin Ghahramani · 2015
Earlier work this paper cites.
Kernel interpolation for scalable structured gaussian processes (kiss-gp)
Andrew Wilson and Hannes Nickisch · 2015
Earlier work this paper cites.
Thoughts on massively scalable gaussian processes
Andrew Gordon Wilson, Christoph Dann, and Hannes Nickisch · 2015
Earlier work this paper cites.
The international Genome sample resource (IGSR): A worldwide collection of genome variation incorporating the 1000 Genomes Project data
Laura Clarke, Susan Fairley, Xiangqun Zheng-Bradley, Ian Streeter, Emily Perry, Ernesto Lowy, Anne-Marie Tassé, and Paul Flicek · 2016
Cited alongside, same era.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2016
Cited alongside, same era.
Faster variational inducing input gaussian process classification
Pavel Izmailov and Dmitry Kropotov · 2016
Cited alongside, same era.
Meta-learning with memory-augmented neural networks
Adam Santoro, Sergey Bartunov, Matthew Botvinick, Daan Wierstra, and Timothy Lillicrap · 2016
Cited alongside, same era.
Sequential inference for deep gaussian process
Yali Wang, Marcus Brubaker, Brahim Chaib-Draa, and Raquel Urtasun · 2016
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Later among the works it cites.
Bootstrapping neural processes
Juho Lee, Yoonho Lee, Jungtaek Kim, Eunho Yang, Sung Ju Hwang, and Yee Whye Teh · 2020
Later among the works it cites.
Sparkbeagle: Scalable genotype imputation from distributed whole-genome reference panels in the cloud
Altti Ilari Maarala, Kalle Pärn, Javier Nuñez-Fontarnau, and Keijo Heljanko · 2020
Later among the works it cites.
Dataset meta-learning from kernel ridge-regression
Timothy Nguyen, Zhourong Chen, and Jaehoon Lee · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep kernel learning
Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P Xing · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
A one-penny imputed genome from next-generation reference panels
Brian L. Browning, Ying Zhou, and Sharon R. Browning · 2018
Cited alongside, same era.
Scalable gaussian processes with grid-structured eigenfunctions (gp-grief), 2018
Trefor W. Evans and Prasanth B. Nair · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Cited alongside, same era.
Marta Garnelo, Jonathan Schwarz, Dan Rosenbaum, Fabio Viola, Danilo J Rezende, SM Eslami, and Yee Whye Teh · 2018
Cited alongside, same era.
Attentive neural processes
Hyunjik Kim, Andriy Mnih, Jonathan Schwarz, Marta Garnelo, Ali Eslami, Dan Rosenbaum, Oriol Vinyals, and Yee Whye Teh · 2018
Cited alongside, same era.
Genotype imputation using the positional burrows wheeler transform
Simone Rubinacci, Olivier Delaneau, and Jonathan Marchini · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Later among the works it cites.
Retrieval-augmented transformer-xl for close-domain dialog generation
Giovanni Bonetta, Rossella Cancelliere, Ding Liu, and Paul Vozila · 2021
Later among the works it cites.
Coordination among neural modules through a shared global workspace
Anirudh Goyal, Aniket Didolkar, Alex Lamb, Kartikeya Badola, Nan Rosemary Ke, Nasim Rahaman, Jonathan Binas, Charles Blundell, Michael Mozer, and Yoshua Bengio · 2021
Later among the works it cites.
Long-range transformers for dynamic spatiotemporal forecasting
Jake Grigsby, Zhe Wang, and Yanjun Qi · 2021
Later among the works it cites.
Self-attention between datapoints: Going beyond individual input-output pairs in deep learning
Jannik Kossen, Neil Band, Clare Lyle, Aidan N. Gomez, Tom Rainforth, and Yarin Gal · 2021
Later among the works it cites.
A beginner’s guide to low-coverage whole genome sequencing for population genomics
Runyang Nicolas Lou, Arne Jacobs, Aryn P Wilder, and Nina Overgaard Therkildsen · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Later among the works it cites.
Mlp-mixer: An all-mlp architecture for vision
Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Nyströmformer: A nystöm-based algorithm for approximating self-attention
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh · 2021
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Closest in time.
Retrieval-augmented reinforcement learning, 2022
Anirudh Goyal, Abram L. Friesen, Andrea Banino, Theophane Weber, Nan Rosemary Ke, Adria Puigdomenech Badia, Arthur Guez, Mehdi Mirza, Peter C. Humphreys, Ksenia Konyushkova, Laurent Sifre, Michal Valko, Simon Osindero, Timothy Lillicrap, Nicolas Heess, and Charles Blundell · 2022
Closest in time.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Closest in time.
Transformer neural processes: Uncertainty-aware meta learning via sequence modeling
Tung Nguyen and Aditya Grover · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen · 2022
Closest in time.
How to apply de bruijn graphs to genome assembly
Phillip E. C. Compeau, Pavel A. Pevzner, and Glenn Tesler · 2023
Closest in time.
Latent bottlenecked attentive neural processes
Leo Feng, Hossein Hajimirsadeghi, Yoshua Bengio, and Mohamed Osama Ahmed · 2023
Closest in time.