Fetching the paper…
Reading the bibliography…
Neural Processes (NPs) are a popular class of approaches for meta-learning.
The application of bayesian methods for seeking the extremum
Mockus, J., Tiesis, V., and Zilinskas, A · 1978
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Schmidhuber, J · 1987
Earlier work this paper cites.
Global versus local search in constrained optimization of computer models
Schonlau, M., Welch, W. J., and Jones, D. R · 1998
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N. and Lugosi, G · 2006
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Taking the human out of the loop: A review of bayesian optimization
Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and De Freitas, N · 2015
Earlier work this paper cites.
Chen, X., Kingma, D. P., Salimans, T., Duan, Y., Dhariwal, P., Schulman, J., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Emnist: Extending mnist to handwritten letters
Cohen, G., Afshar, S., Tapson, J., and Van Schaik, A · 2017
Earlier work this paper cites.
BayesO: A Bayesian optimization framework in Python
Kim, J. and Choi, S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Fixing a broken elbo
Alemi, A., Poole, B., Fischer, I., Dillon, J., Saurous, R. A., and Murphy, K · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
A tutorial on bayesian optimization
Frazier, P. I · 2018
Earlier work this paper cites.
Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration
Gardner, J. R., Pleiss, G., Bindel, D., Weinberger, K. Q., and Wilson, A. G · 2018
Earlier work this paper cites.
Empirical evaluation of neural process objectives
Le, T. A., Kim, H., Garnelo, M., Rosenbaum, D., Schwarz, J., and Teh, Y. W · 2018
Earlier work this paper cites.
Large-scale celebfaces attributes (celeba) dataset
Liu, Z., Luo, P., Wang, X., and Tang, X · 2018
Earlier work this paper cites.
Janossy pooling: Learning deep permutation-invariant functions for variable-size inputs
Murphy, R. L., Srinivasan, B., Rao, V., and Ribeiro, B · 2018
Cited alongside, same era.
Improving language understanding with unsupervised learning
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Riquelme, C., Tucker, G., and Snoek, J · 2018
Cited alongside, same era.
Vanschoren, J · 2018
Cited alongside, same era.
Meta-learning surrogate models for sequential decision making
Galashov, A., Schwarz, J., Kim, H., Garnelo, M., Saxton, D., Kohli, P., Eslami, S., and Teh, Y. W · 2019
Benchmark functions for bayesian optimization
Kim, J · 2020
Later among the works it cites.
Bootstrapping neural processes
Lee, J., Lee, Y., Kim, J., Yang, E., Hwang, S. J., and Teh, Y. W · 2020
Later among the works it cites.
Variational transformers for diverse response generation
Lin, Z., Winata, G. I., Xu, P., Liu, Z., and Fung, P · 2020
Later among the works it cites.
Metafun: Meta-learning with iterative functional updates
Xu, J., Ton, J.-F., Kim, H., Kosiorek, A., and Teh, Y. W · 2020
Later among the works it cites.
Bruinsma, W. P., Requeima, J., Foong, A. Y., Gordon, J., and Turner, R. E · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Convolutional conditional neural processes
Gordon, J., Bruinsma, W. P., Foong, A. Y., Requeima, J., Dubois, Y., and Turner, R. E · 2019
Cited alongside, same era.
Kim, H., Mnih, A., Schwarz, J., Garnelo, M., Eslami, A., Rosenbaum, D., Vinyals, O., and Teh, Y. W · 2019
Cited alongside, same era.
Set transformer: A framework for attention-based permutation-invariant neural networks
Lee, J., Lee, Y., Kim, J., Kosiorek, A., Choi, S., and Teh, Y. W · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Singh, G., Yoon, J., Son, Y., and Ahn, S · 2019
Cited alongside, same era.
Botorch: A framework for efficient monte-carlo bayesian optimization
Balandat, M., Karrer, B., Jiang, D., Daulton, S., Letham, B., Wilson, A. G., and Bakshy, E · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
Chung, Y., Char, I., Guo, H., Schneider, J., and Neiswanger, W · 2021
Later among the works it cites.
Jumbo: Scalable multi-task bayesian optimization using offline data
Hakhamaneshi, K., Abbeel, P., Stojanovic, V., and Grover, A · 2021
Later among the works it cites.
Equivariant learning of stochastic fields: Gaussian processes and steerable conditional neural processes
Holderrieth, P., Hutchinson, M. J., and Teh, Y. W · 2021
Later among the works it cites.
Self-attention between datapoints: Going beyond individual input-output pairs in deep learning
Kossen, J., Band, N., Lyle, C., Gomez, A. N., Rainforth, T., and Gal, Y · 2021
Later among the works it cites.
Pretrained transformers as universal computation engines
Lu, K., Grover, A., Abbeel, P., and Mordatch, I · 2021
Later among the works it cites.
On contrastive representations of stochastic processes
Mathieu, E., Foster, A., and Teh, Y. W · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Data distributional properties drive emergent in-context learning in transformers
Chan, S. C., Santoro, A., Lampinen, A. K., Wang, J. X., Singh, A., Richemond, P. H., McClelland, J., DeepMind, S., and Hill, F · 2022
Closest in time.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A · 2022
Closest in time.