Fetching the paper…
Reading the bibliography…
This paper introduces the Sequential Monte Carlo Transformer, an original approach that naturally captures the observations distribution in a transformer architecture.
On the expressiveness of approximate inference in bayesian neural networks
Foong, A. Y., Burt, D. R., Li, Y., and Turner, R. E. (2019) · 1909
Earlier work this paper cites.
Particle smoothing variational objectives
Moretti, A. K., Wang, Z., Wu, L., Drori, I., and Pe’er, I. (2019) · 1909
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
Dempster, A. P., Laird, N. M., and Rubin, D. B. (1977) · 1977
Earlier work this paper cites.
A practical bayesian framework for backpropagation networks
MacKay, D. J. (1992) · 1992
Earlier work this paper cites.
Novel approach to nonlinear/non-Gaussian bayesian state estimation
Gordon, N., Salmond, D., and Smith, A. (1993) · 1993
Earlier work this paper cites.
Monte-Carlo filter and smoother for non-Gaussian nonlinear state space models
Kitagawa, G. (1996) · 1996
Earlier work this paper cites.
Efficient particle-based online smoothing in general hidden Markov models: the PaRIS algorithm
Olsson, J., Westerborn, J., et al. (2017) · 1996
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Sequential Monte Carlo methods for dynamic systems
Liu, J. and Chen, R. (1998) · 1998
Earlier work this paper cites.
Filtering via simulation: Auxiliary particle filters
Pitt, M. K. and Shephard, N. (1999) · 1999
Earlier work this paper cites.
Query learning with large margin classifiers
Campbell, C., Cristianini, N., and Smola, A. J. (2000) · 2000
Earlier work this paper cites.
On sequential Monte-Carlo sampling methods for Bayesian filtering
Doucet, A., Godsill, S., and Andrieu, C. (2000) · 2000
Earlier work this paper cites.
Monte carlo smoothing and self-organizing state-space model
Kitagawa, G. and Sato, S. (2001) · 2001
Earlier work this paper cites.
Variational hyper rnn for sequence modeling
Deng, R., Cao, Y., Chang, B., Sigal, L., Mori, G., and Brubaker, M. A. (2020) · 2002
Earlier work this paper cites.
Feynman-Kac Formulae: Genealogical and Interacting Particle Systems With Applications
Del Moral, P. (2004) · 2004
Earlier work this paper cites.
Monte Carlo smoothing for non-linear time series
Godsill, S. J., Doucet, A., and West, M. (2004) · 2004
Earlier work this paper cites.
Inference in Hidden Markov Models
Cappé, O., Moulines, E., and Rydén, T. (2005) · 2005
Earlier work this paper cites.
Deep and confident prediction for time series at uber
Zhu, L. and Laptev, N. (2017) · 2007
Earlier work this paper cites.
Sequential monte carlo smoothing with application to parameter estimation in nonlinear state space models
Olsson, J., Cappe, O., Douc, R., and Moulines, E. (2008) · 2008
Earlier work this paper cites.
A backward particle interpretation of feynman-kac formulae
Del Moral, P., Doucet, A., and Singh, S. S. (2010) · 2010
Cited alongside, same era.
A sequential smoothing algorithm with linear computational cost
Fearnhead, P., Wyncoll, D., and Tawn, J. (2010) · 2010
Cited alongside, same era.
Kalman temporal differences
Geist, M. and Pietquin, O. (2010) · 2010
Cited alongside, same era.
Particle approximations of the score and observed information matrix in state space models with application to parameter estimation
Poyiadjis, G., Doucet, A., and Singh, S. (2011) · 2011
Cited alongside, same era.
Bayesian learning for neural networks
Neal, R. M. (2012) · 2012
Cited alongside, same era.
Non-asymptotic deviation inequalities for smoothed additive functionals in nonlinear state-space models
Dubarry, C. and Le Corff, S. (2013) · 2013
Bayesian recurrent neural networks
Fortunato, M., Blundell, C., and Vinyals, O. (2017) · 2017
Later among the works it cites.
Snapshot ensembles: Train 1, get M for free
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J. E., and Weinberger, K. Q. (2017) · 2017
Later among the works it cites.
Bayesian segnet: Model uncertainty in deep convolutional encoder-decoder architectures for scene understanding
Kendall, A., Badrinarayanan, V., and Cipolla, R. (2017) · 2017
Later among the works it cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C. (2017) · 2017
Later among the works it cites.
Filtering variational objectives
Maddison, C. J., Lawson, J., Tucker, G., Heess, N., Norouzi, M., Mnih, A., Doucet, A., and Teh, Y. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2014) · 2014
Cited alongside, same era.
A structured self-attentive sentence embedding
Lin, Z., Feng, M., dos Santos, C. N., Yu, M., Xiang, B., Zhou, B., and Bengio, Y. (2014) · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2015) · 2015
Cited alongside, same era.
Weight uncertainty in neural networks
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. (2015) · 2015
Cited alongside, same era.
A recurrent latent variable model for sequential data
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A. C., and Bengio, Y. (2015) · 2015
Cited alongside, same era.
Variational recurrent auto-encoders
Fabius, O. and van Amersfoort, J. R. (2015) · 2015
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
Noisy networks for exploration
Fortunato, M., Azar, M. G., Piot, B., Menick, J., Osband, I., Graves, A., Mnih, V., Munos, R., Hassabis, D., Pietquin, O., Blundell, C., and Legg, S. (2018) · 2018
Later among the works it cites.
Deep variational reinforcement learning for pomdps
Igl, M., Zintgraf, L., Le, T. A., Wood, F., and Whiteson, S. (2018) · 2018
Later among the works it cites.
Fast and scalable bayesian deep learning by weight-perturbation in adam
Khan, M. E., Nielsen, D., Tangkaratt, V., Lin, W., Gal, Y., and Srivastava, A. (2018) · 2018
Later among the works it cites.
Twisted variational sequential monte carlo
Lawson, D., Tucker, G., Naesseth, C. A., Maddison, C. J., Adams, R. P., and Teh, Y. W. (2018) · 2018
Later among the works it cites.
Auto-encoding sequential monte carlo
Le, T. A., Igl, M., Rainforth, T., Jin, T., and Wood, F. (2018) · 2018
Later among the works it cites.
Variational sequential monte carlo
Naesseth, C., Linderman, S., Ranganath, R., and Blei, D. (2018) · 2018
Later among the works it cites.
High-quality prediction intervals for deep learning: A distribution-free, ensembled approach
Pearce, T., Brintrup, A., Zaki, M., and Neely, A. (2018) · 2018
Later among the works it cites.
Bayesian uncertainty estimation for batch normalized deep networks
Teye, M., Azizpour, H., and Smith, K. (2018) · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2019) · 2019
Later among the works it cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X. (2019) · 2019
Later among the works it cites.
Single-model uncertainties for deep learning
Tagasovska, N. and Lopez-Paz, D. (2019) · 2019
Later among the works it cites.
Pitfalls of in-domain uncertainty estimation and ensembling in deep learning
Ashukha, A., Lyzhov, A., Molchanov, D., and Vetrov, D. (2020) · 2020
Closest in time.
Deeppipe: A distribution-free uncertainty quantification approach for time series forecasting
Wang, B., Li, T., Yan, Z., Zhang, G., and Lu, J. (2020) · 2020
Closest in time.