Fetching the paper…
Reading the bibliography…
Mixture models trained via EM are among the simplest, most widely used and well understood latent variable models in the machine learning literature.
Maximum likelihood from incomplete data via the em algorithm
Dempster, A. P., Laird, N. M., and Rubin, D. B · 1977
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S., , and Hinton, G. E · 1991
Earlier work this paper cites.
An information-theoretic analysis of hard and soft assignment methods for clustering
Kearns, M., Mansour, Y., and Ng, A. Y · 1998
Earlier work this paper cites.
A view of the em algorithm that justifies incremental, sparse, and other variants
Neal, R. M. and Hinton, G. E · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Koehn, P., Hoang, H., Birch, A., Callison-Burch, C., Federico, M., Bertoldi, N., Cowan, B., Shen, W., Moran, C., Zens, R., Dyer, C., Bojar, O., Constantin, A., and Herbst, E · 2007
Earlier work this paper cites.
Hyter: Meaning-equivalent semantics for translation evaluation
Dreyer, M. and Marcu, D · 2012
Earlier work this paper cites.
Multiple choice learning: Learning to produce multiple structured outputs
Guzman-Rivera, A., Batra, D., and Kohli, P · 2012
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Graves, A · 2013
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
Eigen, D., Ranzato, M., and Sutskever, I · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
deltableu: A discriminative metric for generation tasks with intrinsically diverse targets
Galley, M., Brockett, C., Sordoni, A., Ji, Y., Auli, M., Quirk, C., Mitchell, M., Gao, J., and Dolan, B · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Generating sentences from a continuous space
Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A. M., Jozefowicz, R., and Bengio, S · 2016
Cited alongside, same era.
Stochastic multiple choice learning for training diverse deep ensembles
Lee, S., Purushwalkam, S., Cogswell, M., Ranjan, V., Crandall, D., and Batra, D · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Cited alongside, same era.
Variational neural machine translation
Zhang, B., Xiong, D., Su, J., Duan, H., and Zhang, M · 2016
Cited alongside, same era.
Latent variable dialogue models and their diversity
Cao, K. and Clark, S · 2017
Cited alongside, same era.
Diverse and accurate image description using a variational auto-encoder with an additive gaussian encoding space
Wang, L., Schwing, A., and Lazebnik, S · 2017
Later among the works it cites.
Latent intention dialogue models
Wen, T.-H., Miao, Y., Blunsom, P., and Young, S · 2017
Later among the works it cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., and Yu, Y · 2017
Later among the works it cites.
Understanding back-translation at scale
Edunov, S., Ott, M., Auli, M., and Grangier, D · 2018
Later among the works it cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Later among the works it cites.
Achieving human parity on automatic chinese to english news translation
Hassan, H., Aue, A., Chen, C., Chowdhary, V., Clark, J., Federmann, C., Huang, X., Junczys-Dowmunt, M., Lewis, W., Li, M., et al · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards diverse and natural image descriptions via a conditional gan
Dai, B., Fidler, S., Urtasun, R., and Lin, D · 2017
Cited alongside, same era.
Convolutional Sequence to Sequence Learning
Gehring, J., Auli, M., Grangier, D., Yarats, D., and Dauphin, Y. N · 2017
Cited alongside, same era.
A simple, fast diverse decoding algorithm for neural generation
Li, J., Monroe, W., and Jurafsky, D · 2017
Cited alongside, same era.
A hierarchical latent variable encoder-decoder model for generating dialogues
Serban, I. V., Sordoni, A., Lowe, R., Charlin, L., Pineau, J., Courville, A. C., and Bengio, Y · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Later among the works it cites.
Sequence to sequence mixture model for diverse machine translation
He, X., Haffari, G., and Norouzi, M · 2018
Later among the works it cites.
Fast decoding in sequence models using discrete latent variables
Kaiser, L., Bengio, S., Roy, A., Vaswani, A., Parmar, N., Uszkoreit, J., and Shazeer, N · 2018
Later among the works it cites.
A stochastic decoder for neural machine translation
Schulz, P., Aziz, W., and Cohn, T · 2018
Later among the works it cites.
Diverse beam search for improved description of complex scenes
Vijayakumar, A. K., Cogswell, M., Selvaraju, R. R., Sun, Q., Lee, S., Crandall, D., and Batra, D · 2018
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Closest in time.