Fetching the paper…
Reading the bibliography…
The field of machine translation (MT), the automatic translation of written text from one natural language into another, has experienced a major paradigm shift in recent years.
Modeling latent sentence structure in neural machine translation
Bastings, J., Aziz, W., Titov, I., & Sima’an, K. (2019) · 1901
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., & Salakhutdinov, R. (2019) · 1901
Earlier work this paper cites.
Self-attentive model for headline generation
Daniil, G., Kalaidin, P., & Malykh, V. (2019) · 1901
Earlier work this paper cites.
Assessing BERT’s syntactic abilities
Goldberg, Y. (2019) · 1901
Earlier work this paper cites.
Learning efficient lexically-constrained neural machine translation with external memory
Li, Y., Liu, X., Liu, D., Zhang, X., & Liu, J. (2019) · 1901
Earlier work this paper cites.
Context in neural machine translation: A review of models and evaluations
Popescu-Belis, A. (2019) · 1901
Earlier work this paper cites.
Unsupervised neural machine translation with SMT as posterior regularization
Ren, S., Zhang, Z., Liu, S., Zhou, M., & Ma, S. (2019) · 1901
Earlier work this paper cites.
So, D. R., Liang, C., & Le, Q. V. (2019) · 1901
Earlier work this paper cites.
No training required: Exploring random encoders for sentence classification
Wieting, J., & Kiela, D. (2019) · 1901
Earlier work this paper cites.
Adding interpretable attention to neural translation models improves word alignment
Zenkel, T., Wuebker, J., & DeNero, J. (2019) · 1901
Earlier work this paper cites.
Generating textual adversarial examples for deep learning models: A survey
Zhang, W. E., Sheng, Q. Z., & Alhazmi, A. A. F. (2019) · 1901
Earlier work this paper cites.
A fully differentiable beam search decoder
Collobert, R., Hannun, A., & Synnaeve, G. (2019) · 1902
Earlier work this paper cites.
Insertion-based decoding with automatically inferred generation order
Gu, J., Liu, Q., & Cho, K. (2019a) · 1902
Earlier work this paper cites.
Guo, Q., Qiu, X., Liu, P., Shao, Y., Xue, X., & Zhang, Z. (2019) · 1902
Earlier work this paper cites.
Training on synthetic noise improves robustness to natural noise in machine translation
Karpukhin, V., Levy, O., Eisenstein, J., & Ghazvininejad, M. (2019) · 1902
Earlier work this paper cites.
Augmenting neural machine translation with knowledge graphs
Moussallem, D., Arčan, M., Ngomo, A.-C. N., & Buitelaar, P. (2019) · 1902
Earlier work this paper cites.
Insertion Transformer: Flexible sequence generation via insertion operations
Stern, M., Chan, W., Kiros, J. R., & Uszkoreit, J. (2019) · 1902
Earlier work this paper cites.
Non-autoregressive machine translation with auxiliary regularization
Wang, Y., Tian, F., He, D., Qin, T., Zhai, C., & Liu, T.-Y. (2019b) · 1902
Earlier work this paper cites.
Non-monotonic sequential text generation
Welleck, S., Brantley, K., Daumé III, H., & Cho, K. (2019) · 1902
Earlier work this paper cites.
Calibration of encoder decoder models for neural machine translation
Kumar, A., & Sarawagi, S. (2019) · 1903
Earlier work this paper cites.
Train, sort, explain: Learning to diagnose translation models
Schwarzenberg, R., Harbecke, D., Macketanz, V., Avramidis, E., & Möller, S. (2019) · 1903
Earlier work this paper cites.
Neutron: An implementation of the Transformer translation model and its variants
Xu, H., & Liu, Q. (2019) · 1903
Earlier work this paper cites.
Modeling recurrence for Transformer
Hao, J., Wang, X., Yang, B., Wang, L., Zhang, J., & Tu, Z. (2019) · 1904
Earlier work this paper cites.
Unsupervised recurrent neural network grammars
Kim, Y., Rush, A. M., Yu, L., Kuncoro, A., Dyer, C., & Melis, G. (2019b) · 1904
Earlier work this paper cites.
Dynamic evaluation of Transformer language models
Krause, B., Kahembwe, E., Murray, I., & Renals, S. (2019) · 1904
Earlier work this paper cites.
End-to-end speech translation with knowledge distillation
Liu, Y., Xiong, H., He, Z., Zhang, J., Wu, H., Wang, H., & Zong, C. (2019) · 1904
Earlier work this paper cites.
Language models with Transformers
Wang, C., Li, M., & Smola, A. (2019a) · 1904
Earlier work this paper cites.
A survey of multilingual neural machine translation
Dabre, R., Chu, C., & Kunchukuttan, A. (2019) · 1905
Earlier work this paper cites.
Gu, J., Wang, C., & Zhao, J. (2019b) · 1905
Earlier work this paper cites.
On the validity of self-attention as explanation in transformer models
Brunner, G., Liu, Y., Pascual, D., Richter, O., & Wattenhofer, R. (2019) · 1908
Earlier work this paper cites.
Chunk-based decoder for neural machine translation
Ishiwatari, S., Yao, J., Liu, S., Li, M., Zhou, M., Yoshinaga, N., Kitsuregawa, M., & Jia, W. (2017) · 1912
Earlier work this paper cites.
Improving robustness of machine translation with synthetic noise
Vaibhav, V., Singh, S., Stewart, C., & Neubig, G. (2019) · 1920
Earlier work this paper cites.
Improved neural machine translation with a syntax-aware encoder and decoder
Chen, H., Huang, S., Chiang, D., & Chen, J. (2017b) · 1945
Earlier work this paper cites.
The psychology of language
Zipf, G. K. (1946) · 1946
Earlier work this paper cites.
Unfolding and shrinking neural machine translation ensembles
Stahlberg, F., & Byrne, B. (2017) · 1956
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G. E., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Graph convolutional encoders for syntax-aware neural machine translation
Bastings, J., Titov, I., Aziz, W., Marcheggiani, D., & Simaan, K. (2017) · 1967
Earlier work this paper cites.
Semi-supervised learning for neural machine translation
Cheng, Y., Xu, W., He, Z., He, W., Wu, H., Sun, M., & Liu, Y. (2016c) · 1974
Earlier work this paper cites.
Trainable greedy decoding for neural machine translation
Gu, J., Cho, K., & Li, V. O. (2017b) · 1978
Earlier work this paper cites.
Neurocomputing: Foundations of research
Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1988) · 1988
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W. E., & Jackel, L. D. (1989a) · 1989
Earlier work this paper cites.
Phoneme recognition using time-delay neural networks
Waibel, A., Hanazawa, T., Hinton, G. E., Shikano, K., & Lang, K. J. (1989) · 1989
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Williams, R. J., & Zipser, D. (1989) · 1989
Earlier work this paper cites.
Neural network ensembles
Hansen, L. K., & Salamon, P. (1990) · 1990
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W. E., & Jackel, L. D. (1990) · 1990
Earlier work this paper cites.
Recursive distributed representations
Pollack, J. B. (1990) · 1990
Earlier work this paper cites.
Connectionist pushdown automata that learn context-free grammars
Sun, G.-Z., Chen, H.-H., Giles, C. L., Lee, Y.-C., & Chen, D. (1990) · 1990
Earlier work this paper cites.
Generalization and maximum likelihood from small data sets
Byrne, B. (1993) · 1993
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B., Stork, D. G. et al. (1993) · 1993
Earlier work this paper cites.
The Neural Network Pushdown Automation: Model, Stack and Learning Simulations
Sun, G.-Z., Giles, C. L., Chen, H.-H., & Lee, Y.-C. (1993) · 1993
Earlier work this paper cites.
A new algorithm for data compression
Gage, P. (1994) · 1994
Earlier work this paper cites.
On the computational power of neural nets
Siegelmann, H. T., & Sontag, E. D. (1995) · 1995
Earlier work this paper cites.
HMM-based word alignment in statistical translation
Vogel, S., Ney, H., & Tillmann, C. (1996) · 1996
Earlier work this paper cites.
A latent semantic analysis framework for large-span language modeling
Bellegarda, J. R. (1997) · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., & Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Schuster, M., & Paliwal, K. K. (1997) · 1997
Earlier work this paper cites.
Stochastic inversion transduction grammars and bilingual parsing of parallel corpora
Wu, D. (1997) · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., & Haffner, P. (1998) · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M. (1999) · 1999
Earlier work this paper cites.
Mining the web for bilingual text
Resnik, P. (1999) · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
Dietterich, T. G. (2000) · 2000
Earlier work this paper cites.
Segmental minimum Bayes-risk ASR voting strategies
Goel, V., Kumar, S., & Byrne, B. (2000) · 2000
Earlier work this paper cites.
Gradient flow in recurrent nets: The difficulty of learning long-term dependencies
Hochreiter, S., Bengio, Y., Frasconi, P., & Schmidhuber, J. (2001) · 2001
Earlier work this paper cites.
Statistical multi-source translation
Och, F. J., & Ney, H. (2001) · 2001
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002) · 2002
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., & Jauvin, C. (2003) · 2003
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Brown, P. F., Della Pietra, S. A., Della Pietra, V. J., & Mercer, R. L. (1993) · 2003
Earlier work this paper cites.
Minimum error rate training in statistical machine translation
Och, F. J. (2003) · 2003
Earlier work this paper cites.
The web as a parallel corpus
Resnik, P., & Smith, N. A. (2003) · 2003
Earlier work this paper cites.
Neural lattice search for domain adaptation in machine translation
Khayrallah, H., Kumar, G., Duh, K., Post, M., & Koehn, P. (2017) · 2004
Earlier work this paper cites.
Minimum Bayes-risk decoding for statistical machine translation
Kumar, S., & Byrne, B. (2004) · 2004
Earlier work this paper cites.
Adaptation of the translation model for statistical machine translation based on information retrieval
Hildebrand, A. S., Eck, M., Vogel, S., & Waibel, A. (2005) · 2005
Earlier work this paper cites.
SGNMT – a flexible NMT decoding platform for quick prototyping of new models and search strategies
Stahlberg, F., Hasler, E., Saunders, D., & Byrne, B. (2017b) · 2005
Earlier work this paper cites.
Word-level confidence estimation for machine translation using phrase-based translation models
Ueffing, N., & Ney, H. (2005) · 2005
Earlier work this paper cites.
Neural probabilistic language models
Bengio, Y., Schwenk, H., Senécal, J.-S., Morin, F., & Gauvain, J.-L. (2006) · 2006
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., & Niculescu-Mizil, A. (2006) · 2006
Earlier work this paper cites.
Hierarchical phrase-based translation
Chiang, D. (2007) · 2007
Earlier work this paper cites.
Simultaneous translation of lectures and speeches
Fügen, C., Waibel, A., & Kolss, M. (2007) · 2007
Earlier work this paper cites.
Factored translation models
Koehn, P., & Hoang, H. (2007) · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R., & Weston, J. (2008) · 2008
Earlier work this paper cites.
Investigations on large-scale lightly-supervised training for statistical machine translation
Schwenk, H. (2008) · 2008
Earlier work this paper cites.
Lattice Minimum Bayes-Risk decoding for statistical machine translation
Tromble, R., Kumar, S., Och, F. J., & Macherey, W. (2008) · 2008
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., & Weston, J. (2009) · 2009
Earlier work this paper cites.
Automatic translation from parallel speech: Simultaneous interpretation as mt training data
Paulik, M., & Waibel, A. (2009) · 2009
Earlier work this paper cites.
Discriminative instance weighting for domain adaptation in statistical machine translation
Foster, G., Goutte, C., & Kuhn, R. (2010) · 2010
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Gutmann, M., & Hyvärinen, A. (2010) · 2010
Earlier work this paper cites.
Statistical Machine Translation
Koehn, P. (2010) · 2010
Earlier work this paper cites.
Learning to combine foveal glimpses with a third-order Boltzmann machine
Larochelle, H., & Hinton, G. E. (2010) · 2010
Earlier work this paper cites.
Data-intensive text processing with MapReduce
Lin, J., & Dyer, C. (2010) · 2010
Earlier work this paper cites.
Ensemble-based classifiers
Rokach, L. (2010) · 2010
Earlier work this paper cites.
N-gram-based machine translation enhanced with neural networks for the French-English BTEC-IWSLT’10 task
Zamora-Martinez, F., Castro-Bleda, M. J., & Schwenk, H. (2010) · 2010
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Axelrod, A., He, X., & Gao, J. (2011) · 2011
Earlier work this paper cites.
Goodness: A method for measuring machine translation confidence
Bach, N., Huang, F., & Al-Onaizan, Y. (2011) · 2011
Earlier work this paper cites.
Natural language processing (almost) from scratch
Collobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., & Kuksa, P. (2011) · 2011
Earlier work this paper cites.
Training speech translation from audio recordings of interpreter-mediated communication
Paulik, M., & Waibel, A. (2013) · 2011
Earlier work this paper cites.
MT detection in web-scraped parallel corpora
Rarrick, S., Quirk, C., & Lewis, W. D. (2011) · 2011
Earlier work this paper cites.
Semi-supervised recursive Autoencoders for predicting sentiment distributions
Socher, R., Pennington, J., Huang, E. H., Ng, A. Y., & Manning, C. D. (2011) · 2011
Earlier work this paper cites.
Exploiting objective annotations for measuring translation post-editing effort
Specia, L. (2011) · 2011
Earlier work this paper cites.
Parallel corpus refinement as an outlier detection algorithm
Taghipour, K., Khadivi, S., & Xu, J. (2011) · 2011
Earlier work this paper cites.
Theano: New features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I., Bergeron, A., Bouchard, N., Warde-Farley, D., & Bengio, Y. (2012) · 2012
Earlier work this paper cites.
Learning to parse and translate improves neural machine translation
Eriguchi, A., Tsuruoka, Y., & Cho, K. (2017) · 2012
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T., & Richardson, J. (2018) · 2012
Earlier work this paper cites.
Continuous space translation models with neural networks
Le, H.-S., Allauzen, A., & Yvon, F. (2012) · 2012
Earlier work this paper cites.
Conversion of recurrent neural network language models to weighted finite state transducers for automatic speech recognition
Lecorvé, G., & Motlicek, P. (2012) · 2012
Earlier work this paper cites.
A fast and simple algorithm for training neural probabilistic language models
Mnih, A., & Teh, Y. W. (2012) · 2012
Earlier work this paper cites.
Japanese and korean voice search
Schuster, M., & Nakajima, K. (2012) · 2012
Earlier work this paper cites.
Continuous space translation models for phrase-based statistical machine translation
Schwenk, H. (2012) · 2012
Earlier work this paper cites.
ADADELTA: An adaptive learning rate method
Zeiler, M. D. (2012) · 2012
Earlier work this paper cites.
Machine translation detection from monolingual web-text
Arase, Y., & Zhou, M. (2013) · 2013
Earlier work this paper cites.
Pruning algorithms of neural networks — a comparative study
Augasta, M. G., & Kathirvalavakumar, T. (2013) · 2013
Earlier work this paper cites.
Basho: the complete haiku
Basho, & Reichhold, J. (2013) · 2013
Earlier work this paper cites.
Better mixing via deep representations
Bengio, Y., Mesnil, G., Dauphin, Y. N., & Rifai, S. (2013) · 2013
Earlier work this paper cites.
Audio chord recognition with recurrent neural networks
Boulanger-Lewandowski, N., Bengio, Y., & Vincent, P. (2013) · 2013
Earlier work this paper cites.
Predicting parameters in deep learning
Denil, M., Shakibi, B., Dinh, L., Ranzato, M., & de Freitas, N. (2013) · 2013
Earlier work this paper cites.
A systematic exploration of diversity in machine translation
Gimpel, K., Batra, D., Dyer, C., & Shakhnarovich, G. (2013) · 2013
Earlier work this paper cites.
N-gram posterior probability confidence measures for statistical machine translation: An empirical study
de Gispert, A., Blackwood, G., Iglesias, G., & Byrne, B. (2013) · 2013
Earlier work this paper cites.
Scalable modified Kneser-Ney language model estimation
Heafield, K., Pouzyrevsky, I., Clark, J. H., & Koehn, P. (2013) · 2013
Earlier work this paper cites.
Recurrent continuous translation models
Kalchbrenner, N., & Blunsom, P. (2013) · 2013
Earlier work this paper cites.
Recursive autoencoders for ITG-based translation
Li, P., Liu, Y., & Sun, M. (2013) · 2013
Earlier work this paper cites.
Learning word embeddings efficiently with noise-contrastive estimation
Mnih, A., & Kavukcuoglu, K. (2013) · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., & Bengio, Y. (2013) · 2013
Earlier work this paper cites.
Sentence simplification with memory-augmented neural networks
Vu, T., Hu, B., Munkhdalai, T., & Yu, H. (2018) · 2013
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
Xue, J., Li, J., & Gong, Y. (2013) · 2013
Earlier work this paper cites.
Multiple object recognition with visual attention
Ba, J. L., Mnih, V., & Kavukcuoglu, K. (2014) · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., & Bengio, Y. (2014b) · 2014
Earlier work this paper cites.
End-to-end continuous speech recognition using attention-based recurrent NN: First results
Chorowski, J., Bahdanau, D., Cho, K., & Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
Denton, E. L., Zaremba, W., Bruna, J., LeCun, Y., & Fergus, R. (2014) · 2014
Earlier work this paper cites.
Fast and robust neural network joint models for statistical machine translation
Devlin, J., Zbib, R., Huang, Z., Lamar, T., Schwartz, R., & Makhoul, J. (2014) · 2014
Earlier work this paper cites.
Notes on noise contrastive estimation and negative sampling
Dyer, C. (2014) · 2014
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., & Danihelka, I. (2014) · 2014
Earlier work this paper cites.
Don’t until the final verb wait: Reinforcement learning for simultaneous machine translation
Grissom II, A., He, H., Boyd-Graber, J., Morgan, J., & Daumé III, H. (2014) · 2014
Earlier work this paper cites.
A convolutional neural network for modelling sentences
Kalchbrenner, N., Grefenstette, E., & Blunsom, P. (2014) · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Kim, Y. (2014) · 2014
Earlier work this paper cites.
Efficient lattice rescoring using recurrent neural network language models
Liu, X., Wang, Y., Chen, X., Gales, M. J., & Woodland, P. C. (2014) · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
Livni, R., Shalev-Shwartz, S., & Shamir, O. (2014) · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Mnih, V., Heess, N., Graves, A. et al. (2014) · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J., Socher, R., & Manning, C. D. (2014) · 2014
Earlier work this paper cites.
Overcoming the curse of sentence length for neural machine translation using automatic segmentation
Pouget-Abadie, J., Bahdanau, D., van Merrienboer, B., Cho, K., & Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Deep convolutional neural networks for sentiment analysis of short texts
dos Santos, C., & Gatti, M. (2014) · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., & Le, Q. V. (2014) · 2014
Earlier work this paper cites.
Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models
Swietojanski, P., & Renals, S. (2014) · 2014
Earlier work this paper cites.
Morphological inflection generation with hard monotonic attention
Aharoni, R., & Goldberg, Y. (2017a) · 2015
Earlier work this paper cites.
When and why are log-linear models self-normalizing?
Andreas, J., & Klein, D. (2015) · 2015
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., & Samek, W. (2015) · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., & Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., & Shazeer, N. (2015) · 2015
Earlier work this paper cites.
Multi-task learning for multiple language translation
Dong, D., Wu, H., He, W., Yu, D., & Wang, H. (2015) · 2015
Earlier work this paper cites.
Multilingual image description with neural sequence models
Elliott, D., Frank, S., & Hasler, E. (2015) · 2015
Earlier work this paper cites.
Learning to transduce with unbounded memory
Grefenstette, E., Hermann, K. M., Suleyman, M., & Blunsom, P. (2015) · 2015
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Gulcehre, C., Firat, O., Xu, K., Cho, K., Barrault, L., Lin, H.-C., Bougares, F., Schwenk, H., & Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., & Dally, W. (2015) · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., & Blunsom, P. (2015) · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G. E., Vinyals, O., & Dean, J. (2015) · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S., & Szegedy, C. (2015) · 2015
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Joulin, A., & Mikolov, T. (2015) · 2015
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Karpathy, A., Johnson, J., & Fei-Fei, L. (2015) · 2015
Earlier work this paper cites.
Kurach, K., Andrychowicz, M., & Sutskever, I. (2015) · 2015
Earlier work this paper cites.
Skype translator: Breaking down language and hearing barriers
Lewis, W. D. (2015) · 2015
Earlier work this paper cites.
Character-based neural machine translation
Ling, W., Trancoso, I., Dyer, C., & Black, A. W. (2015) · 2015
Earlier work this paper cites.
Stanford neural machine translation systems for spoken language domains
Luong, M.-T., & Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Luong, M.-T., Pham, H., & Manning, C. D. (2015b) · 2015
Earlier work this paper cites.
Speed or accuracy? A study in evaluation of simultaneous speech translation
Mieno, T., Neubig, G., Sakti, S., Toda, T., & Nakamura, S. (2015) · 2015
Earlier work this paper cites.
Neural reranking improves subjective quality of machine translation: NAIST at WAT2015
Neubig, G., Morishita, M., & Nakamura, S. (2015) · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Ranzato, M., Chopra, S., Auli, M., & Zaremba, W. (2015) · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
Rush, A. M., Chopra, S., & Weston, J. (2015) · 2015
Earlier work this paper cites.
Neural responding machine for short-text conversation
Shang, L., Lu, Z., & Li, H. (2015) · 2015
Earlier work this paper cites.
Convolutional LSTM networks for subcellular localization of proteins
Sønderby, S. K., Sønderby, C. K., Nielsen, H., & Winther, O. (2015) · 2015
Earlier work this paper cites.
Data-free parameter pruning for deep neural networks
Srinivas, S., & Babu, R. V. (2015) · 2015
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Szlam, A., Weston, J., & Fergus, R. (2015) · 2015
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Tai, K. S., Socher, R., & Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Grammar as a foreign language
Vinyals, O., Kaiser, Ł., Koo, T., Petrov, S., Sutskever, I., & Hinton, G. E. (2015) · 2015
Earlier work this paper cites.
Describing videos by exploiting temporal structure
Yao, L., Torabi, A., Cho, K., Ballas, N., Pal, C., Larochelle, H., & Courville, A. (2015) · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., Wicke, M., Yu, Y., & Zheng, X. (2016) · 2016
Earlier work this paper cites.
Alignment-based neural machine translation
Alkhouli, T., Bretschner, G., Peter, J.-T., Hethnawi, M., Guta, A., & Ney, H. (2016) · 2016
Earlier work this paper cites.
Incorporating discrete translation lexicons into neural machine translation
Arthur, P., Neubig, G., & Nakamura, S. (2016) · 2016
Earlier work this paper cites.
Deeper machine translation and evaluation for German
Avramidis, E., Macketanz, V., Burchardt, A., Helcl, J., & Uszkoreit, H. (2016) · 2016
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016) · 2016
Earlier work this paper cites.
NoiseOut: A simple way to prune neural networks
Babaeizadeh, M., Smaragdis, P., & Campbell, R. H. (2016) · 2016
Earlier work this paper cites.
Neural versus phrase-based machine translation quality: A case study
Bentivogli, L., Bisazza, A., Cettolo, M., & Federico, M. (2016) · 2016
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Bojar, O., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huck, M., Jimeno Yepes, A., Koehn, P., Logacheva, V., Monz, C., Negri, M., Neveol, A., Neves, M., Popel, M., Post, M., Rubino, R., Scarton, C., Specia, L., Turchi, M., Verspoor, K., & Zampieri, M. (2016) · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
Chan, W., Jaitly, N., Le, Q. V., & Vinyals, O. (2016) · 2016
Earlier work this paper cites.
Chandar, S., Ahn, S., Larochelle, H., Vincent, P., Tesauro, G., & Bengio, Y. (2016) · 2016
Earlier work this paper cites.
Guided alignment training for topic-aware neural machine translation
Chen, W., Matusov, E., Khadivi, S., & Peter, J.-T. (2016) · 2016
Earlier work this paper cites.
Long short-term memory-networks for machine reading
Cheng, J., Dong, L., & Lapata, M. (2016a) · 2016
Cited alongside, same era.
Noisy parallel approximate decoding for conditional recurrent language model
Cho, K. (2016) · 2016
Cited alongside, same era.
Can neural machine translation do simultaneous translation?
Cho, K., & Esipova, M. (2016) · 2016
Cited alongside, same era.
A character-level decoder without explicit segmentation for neural machine translation
Chung, J., Cho, K., & Bengio, Y. (2016) · 2016
Cited alongside, same era.
Incorporating structural alignment biases into an attentional neural translation model
Cohn, T., Hoang, C. D. V., Vymolova, E., Yao, K., Dyer, C., & Haffari, G. (2016) · 2016
Cited alongside, same era.
Instance weighting for neural machine translation domain adaptation
Wang, R., Utiyama, M., Liu, L., Chen, K., & Sumita, E. (2017c) · 2017
Later among the works it cites.
Dynamic data selection for neural machine translation
van der Wees, M., Bisazza, A., & Monz, C. (2017) · 2017
Later among the works it cites.
Dual supervised learning
Xia, Y., Qin, T., Chen, W., Bian, J., Yu, N., & Liu, T.-Y. (2017) · 2017
Later among the works it cites.
Towards bidirectional hierarchical representations for attention-based neural machine translation
Yang, B., Wong, D. F., Xiao, T., Chao, L. S., & Zhu, J. (2017a) · 2017
Later among the works it cites.
SeqGAN: Sequence generative adversarial nets with policy gradient
Yu, L., Zhang, W., Wang, J., & Yu, Y. (2017) · 2017
Later among the works it cites.
To prune, or not to prune: exploring the efficacy of pruning for model compression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Crego, J., Kim, J., Klein, G., Rebollo, A., Yang, K., Senellart, J., Akhanov, E., Brunelle, P., Coquard, A., Deng, Y. et al. (2016) · 2016
Cited alongside, same era.
Kyoto university participation to WAT 2016
Cromieres, F., Chu, C., Nakazawa, T., & Kurohashi, S. (2016) · 2016
Cited alongside, same era.
Language to logical form with neural attention
Dong, L., & Lapata, M. (2016) · 2016
Cited alongside, same era.
An attentional model for speech translation without transcription
Duong, L., Anastasopoulos, A., Chiang, D., Bird, S., & Cohn, T. (2016) · 2016
Cited alongside, same era.
QCRI machine translation systems for IWSLT 16
Durrani, N., Dalvi, F., Sajjad, H., & Vogel, S. (2016) · 2016
Cited alongside, same era.
Recurrent neural network grammars
Dyer, C., Kuncoro, A., Ballesteros, M., & Smith, N. A. (2016) · 2016
Cited alongside, same era.
Attention pooling-based convolutional neural network for sentence modelling
Er, M. J., Zhang, Y., Wang, N., & Pratama, M. (2016) · 2016
Cited alongside, same era.
Zhu, M., & Gupta, S. (2017) · 2017
Later among the works it cites.
Differentiable lower bound for expected bleu score
Zhukov, V., Golikov, E., & Kretov, M. (2017) · 2017
Later among the works it cites.
On the alignment problem in multi-head attention-based neural machine translation
Alkhouli, T., Bretschner, G., & Ney, H. (2018) · 2018
Later among the works it cites.
Towards two-dimensional sequence to sequence model in neural machine translation
Bahar, P., Brix, C., & Ney, H. (2018) · 2018
Later among the works it cites.
Findings of the third shared task on multimodal machine translation
Barrault, L., Bougares, F., Specia, L., Lala, C., Elliott, D., & Frank, S. (2018) · 2018
Later among the works it cites.
Identifying and controlling important neurons in neural machine translation
Bau, A., Belinkov, Y., Sajjad, H., Durrani, N., Dalvi, F., & Glass, J. (2018) · 2018
Later among the works it cites.
Evaluating discourse phenomena in neural machine translation
Bawden, R., Sennrich, R., Birch, A., & Haddow, B. (2018) · 2018
Later among the works it cites.
Findings of the 2018 conference on machine translation (WMT18)
Bojar, O., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Koehn, P., & Monz, C. (2018) · 2018
Later among the works it cites.
Looking for ELMo’s friends: Sentence-level pretraining beyond language modeling
Bowman, S., Pavlick, E., Grave, E., van Durme, B., Wang, A., Hula, J., Xia, P., Pappagari, R., McCoy, R. T., Patel, R. et al. (2018) · 2018
Later among the works it cites.
Using monolingual data in neural machine translation: A systematic study
Burlot, F., & Yvon, F. (2018) · 2018
Later among the works it cites.
Caccia, M., Caccia, L., Fedus, W., Larochelle, H., Pineau, J., & Charlin, L. (2018) · 2018
Later among the works it cites.
RNNbow: Visualizing learning via backpropagation gradients in RNNs
Cashman, D., Patterson, G., Mosca, A., Watts, N., Robinson, S., & Chang, R. (2018) · 2018
Later among the works it cites.
A stable and effective learning strategy for trainable greedy decoding
Chen, Y., Li, V. O., Cho, K., & Bowman, S. (2018c) · 2018
Later among the works it cites.
Towards robust neural machine translation
Cheng, Y., Tu, Z., Meng, F., Zhai, J., & Liu, Y. (2018) · 2018
Later among the works it cites.
Revisiting character-based neural machine translation with capacity and compression
Cherry, C., Foster, G., Bapna, A., Firat, O., & Macherey, W. (2018) · 2018
Later among the works it cites.
Fine-grained attention mechanism for neural machine translation
Choi, H., Cho, K., & Bengio, Y. (2018b) · 2018
Later among the works it cites.
A comprehensive empirical comparison of domain adaptation methods for neural machine translation
Chu, C., Dabre, R., & Kurohashi, S. (2018) · 2018
Later among the works it cites.
A survey of domain adaptation for neural machine translation
Chu, C., & Wang, R. (2018) · 2018
Later among the works it cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., & Baroni, M. (2018) · 2018
Later among the works it cites.
Multi-source syntactic neural machine translation
Currey, A., & Heafield, K. (2018) · 2018
Later among the works it cites.
NeuroX: A toolkit for analyzing individual neurons in neural networks
Dalvi, F., Nortonsmith, A., Bau, A., Belinkov, Y., Sajjad, H., Durrani, N., & Glass, J. (2018) · 2018
Later among the works it cites.
Deep neural machine translation with weakly-recurrent units
Di Gangi, M. A., & Federico, M. (2018) · 2018
Later among the works it cites.
How much attention do you need? A granular analysis of neural machine translation architectures
Domhan, T. (2018) · 2018
Later among the works it cites.
How much does tokenization affect in neural machine translation?
Domingo, M., Garcıa-Martınez, M., Helle, A., & Casacuberta, F. (2018) · 2018
Later among the works it cites.
SMT versus NMT: Preliminary comparisons for irish
Dowling, M., Lynn, T., Poncelas, A., & Way, A. (2018) · 2018
Later among the works it cites.
What is in a translation unit? Comparing character and subword representations beyond translation
Durrani, N., Dalvi, F., Sajjad, H., Belinkov, Y., & Nakov, P. (2018) · 2018
Later among the works it cites.
Understanding back-translation at scale
Edunov, S., Ott, M., Auli, M., & Grangier, D. (2018a) · 2018
Later among the works it cites.
Classical structured prediction losses for sequence to sequence learning
Edunov, S., Ott, M., Auli, M., Grangier, D., & Ranzato, M. (2018b) · 2018
Later among the works it cites.
(self-attentive) autoencoder-based universal language representation for machine translation
Escolano, C., Costa-jussà, M. R., & Fonollosa, J. A. (2018) · 2018
Later among the works it cites.
Controllable abstractive summarization
Fan, A., Grangier, D., & Auli, M. (2018) · 2018
Later among the works it cites.
Pathologies of neural models make interpretations difficult
Feng, S., Wallace, E., Grissom II, A., Iyyer, M., Rodriguez, P., & Boyd-Graber, J. (2018b) · 2018
Later among the works it cites.
Adaptive multi-pass decoder for neural machine translation
Geng, X., Feng, X., Qin, B., & Liu, T. (2018) · 2018
Later among the works it cites.
A continuous relaxation of beam search for end-to-end training of neural sequence models
Goyal, K., Neubig, G., Dyer, C., & Berg-Kirkpatrick, T. (2018) · 2018
Later among the works it cites.
Non-autoregressive neural machine translation with enhanced decoder input
Guo, J., Tan, X., He, D., Qin, T., Xu, L., & Liu, T.-Y. (2018) · 2018
Later among the works it cites.
Achieving human parity on automatic Chinese to English news translation
Hassan, H., Aue, A., Chen, C., Chowdhary, V., Clark, J. H., Federmann, C., Huang, X., Junczys-Dowmunt, M., Lewis, W. D., Li, M. et al. (2018) · 2018
Later among the works it cites.
Non-adversarial unsupervised word translation
Hoshen, Y., & Wolf, L. (2018) · 2018
Later among the works it cites.
Reinforced mnemonic reader for machine reading comprehension
Hu, M., Peng, Y., Huang, Z., Qiu, X., Wei, F., & Zhou, M. (2018) · 2018
Later among the works it cites.
Accelerating NMT batched beam decoding with LMBR posteriors for deployment
Iglesias, G., Tambellini, W., de Gispert, A., Hasler, E., & Byrne, B. (2018) · 2018
Later among the works it cites.
Enhancement of encoder and attention using target monolingual corpora in neural machine translation
Imamura, K., Fujita, A., & Sumita, E. (2018) · 2018
Later among the works it cites.
English-Basque statistical and neural machine translation
Jauregi Unanue, I., Garmendia Arratibel, L., Zare Borzeshi, E., & Piccardi, M. (2018) · 2018
Later among the works it cites.
Fast decoding in sequence models using discrete latent variables
Kaiser, Ł., Roy, A., Vaswani, A., Pamar, N., Bengio, S., Uszkoreit, J., & Shazeer, N. (2018) · 2018
Later among the works it cites.
Deep reinforcement learning for sequence to sequence models
Keneshloo, Y., Shi, T., Reddy, C. K., & Ramakrishnan, N. (2018) · 2018
Later among the works it cites.
On the impact of various types of noise on neural machine translation
Khayrallah, H., & Koehn, P. (2018) · 2018
Later among the works it cites.
Regularized training objective for continued training for domain adaptation in neural machine translation
Khayrallah, H., Thompson, B., Duh, K., & Koehn, P. (2018) · 2018
Later among the works it cites.
Findings of the WMT 2018 shared task on parallel corpus filtering
Koehn, P., Khayrallah, H., Heafield, K., & Forcada, M. L. (2018) · 2018
Later among the works it cites.
Neural machine translation with adequacy-oriented learning
Kong, X., Tu, Z., Shi, S., Hovy, E., & Zhang, T. (2018) · 2018
Later among the works it cites.
Optimally segmenting inputs for NMT shows preference for character-level processing
Kreutzer, J., & Sokolov, A. (2018) · 2018
Later among the works it cites.
OpenSeq2Seq: Extensible toolkit for distributed and mixed precision training of sequence-to-sequence models
Kuchaiev, O., Ginsburg, B., Gitman, I., Lavrukhin, V., Case, C., & Micikevicius, P. (2018) · 2018
Later among the works it cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Kudo, T. (2018) · 2018
Later among the works it cites.
A comparison of Transformer and recurrent neural networks on multilingual neural machine translation
Lakew, S. M., Cettolo, M., & Federico, M. (2018) · 2018
Later among the works it cites.
Phrase-based & neural unsupervised machine translation
Lample, G., Ott, M., Conneau, A., Denoyer, L., & Ranzato, M. (2018) · 2018
Later among the works it cites.
Has machine translation achieved human parity? a case for document-level evaluation
Läubli, S., Sennrich, R., & Volk, M. (2018) · 2018
Later among the works it cites.
Learning hard alignments with variational inference
Lawson, D., Chiu, C.-C., Tucker, G., Raffel, C., Swersky, K., & Jaitly, N. (2018) · 2018
Later among the works it cites.
Deterministic non-autoregressive neural sequence modeling by iterative refinement
Lee, J., Mansimov, E., & Cho, K. (2018) · 2018
Later among the works it cites.
End-to-end non-autoregressive neural machine translation with connectionist temporal classification
Libovický, J., & Helcl, J. (2018) · 2018
Later among the works it cites.
Learning when to concentrate or divert attention: Self-adaptive attention temperature for neural machine translation
Lin, J., Sun, X., Ren, X., Li, M., & Su, Q. (2018a) · 2018
Later among the works it cites.
The mythos of model interpretability
Lipton, Z. C. (2018) · 2018
Later among the works it cites.
Controlling length in abstractive summarization using a convolutional neural network
Liu, Y., Luo, Z., & Zhu, K. (2018b) · 2018
Later among the works it cites.
A neural interlingua for multilingual machine translation
Lu, Y., Keung, P., Ladhak, F., Bhardwaj, V., Zhang, S., & Sun, J. (2018) · 2018
Later among the works it cites.
Morphological and language-agnostic word segmentation for NMT
Macháček, D., Vidra, J., & Bojar, O. (2018) · 2018
Later among the works it cites.
SMT vs NMT: a comparison over Hindi & Bengali simple sentences
Mahata, S. K., Mandal, S., Das, D., & Bandyopadhyay, S. (2018) · 2018
Later among the works it cites.
A smorgasbord of features to combine phrase-based and neural machine translation
Marie, B., & Fujita, A. (2018) · 2018
Later among the works it cites.
Document context neural machine translation with memory networks
Maruf, S., & Haffari, G. (2018) · 2018
Later among the works it cites.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., & Team, O. D. (2018) · 2018
Later among the works it cites.
Parallel attention mechanisms in neural machine translation
Medina, J. R., & Kalita, J. (2018) · 2018
Later among the works it cites.
Middle-out decoding
Mehri, S., & Sigal, L. (2018) · 2018
Later among the works it cites.
Neural machine translation with key-value memory-augmented attention
Meng, F., Tu, Z., Cheng, Y., Wu, H., Zhai, J., Yang, Y., & Wang, D. (2018) · 2018
Later among the works it cites.
MTNT: A testbed for machine translation of noisy text
Michel, P., & Neubig, G. (2018) · 2018
Later among the works it cites.
Self-attentive residual decoder for neural machine translation
Miculicich, L., Pappas, N., Ram, D., & Popescu-Belis, A. (2018a) · 2018
Later among the works it cites.
Document-level neural machine translation with hierarchical attention networks
Miculicich, L., Ram, D., Pappas, N., & Henderson, J. (2018b) · 2018
Later among the works it cites.
A large-scale test set for the evaluation of context-aware pronoun translation in neural machine translation
Müller, M., Rios, A., Voita, E., & Sennrich, R. (2018) · 2018
Later among the works it cites.
Correcting length bias in neural machine translation
Murray, K., & Chiang, D. (2018) · 2018
Later among the works it cites.
Rapid adaptation of neural machine translation to new languages
Neubig, G., & Hu, J. (2018) · 2018
Later among the works it cites.
XNMT: The eXtensible neural machine translation toolkit
Neubig, G., Sperber, M., Wang, X., Felix, M., Matthews, A., Padmanabhan, S., Qi, Y., Sachan, D., Arthur, P., Godard, P., Hewitt, J., Riad, R., & Wang, L. (2018) · 2018
Later among the works it cites.
Improving lexical choice in neural machine translation
Nguyen, T. Q., & Chiang, D. (2018) · 2018
Later among the works it cites.
Multi-source neural machine translation with data augmentation
Nishimura, Y., Sudoh, K., Neubig, G., & Nakamura, S. (2018) · 2018
Later among the works it cites.
Bi-Directional neural machine translation with synthetic parallel data
Niu, X., Denkowski, M., & Carpuat, M. (2018) · 2018
Later among the works it cites.
The RGNLP machine translation systems for WAT 2018
Ojha, A. K., Chowdhury, K. D., Liu, C.-H., & Saxena, K. (2018) · 2018
Later among the works it cites.
Recurrent relational networks
Palm, R., Paquet, U., & Winther, O. (2018) · 2018
Later among the works it cites.
NMT-Keras: A very flexible toolkit with a focus on interactive NMT and online learning
Álvaro Peris, & Casacuberta, F. (2018) · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., & Zettlemoyer, L. (2018) · 2018
Later among the works it cites.
Investigating backtranslation in neural machine translation
Poncelas, A., Shterionov, D., Way, A., Wenniger, G. M. d. B., & Passban, P. (2018) · 2018
Later among the works it cites.
Training tips for the Transformer model
Popel, M., & Bojar, O. (2018) · 2018
Later among the works it cites.
Pieces of eight: 8-bit neural machine translation
Quinn, J., & Ballesteros, M. (2018) · 2018
Later among the works it cites.
Improving language understanding with unsupervised learning
Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018) · 2018
Later among the works it cites.
Triangular architecture for rare language translation
Ren, S., Chen, W., Liu, S., Li, M., Zhou, M., & Ma, S. (2018) · 2018
Later among the works it cites.
Debugging neural machine translations
Rikters, M. (2018) · 2018
Later among the works it cites.
The RWTH Aachen University filtering system for the WMT 2018 parallel corpus filtering task
Rossenbach, N., Rosendahl, J., Kim, Y., Graça, M., Gokrani, A., & Ney, H. (2018) · 2018
Later among the works it cites.
Optimizing segmentation granularity for neural machine translation
Salesky, E., Runge, A., Coda, A., Niehues, J., & Neubig, G. (2018) · 2018
Later among the works it cites.
How to move to neural machine translation for enterprise-scale programs—an early adoption case study
Schmidt, T., & Marg, L. (2018) · 2018
Later among the works it cites.
Generative neural machine translation
Shah, H., & Barber, D. (2018) · 2018
Later among the works it cites.
Semi-supervised neural machine translation with language models
Skorokhodov, I., Rykachevskiy, A., Emelyanenko, D., Slotin, S., & Ponkratov, A. (2018) · 2018
Later among the works it cites.
Hybrid self-attention network for machine translation
Song, K., Xu, T., Peng, F., & Lu, J. (2018) · 2018
Later among the works it cites.
Findings of the WMT 2018 shared task on quality estimation
Specia, L., Blain, F., Logacheva, V., Astudillo, R., & Martins, A. F. T. (2018) · 2018
Later among the works it cites.
An operation sequence model for explainable neural machine translation
Stahlberg, F., Saunders, D., & Byrne, B. (2018c) · 2018
Later among the works it cites.
Blockwise parallel decoding for deep autoregressive models
Stern, M., Shazeer, N., & Uszkoreit, J. (2018) · 2018
Later among the works it cites.
Variational recurrent neural machine translation
Su, J., Wu, S., Xiong, D., Lu, Y., Han, X., & Zhang, B. (2018) · 2018
Later among the works it cites.
Lattice-to-sequence attentional neural machine translation models
Tan, Z., Su, J., Wang, B., Chen, Y., & Shi, X. (2018) · 2018
Later among the works it cites.
Why self-attention? A targeted evaluation of neural machine translation architectures
Tang, G., Müller, M., Rios, A., & Sennrich, R. (2018a) · 2018
Later among the works it cites.
Multi-domain neural machine translation
Tars, S., & Fishel, M. (2018) · 2018
Later among the works it cites.
Freezing subnetworks to analyze domain adaptation in neural machine translation
Thompson, B., Khayrallah, H., Anastasopoulos, A., McCarthy, A. D., Duh, K., Marvin, R., McNamee, P., Gwinnup, J., Anderson, T., & Koehn, P. (2018) · 2018
Later among the works it cites.
The importance of being recurrent for modeling hierarchical structure
Tran, K., Bisazza, A., & Monz, C. (2018) · 2018
Later among the works it cites.
Learning to remember translation history with a continuous cache
Tu, Z., Liu, Y., Shi, S., & Zhang, T. (2018) · 2018
Later among the works it cites.
Tensor2Tensor for neural machine translation
Vaswani, A., Bengio, S., Brevdo, E., Chollet, F., Gomez, A. N., Gouws, S., Jones, L., Kaiser, Ł., Kalchbrenner, N., Parmar, N., Sepassi, R., Shazeer, N., & Uszkoreit, J. (2018) · 2018
Later among the works it cites.
Context-aware neural machine translation learns anaphora resolution
Voita, E., Serdyukov, P., Sennrich, R., & Titov, I. (2018) · 2018
Later among the works it cites.
Statistical vs. neural machine translation: A comparison of mth and deepl at swiss post’s language service
Volkart, L., Bouillon, P., & Girletti, S. (2018) · 2018
Later among the works it cites.
Semi-autoregressive neural machine translation
Wang, C., Zhang, J., & Chen, H. (2018a) · 2018
Later among the works it cites.
Sentence selection and weighting for neural machine translation domain adaptation
Wang, R., Utiyama, M., Finch, A., Liu, L., Chen, K., & Sumita, E. (2018c) · 2018
Later among the works it cites.
SwitchOut: An efficient data augmentation algorithm for neural machine translation
Wang, X., Pham, H., Dai, Z., & Neubig, G. (2018e) · 2018
Later among the works it cites.
A tree-based decoder for neural machine translation
Wang, X., Pham, H., Yin, P., & Neubig, G. (2018f) · 2018
Later among the works it cites.
Incorporating statistical machine translation word knowledge into neural machine translation
Wang, X., Tu, Z., & Zhang, M. (2018g) · 2018
Later among the works it cites.
Global-context neural machine translation through target-side attentive residual connections
Werlen, L. M., Pappas, N., Ram, D., & Popescu-Belis, A. (2018) · 2018
Later among the works it cites.
Do latent tree learning models identify meaningful structure in sentences?
Williams, A., Drozdov, A., & Bowman, S. (2018) · 2018
Later among the works it cites.
A study of reinforcement learning for neural machine translation
Wu, L., Tian, F., Qin, T., Lai, J., & Liu, T.-Y. (2018a) · 2018
Later among the works it cites.
Phrase-level self-attention networks for universal sentence encoding
Wu, W., Wang, H., Liu, T., & Ma, S. (2018b) · 2018
Later among the works it cites.
Finding better subword segmentation for neural machine translation
Wu, Y., & Zhao, H. (2018) · 2018
Later among the works it cites.
Two effective approaches to data reduction for neural machine translation: Static and dynamic sentence selection
Xu, X., Kuang, S., & Xiong, D. (2018) · 2018
Later among the works it cites.
Breaking the beam search curse: A study of (re-)scoring methods and stopping criteria for neural machine translation
Yang, Y., Huang, L., & Ma, M. (2018b) · 2018
Later among the works it cites.
Regularizing forward and backward decoding to improve neural machine translation
Yang, Z., Chen, L., & Le Nguyen, M. (2018c) · 2018
Later among the works it cites.
Improving neural machine translation with conditional sequence generative adversarial nets
Yang, Z., Chen, W., Wang, F., & Xu, B. (2018d) · 2018
Later among the works it cites.
Sentence encoding with tree-constrained relation networks
Yu, L., d’Autume, C. d. M., Dyer, C., Blunsom, P., Kong, L., & Ling, W. (2018) · 2018
Later among the works it cites.
Incorporating syntactic uncertainty in neural machine translation with a forest-to-sequence model
Zaremoodi, P., & Haffari, G. (2018) · 2018
Later among the works it cites.
Improving the Transformer translation model with document-level context
Zhang, J., Luan, H., Sun, M., Zhai, F., Xu, J., Zhang, M., & Liu, Y. (2018b) · 2018
Later among the works it cites.
Exploring recombination for efficient decoding of neural machine translation
Zhang, Z., Wang, R., Utiyama, M., Sumita, E., & Zhao, H. (2018g) · 2018
Later among the works it cites.
Massively multilingual neural machine translation
Aharoni, R., Johnson, M., & Firat, O. (2019) · 2019
Closest in time.
Syntactically supervised Transformers for faster neural machine translation
Akoury, N., Krishna, K., & Iyyer, M. (2019) · 2019
Closest in time.
Analyzing and interpreting neural networks for NLP: A report on the first BlackboxNLP workshop
Alishahi, A., Chrupala, G., & Linzen, T. (2019) · 2019
Closest in time.
Findings of the 2019 conference on machine translation (WMT19)
Bojar, O. et al. (2019) · 2019
Closest in time.
An error analysis for image-based multi-modal neural machine translation
Calixto, I., & Liu, Q. (2019) · 2019
Closest in time.
What is one grain of sand in the desert? Analyzing individual neurons in deep NLP models
Dalvi, F., Durrani, N., Sajjad, H., Belinkov, Y., Bau, A., & Glass, J. (2019) · 2019
Closest in time.
BERT: Pre-training of deep bidirectional Transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019) · 2019
Closest in time.
Enhancing translation from English to Arabic using two-phase decoder translation
ElMaghraby, A., & Rafea, A. (2019) · 2019
Closest in time.
Incorporating source-side phrase structures into neural machine translation
Eriguchi, A., Hashimoto, K., & Tsuruoka, Y. (2019) · 2019
Closest in time.
Attention is not Explanation
Jain, S., & Wallace, B. C. (2019) · 2019
Closest in time.
Post-editing neural machine translation versus phrase-based machine translation for English–Chinese
Jia, Y., Carl, M., & Wang, X. (2019) · 2019
Closest in time.
Knowledge distillation using output errors for self-attention end-to-end models
Kim, H.-G., Na, H., Lee, H., Lee, J., Kang, T. G., Lee, M.-J., & Choi, Y. S. (2019a) · 2019
Closest in time.
Text Generation from Knowledge Graphs with Graph Transformers
Koncel-Kedziorski, R., Bekal, D., Luan, Y., Lapata, M., & Hajishirzi, H. (2019) · 2019
Closest in time.
Selective attention for context-aware neural machine translation
Maruf, S., Martins, A. F. T., & Haffari, G. (2019) · 2019
Closest in time.
On evaluation of adversarial perturbations for sequence-to-sequence models
Michel, P., Li, X., Neubig, G., & Pino, J. (2019) · 2019
Closest in time.
Addressing word-order divergence in multilingual neural machine translation for extremely low resource languages
Murthy, R., Kunchukuttan, A., & Bhattacharyya, P. (2019) · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., & Auli, M. (2019) · 2019
Closest in time.
Competence-based curriculum learning for neural machine translation
Platanios, E. A., Stretcu, O., Neubig, G., Poczos, B., & Mitchell, T. (2019) · 2019
Closest in time.
Text normalization using memory augmented neural networks
Pramanik, S., & Hussain, A. (2019) · 2019
Closest in time.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019) · 2019
Closest in time.
Domain adaptive inference for neural machine translation
Saunders, D., de Gispert, A., Stahlberg, F., & Byrne, B. (2019) · 2019
Closest in time.
Is attention interpretable?
Serrano, S., & Smith, N. A. (2019) · 2019
Closest in time.
Ordered neurons: Integrating tree structures into recurrent neural networks
Shen, Y., Tan, S., Sordoni, A., & Courville, A. (2019) · 2019
Closest in time.
On NMT search errors and model errors: Cat got your tongue?
Stahlberg, F., & Byrne, B. (2019) · 2019
Closest in time.
CUED@WMT19:EWC&LMs
Stahlberg, F., Saunders, D., de Gispert, A., & Byrne, B. (2019) · 2019
Closest in time.
Positional encoding to control output sequence length
Takase, S., & Okazaki, N. (2019) · 2019
Closest in time.
Extract and edit: An alternative to back-translation for unsupervised neural machine translation
Wu, J., Wang, X., & Wang, W. Y. (2019b) · 2019
Closest in time.
Towards string-to-tree neural machine translation
Aharoni, R., & Goldberg, Y. (2017b) · 2021
Closest in time.
Vocabulary manipulation for neural machine translation
Mi, H., Wang, Z., & Ittycheriah, A. (2016c) · 2021
Closest in time.
Natural language inference by tree-based convolution and heuristic matching
Mou, L., Men, R., Li, G., Xu, Y., Zhang, L., Yan, R., & Jin, Z. (2016) · 2022
Closest in time.
CytonMT: An efficient neural machine translation open-source toolkit implemented in C++
Wang, X., Utiyama, M., & Sumita, E. (2018h) · 2023
Closest in time.
Near human-level performance in grammatical error correction with hybrid machine translation
Grundkiewicz, R., & Junczys-Dowmunt, M. (2018) · 2046
Closest in time.
Boosting neural machine translation
Zhang, D., Kim, J., Crego, J., & Senellart, J. (2017b) · 2046
Closest in time.
A simple and effective approach to coverage-aware neural machine translation
Li, Y., Xiao, T., Li, Y., Wang, Q., Xu, C., & Zhu, J. (2018) · 2047
Closest in time.
Key-value attention mechanism for neural machine translation
Mino, H., Utiyama, M., Sumita, E., & Tokunaga, T. (2017) · 2049
Closest in time.
Syntactically guided neural machine translation
Stahlberg, F., Hasler, E., Waite, A., & Byrne, B. (2016b) · 2049
Closest in time.
Transfer learning across low-resource, related languages for neural machine translation
Nguyen, T. Q., & Chiang, D. (2017) · 2050
Closest in time.
Multi-representation ensembles and delayed SGD updates improve syntax-based NMT
Saunders, D., Stahlberg, F., de Gispert, A., & Byrne, B. (2018) · 2051
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J. L., Kiros, J. R., Cho, K., Courville, A., Salakhutdinov, R., Zemel, R., & Bengio, Y. (2015) · 2057
Closest in time.
Character-based neural machine translation
Costa-jussà, M. R., & Fonollosa, J. A. (2016) · 2058
Closest in time.
Neural machine translation by minimising the Bayes-risk with respect to syntactic translation lattices
Stahlberg, F., de Gispert, A., Hasler, E., & Byrne, B. (2017a) · 2058
Closest in time.
Sparse and constrained attention for neural machine translation
Malaviya, C., Ferreira, P., & Martins, A. F. T. (2018) · 2059
Closest in time.
How grammatical is character-level neural machine translation? Assessing MT quality with contrastive translation pairs
Sennrich, R. (2017) · 2060
Closest in time.
Neural system combination for machine translation
Zhou, L., Hu, W., Zhang, J., & Zong, C. (2017) · 2060
Closest in time.
Bilingual data cleaning for SMT using graph-based random walk
Cui, L., Zhang, D., Liu, S., Li, M., & Zhou, M. (2013) · 2061
Closest in time.
Reinforcement learning based curriculum optimization for neural machine translation
Kumar, G., Foster, G., Cherry, C., & Krikun, M. (2019) · 2061
Closest in time.
Neural machine translation with recurrent attention modeling
Yang, Z., Hu, Z., Deng, Y., Dyer, C., & Smola, A. (2017b) · 2061
Closest in time.
Overcoming catastrophic forgetting during domain adaptation of neural machine translation
Thompson, B., Gwinnup, J., Khayrallah, H., Duh, K., & Koehn, P. (2019) · 2068
Closest in time.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., & Vaswani, A. (2018) · 2074
Closest in time.
Exploiting semantics in neural machine translation with graph convolutional networks
Marcheggiani, D., Bastings, J., & Titov, I. (2018) · 2078
Closest in time.
Learning hidden unit contribution for adapting neural machine translation models
Vilar, D. (2018) · 2080
Closest in time.
Neural machine translation decoding with terminology constraints
Hasler, E., de Gispert, A., Iglesias, G., & Byrne, B. (2018) · 2081
Closest in time.
Sentence embedding for neural machine translation domain adaptation
Wang, R., Finch, A., Utiyama, M., & Sumita, E. (2017b) · 2089
Closest in time.
Variable-length word encodings for neural translation models
Chitnis, R., & DeNero, J. (2015) · 2093
Closest in time.
Continuous space language models for statistical machine translation
Schwenk, H., Dechelotte, D., & Gauvain, J.-L. (2006) · 2093
Closest in time.