Fetching the paper…
Reading the bibliography…
Pre-training and fine-tuning, e.g., BERT, have achieved great success in language understanding by transferring knowledge from rich-resource pre-training task to the low/zero-resource downstream tasks.
Class-based n-gram models of natural language
Brown, P. F., Desouza, P. V., Mercer, R. L., Pietra, V. J. D., and Lai, J. C · 1992
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Y., Ducharme, R., Vincent, P., and Jauvin, C · 2003
Earlier work this paper cites.
English gigaword
Graff, David, Kong, Junbo, Chen, Ke, Maeda, and Kazuaki · 2003
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Tjong Kim Sang, E. F. and De Meulder, F · 2003
Earlier work this paper cites.
A framework for learning predictive structures from multiple tasks and unlabeled data
Ando, R. K. and Zhang, T · 2005
Earlier work this paper cites.
Domain adaptation with structural correspondence learning
Blitzer, J., McDonald, R., and Pereira, F · 2006
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, R. and Weston, J · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Vincent, P., Larochelle, H., Bengio, Y., and Manzagol, P.-A · 2008
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Danescu-Niculescu-Mizil, C. and Lee, L · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T., Sutskever, I., Chen, K., Corrado, G. S., and Dean, J · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., van Merrienboer, B., Gülçehre, Ç., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J · 2014
Earlier work this paper cites.
Distributed representations of sentences and documents
Le, Q. and Mikolov, T · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Dai, A. M. and Le, Q. V · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Cited alongside, same era.
Learning deep representation with large-scale attributes
Ouyang, W., Li, H., Zeng, X., and Wang, X · 2015
Cited alongside, same era.
Neural responding machine for short-text conversation
Shang, L., Lu, Z., and Li, H · 2015
Cited alongside, same era.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2015
Cited alongside, same era.
Vinyals, O. and Le, Q. V · 2015
Cited alongside, same era.
Neural headline generation with minimum risk training
Ayana, Shen, S., Liu, Z., and Sun, M · 2016
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
Machine comprehension by text-to-text neural question generation
Yuan, X., Wang, T., Gulcehre, C., Sordoni, A., Bachman, P., Zhang, S., Subramanian, S., and Trischler, A · 2017
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Later among the works it cites.
Maskgan: Better text generation via filling in the ______
Fedus, W., Goodfellow, I., and Dai, A · 2018
Later among the works it cites.
Achieving human parity on automatic chinese to english news translation
Hassan, H., Aue, A., Chen, C., Chowdhary, V., Clark, J., Federmann, C., Huang, X., Junczys-Dowmunt, M., Lewis, W., Li, M., et al · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zero-resource translation with multi-lingual neural machine translation
Firat, O., Sankaran, B., Al-Onaizan, Y., Vural, F. T. Y., and Cho, K · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Cited alongside, same era.
Unsupervised pretraining for sequence to sequence learning
Ramachandran, P., Liu, P. J., and Le, Q. V · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Wu, Y., Schuster, M., Chen, Z., Le, Q. V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., Macherey, K., Klingner, J., Shah, A., Johnson, M., Liu, X., Kaiser, L., Gouws, S., Kato, Y., Kudo, T., Kazawa, H., Stevens, K., Kurian, G., Patil, N., Wang, W., Young, C., Smith, J., Riesa, J., Rudnick, A., Vinyals, O., Corrado, G., Hughes, M., and Dean, J · 2016
Cited alongside, same era.
Layer-wise coordination between encoder and decoder for neural machine translation
He, T., Tan, X., Xia, Y., He, D., Qin, T., Chen, Z., and Liu, T.-Y · 2018
Later among the works it cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S · 2018
Later among the works it cites.
Phrase-based & neural unsupervised machine translation
Lample, G., Ott, M., Conneau, A., Denoyer, L., and Ranzato, M · 2018
Later among the works it cites.
An efficient framework for learning sentence representations
Logeswaran, L. and Lee, H · 2018
Later among the works it cites.
Deep contextualized word representations
Peters, M., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Later among the works it cites.
Dense information flow for neural machine translation
Shen, Y., Tan, X., He, D., Qin, T., and Liu, T.-Y · 2018
Later among the works it cites.
Double path networks for sequence to sequence learning
Song, K., Tan, X., He, D., Lu, J., Qin, T., and Liu, T.-Y · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Later among the works it cites.
Unsupervised neural machine translation with weight sharing
Yang, Z., Chen, W., Wang, F., and Xu, B · 2018
Later among the works it cites.
Cross-lingual language model pretraining
Lample, G. and Conneau, A · 2019
Closest in time.
Unsupervised pivot translation for distant languages
Leng, Y., Tan, X., Qin, T., Li, X.-Y., and Liu, T.-Y · 2019
Closest in time.
Almost unsupervised text to speech and automatic speech recognition
Ren, Y., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2019
Closest in time.
Multilingual neural machine translation with knowledge distillation
Tan, X., Ren, Y., He, D., Qin, T., and Liu, T.-Y · 2019
Closest in time.