Adaptive input representations for neural language modeling
Original
Alexei Baevski and Michael Auli. 2018 · 2018
Later among the works it cites.
An empirical evaluation of generic convolutional and recurrent networks for sequence modeling
Original
Shaojie Bai, J Zico Kolter, and Vladlen Koltun. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Maskgan: better text generation via filling in the _
Original
William Fedus, Ian Goodfellow, and Andrew M Dai. 2018 · 2018
Later among the works it cites.
Delete, retrieve, generate: a simple approach to sentiment and style transfer
Juncen Li, Robin Jia, He He, and Percy Liang. 2018 · 2018
Later among the works it cites.
Unpaired sentiment-to-sentiment translation: A cycled reinforcement learning approach
Jingjing Xu, Xu Sun, Qi Zeng, Xiaodong Zhang, Xuancheng Ren, Houfeng Wang, and Wenjie Li. 2018 · 2018
Later among the works it cites.
Style transfer as unsupervised machine translation
Original
Zhirui Zhang, Shuo Ren, Shujie Liu, Jianyong Wang, Peng Chen, Mu Li, Ming Zhou, and Enhong Chen. 2018 · 2018
Later among the works it cites.
Restoring ancient text using deep learning: a case study on Greek epigraphy
Yannis Assael, Thea Sommerschield, and Jonathan Prag. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Closest in time.
Decoding as dynamic programming for recurrent autoregressive models
Najam Zaidi, Trevor Cohn, and Gholamreza Haffari. 2020 · 2020
Closest in time.