Fetching the paper…
Reading the bibliography…
The notion of "in-domain data" in NLP is often over-simplistic and vague, as textual data varies in many nuanced linguistic aspects such as topic, style or level of formality.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Data selection with cluster-based language difference models and cynical selection
Lucía Santamaría and Amittai Axelrod. 2019 · 1904
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2019a · 1905
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. 2019 · 1906
Earlier work this paper cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Scalable evaluation and improvement of document set expansion via neural positive-unlabeled learning
Alon Jacovi, Gang Niu, Yoav Goldberg, and Masashi Sugiyama. 2019 · 1910
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, R’emi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Emerging cross-lingual structure in pretrained language models
Shijie Wu, Alexis Conneau, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Curriculum learning for domain adaptation in neural machine translation
Xuan Zhang, Pamela Shapiro, Gaurav Kumar, Paul McNamee, Marine Carpuat, and Kevin Duh. 2019 · 1915
Earlier work this paper cites.
Non-parametric adaptation for neural machine translation
Ankur Bapna and Orhan Firat. 2019 · 1931
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan. 2003 · 2003
Earlier work this paper cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2003
Earlier work this paper cites.
When and why is unsupervised neural machine translation useless?
Yunsu Kim, Miguel Graça, and Hermann Ney. 2020 · 2004
Earlier work this paper cites.
When does unsupervised machine translation work?
Kelly Marchisio, Kevin Duh, and Philipp Koehn. 2020 · 2004
Earlier work this paper cites.
The frequency and use of lexical bundles in conversation and academic prose
Susan M Conrad and Douglas Biber. 2005 · 2005
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
Introduction to information retrieval
Christopher D Manning, Prabhakar Raghavan, and Hinrich Schütze. 2008 · 2008
Earlier work this paper cites.
K-means vs gmm, sum-product vs max-product
Hal Daume. 2009 · 2009
Earlier work this paper cites.
Active learning for statistical phrase-based machine translation
Gholamreza Haffari, Maxim Roy, and Anoop Sarkar. 2009 · 2009
Earlier work this paper cites.
Intelligent selection of language model training data
Robert C. Moore and William Lewis. 2010 · 2010
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Amittai Axelrod, Xiaodong He, and Jianfeng Gao. 2011 · 2011
Earlier work this paper cites.
KenLM: Faster and smaller language model queries
Kenneth Heafield. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Cited alongside, same era.
Does more data always yield better translations?
Guillem Gascó, Martha-Alicia Rocha, Germán Sanchis-Trilles, Jesús Andrés-Ferrer, and Francisco Casacuberta. 2012 · 2012
Cited alongside, same era.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Cited alongside, same era.
Adaptation data selection using neural language models: Experiments in machine translation
Kevin Duh, Graham Neubig, Katsuhito Sudoh, and Hajime Tsukada. 2013 · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Cited alongside, same era.
Dynamic topic adaptation for phrase-based MT
A survey of domain adaptation for neural machine translation
Chenhui Chu and Rui Wang. 2018 · 2018
Later among the works it cites.
Dual conditional cross-entropy filtering of noisy parallel corpora
Marcin Junczys-Dowmunt. 2018 · 2018
Later among the works it cites.
Hallucinations in neural machine translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2018 · 2018
Later among the works it cites.
Data selection for nmt using infrequent n-gram recovery
Zuzanna Parcheta, Germán Sanchis-Trilles, and Francisco Casacuberta. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Data selection with feature decay algorithms using an approximated target side
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eva Hasler, Phil Blunsom, Philipp Koehn, and Barry Haddow. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Survey of data-selection methods in statistical machine translation
Sauleh Eetemadi, William Lewis, Kristina Toutanova, and Hayder Radha. 2015 · 2015
Cited alongside, same era.
What’s in a domain? analyzing genre and topic differences in statistical machine translation
Marlies van der Wees, Arianna Bisazza, Wouter Weerkamp, and Christof Monz. 2015 · 2015
Cited alongside, same era.
Adapting to all domains at once: Rewarding domain invariance in SMT
Hoang Cuong, Khalil Sima’an, and Ivan Titov. 2016 · 2016
Cited alongside, same era.
Data selection for IT texts using paragraph vector
Mirela-Stefania Duma and Wolfgang Menzel. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Alberto Poncelas, Gideon Maillette de Buy Wenniger, and Andy Way. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Extracting in-domain training corpora for neural machine translation using data selection methods
Catarina Cruz Silva, Chao-Hong Liu, Alberto Poncelas, and Andy Way. 2018 · 2018
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
The OPUS resource repository: An open package for creating parallel corpora and machine translation services
Mikko Aulamo and Jörg Tiedemann. 2019 · 2019
Later among the works it cites.
Findings of the 2019 conference on machine translation (WMT19)
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019 · 2019
Later among the works it cites.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Unsupervised domain adaptation for neural machine translation with domain-aware feature embeddings
Zi-Yi Dou, Junjie Hu, Antonios Anastasopoulos, and Graham Neubig. 2019a · 2019
Later among the works it cites.
Domain adaptation of neural machine translation by lexicon induction
Junjie Hu, Mengzhou Xia, Graham Neubig, and Jaime Carbonell. 2019 · 2019
Later among the works it cites.
Domain adaptation with BERT-based domain classification and data selection
Xiaofei Ma, Peng Xu, Zhiguo Wang, Ramesh Nallapati, and Bing Xiang. 2019 · 2019
Later among the works it cites.
Domain robustness in neural machine translation
Mathias Müller, Annette Rios, and Rico Sennrich. 2019 · 2019
Later among the works it cites.
Facebook FAIR’s WMT19 news translation task submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 2019
Later among the works it cites.
Transfer learning in natural language processing
Sebastian Ruder, Matthew E. Peters, Swabha Swayamdipta, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Closest in time.