Fetching the paper…
Reading the bibliography…
We present Latin BERT, a contextual language model for the Latin language, trained on 642.7 million words from a variety of sources spanning the Classical era to the 21st century.
Wanrong Zhu, Zhiting Hu, and Eric Xing. 2019 · 1901
Earlier work this paper cites.
Understanding the behaviors of BERT in ranking
Yifan Qiao, Chenyan Xiong, Zhenghao Liu, and Zhiyuan Liu. 2019 · 1904
Earlier work this paper cites.
Adaptation of deep bidirectional multilingual transformers for Russian language
Yuri Kuratov and Mikhail Arkhipov. 2019 · 1905
Earlier work this paper cites.
On the Feasibility of Automated Detection of Allusive Text Reuse
Enrique Manjavacas, Brian Long, and Mike Kestemont. 2019 · 1905
Earlier work this paper cites.
Pre-training with whole word masking for Chinese BERT
Yiming Cui, Wanxiang Che, Ting Liu, Bing Qin, Ziqing Yang, Shijin Wang, and Guoping Hu. 2019 · 1906
Earlier work this paper cites.
Milan Straka, Jana Straková, and Jan Hajič. 2019 · 1908
Earlier work this paper cites.
Rethinking self-attention: Towards interpretability in neural parsing
Khalil Mrini, Franck Dernoncourt, Trung Bui, Walter Chang, and Ndapa Nakashole. 2019 · 1911
Earlier work this paper cites.
FlauBERT: Unsupervised language model pre-training for French
Hang Le, Loïc Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoît Crabbé, Laurent Besacier, and Didier Schwab. 2019 · 1912
Earlier work this paper cites.
Textual Criticism and Editorial Technique Applicable to Greek and Latin Texts
Martin Litchfield West. 1973 · 1973
Earlier work this paper cites.
Building a digital library: The Perseus Project as a case study in the humanities
Gregory Crane. 1996 · 1996
Earlier work this paper cites.
A Greek-English Lexicon, 9th edition
Henry George Liddell, Robert Scott, Henry Stuart Jones, and Robert McKenzie, editors. 1996 · 1996
Earlier work this paper cites.
The Perseus Project: A digital library for the humanities
David A. Smith, Jeffrey A. Rydberg-Cox, and Gregory Crane. 2000 · 2000
Earlier work this paper cites.
What the [MASK]? making sense of language-specific BERT models
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2020 · 2003
Earlier work this paper cites.
The design and use of a Latin dependency treebank
David Bamman and Gregory Crane. 2006 · 2006
Earlier work this paper cites.
Latein ist tot, es lebe Latein!: kleine Geschichte einer grossen Sprache
Wilfried Stroh. 2007 · 2007
Earlier work this paper cites.
Creating a parallel treebank of the old Indo-European Bible translations
Dag T. T. Haug and Marius L. Jøhndal. 2008 · 2008
Earlier work this paper cites.
Visualizing data using t-SNE
Laurens van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Improvements in Parsing the Index Thomisticus Treebank. Revision, Combination and a Feature Model for Medieval Latin
Marco Passarotti and Felice Dell’Orletta. 2010 · 2010
Earlier work this paper cites.
Extracting two thousand years of Latin from a million book library
David Bamman and David Smith. 2011 · 2011
Earlier work this paper cites.
Intertextuality in the digital age
Neil Coffee, Jean-Pierre Koenig, Shakthi Poornima, Roelant Ossewaarde, Christopher Forstall, and Sarah Jacobson. 2012 · 2012
Earlier work this paper cites.
Latin: Story of a World Language
Jürgen Leonhardt. 2013 · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Cited alongside, same era.
Methods in Latin Computational Linguistics
Barbara McGillivray. 2014 · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Word embeddings pointing the way for late antiquity
Johannes Bjerva and Raf Praet. 2015 · 2015
Cited alongside, same era.
Massively multilingual word embeddings
Waleed Ammar, George Mulcaire, Yulia Tsvetkov, Guillaume Lample, Chris Dyer, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
Rethinking intertextuality through a word-space and social network approach—the case of Cassiodorus
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning. 2019 · 2019
Later among the works it cites.
GlossBERT: BERT for word sense disambiguation with gloss knowledge
Luyao Huang, Chi Sun, Xipeng Qiu, and Xuanjing Huang. 2019 · 2019
Later among the works it cites.
Vector space models of Ancient Greek word meaning, and a case study on Homer
Martina Astrid Rodda, Philomen Probert, and Barbara McGillivray. 2019 · 2019
Later among the works it cites.
Vir is to Moderatus as Mulier is to Intemperans: Lemma embeddings for Latin
Rachele Sprugnoli, Marco Passarotti, and Giovanni Moretti. 2019 · 2019
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Data-driven choices in neural part-of-speech tagging for Latin
Geoff Bacon, Clayton Marr, and David Mortensen. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Johannes Bjerva and Raf Praet. 2016 · 2016
Cited alongside, same era.
Non-literal text reuse in historical texts: An approach to identify reuse transformations and its application to Bible reuse
Maria Moritz, Andreas Wiederhold, Barbara Pavlek, Yuri Bizzoni, and Marco Büchler. 2016 · 2016
Cited alongside, same era.
Quantitative criticism of literary relationships
Joseph P. Dexter, Theodore Katz, Nilesh Tripuraneni, Tathagata Dasgupta, Ajay Kannan, James A. Brofos, Jorge A. Bonilla Lopez, Lea A. Schroeder, Adriana Casarez, Maxim Rabinovich, Ayelet Haimson Lushkov, and Pramit Chaudhuri. 2017 · 2017
Cited alongside, same era.
Word sense disambiguation: A unified evaluation framework and empirical comparison
Alessandro Raganato, Jose Camacho-Collados, and Roberto Navigli. 2017 · 2017
Cited alongside, same era.
NLP-cube: End-to-end raw text processing with neural networks
Tiberiu Boros, Stefan Daniel Dumitrescu, and Ruxandra Burtica. 2018 · 2018
Cited alongside, same era.
An Agenda for the Study of Intertextuality
Neil Coffee. 2018 · 2018
Cited alongside, same era.
Learning word vectors for 157 languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov. 2018 · 2018
Cited alongside, same era.
Closest in time.
Distributional semantics for Neo-Latin
Jelke Bloem, Maria Chiara Parisi, Martin Reynaert, Yvette Oortwijn, Arianna Betti, Clayton Marr, and David Mortensen. 2020 · 2020
Closest in time.
Spanish pre-trained BERT model and evaluation data
José Cañete, Gabriel Chaperon, Rodrigo Fuentes, and Jorge Pérez. 2020 · 2020
Closest in time.
A new Latin treebank for Universal Dependencies: Charters between ancient Latin and Romance languages
Flavio Massimiliano Cecchini, Timo Korkiakangas, and Marco Passarotti. 2020 · 2020
Closest in time.
A gradient boosting–Seq2Seq system for Latin POS tagging and lemmatization
Giuseppe G. A. Celano, Clayton Marr, and David Mortensen. 2020 · 2020
Closest in time.
Enabling language models to fill in the blanks
Chris Donahue, Mina Lee, and Percy Liang. 2020 · 2020
Closest in time.
The Classical Language Toolkit
Kyle P Johnson. 2020 · 2020
Closest in time.
SpanBERT: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Closest in time.
Measuring philosophy in the first thousand years of Greek literature
Thomas Köntges. 2020 · 2020
Closest in time.
LiLa: Linking Latin. Risorse linguistiche per il latino nel Semantic Web
Francesco Mambrini, Flavio Massimiliano Cecchini, Greta Franzini, Eleonora Litta, Marco Carlo Passarotti, and Paolo Ruffolo. 2020 · 2020
Closest in time.
CamemBERT: a tasty French language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020 · 2020
Closest in time.
Latin WordNet
William M. Short. 2020 · 2020
Closest in time.
Overview of the EvaLatin 2020 evaluation campaign
Rachele Sprugnoli, Marco Passarotti, Flavio M Cecchini, Matteo Pellegrini, Clayton Marr, and David Mortensen. 2020 · 2020
Closest in time.
Like Two Pis in a Pod: Author Similarity Across Time in the Ancient Greek Corpus
Grant Storey and David Mimno. 2020 · 2020
Closest in time.
UDPipe at EvaLatin 2020: Contextualized Embeddings and Treebank Embeddings
Milan Straka, Jana Straková, Clayton Marr, and David Mortensen. 2020 · 2020
Closest in time.