Fetching the paper…
Reading the bibliography…
The IMPACT-es diachronic corpus of historical Spanish compiles over one hundred books --containing approximately 8 million words-- in addition to a complementary lexicon which links more than 10 thousand lemmas with attestations of the different variants found in the documents.
Levenshtein VI (1965) Binary codes capable of correcting deletions, insertions, and reversals. Doklady Akademii Nauk SSSR 163(4):845–848, English translation in Soviet Physics Doklady, 10(8), 707–710, 1966
1966
Earlier work this paper cites.
Hirschberg DS (1975) A linear space algorithm for computing maximal common subsequences. Communications of the ACM 18(6):341–343
1975
Earlier work this paper cites.
Lowrance R, Wagner RA (1975) An extension of the string-to-string correction problem. Journal of the Association for Computing Machinery 22(2):177–183
1975
Earlier work this paper cites.
Francis WN, Kucera H (1979) Brown corpus manual. Online at http://www.hit.uib.no/icame/brown/bcm.html
1979
Earlier work this paper cites.
Oncina J, Garcia P, Vidal E (1993) Learning subsequential transducers for pattern recognition interpretation tasks. IEEE Transactions on Pattern Analysis and Machine Intelligence 15(5):448–458
1993
Earlier work this paper cites.
Bengio Y, Frasconi P (1994) An input output HMM architecture. In: Advances in Neural Information Processing Systems 7, [NIPS Conference, Denver, Colorado, USA, 1994], MIT Press, pp 427–434
1994
Earlier work this paper cites.
Bengio S, Bengio Y (1996) An EM algorithm for asynchronous input/output hidden Markov models. In: Proceedings of the International Conference on Neural Information Processing, ICONIP, Hong Kong, pp 328–334
1996
Earlier work this paper cites.
Matthews PH (1997) The Concise Oxford Dictionary of Linguistics. Oxford University Press
1997
Earlier work this paper cites.
Manning CD, Schütze H (1999) Foundations of Statistical Natural Language Processing. MIT Press, Cambridge, MA, USA
1999
Earlier work this paper cites.
Davies M (2002) Un corpus anotado de 100.000.000 palabras del español histórico y moderno. Procesamiento del Lenguaje Natural 29:21–27
2002
Earlier work this paper cites.
Papineni K, Roukos S, Ward T, Zhu WJ (2002) BLEU: a method for automatic evaluation of machine translation. In: Proceedings of the 40th Annual meeting of the Association for Computational Linguistics, Philadelphia, USA, pp 311–318
2002
Earlier work this paper cites.
Zens R, Och FJ, Ney H (2002) Phrase-based statistical machine translation. In: Proceedings 25th Annual German Conference on Advances in Artificial Intelligence, Lecture Notes in Computer Science, vol 2479, Springer-Verlag, pp 18–32
2002
Earlier work this paper cites.
Och FJ (2003) Minimum error rate training in statistical machine translation. In: Proceedings of the 41st Annual Meeting of the Association for Computational Linguistics, Sapporo, Japan, pp 160–167
2003
Cited alongside, same era.
Och FJ, Ney H (2003) A systematic comparison of various statistical alignment models. Computational Linguistics 29(1):19–51
2003
Cited alongside, same era.
Carreras X, Chao I, Padró L, Padró M (2004) FreeLing: an open-source suite of language analyzers. In: Proceedings of the 4th International Conference on Language Resources and Evaluation, Lisbon, Portugal, pp 239–242
2004
Cited alongside, same era.
Procházková P (2006) Fundamentos de la lingüística de corpus. Concepción de los corpus y métodos de investigación con corpus. Available online at http://prochazkova.de/fundamentos_de_la_lingüística_de_corpus.pdf
2006
Cited alongside, same era.
Sánchez-Marco C, Boleda G, Fontana JM, Domingo J (2010) Annotation and representation of a diachronic corpus of Spanish. In: Proceedings of the 7th International Conference on Language Resources and Evaluation, La Valleta, Malta, pp 2713–2718
2010
Later among the works it cites.
Erjavec T (2011) Automatic linguistic annotation of historical language: ToTrTaLe and XIX century Slovene. In: Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities, Association for Computational Linguistics, Portland, OR, USA, pp 33–38
2011
Later among the works it cites.
Forcada ML, Ginestí-Rosell M, Nordfalk J, O’Regan J, Ortiz-Rojas S, Pérez-Ortiz JA, Sánchez-Martínez F, Ramírez-Sánchez G, Tyers FM (2011) Apertium: a free/open-source platform for rule-based machine translation. Machine Translation 25(2):127–144
2011
Later among the works it cites.
Medina Urrea A, Méndez Cruz CF (2011) El corpus histórico del español en México. Revista Digital Universitaria 12(7):3–25
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mihov S, Schulz KU (2007) Efficient dictionary-based text rewriting using subsequential transducers. Natural Language Engineering 13(4):353–381
2007
Cited alongside, same era.
Baron A, Rayson P (2008) VARD 2: A tool for dealing with spelling variation in historical corpora. In: Proceedings of the Postgraduate Conference in Corpus Linguistics, Birmingham, UK
2008
Cited alongside, same era.
World Wide Web Consortium (2008) Extensible markup language (XML) 1.0 (fifth edition). Online at http://www.w3.org/TR/2008/REC-xml-20081126
2008
Cited alongside, same era.
Depuydt K, de Does J (2009) Fons Verborum. Feestbundel voor prof. dr. A.M.F.J. (Fons) Moerdijk, aangeboden door vrienden en collega’s bij zijn afscheid van het INL, Instituut voor Nederlandse Lexicologie, Leiden/Amsterdam, chap Computational Tools and Lexica to Improve Access to Text., pp 187–199
2009
Cited alongside, same era.
Kocjančič P (2009) Internet y los recursos lingüísticos para la lengua española: diccionarios y corpus. Verba hispanica: anuario del Departamento de la Lengua y Literatura Españolas de la Facultad de Filosofía y Letras de la Universidad de Ljubljana 17:145–164
2009
Cited alongside, same era.
Montgomery DC (2009) Introduction To Statistical Quality Control. John Wiley & Sons
2009
Cited alongside, same era.
Sánchez Marco C, Boleda G, Fontana JM (2009) Propuesta de codificación de la información paleográfica y lingüística para textos diacrónicos del español. uso del estándar TEI. In: Proceedings of the Congreso Internacional Tradición e innovación: Nuevas perspectivas para la edición y el estudio de documentos antiguos, Madrid, Spain
2009
Cited alongside, same era.
Koehn P (2010) Statistical Machine Translation. Cambridge University Press
2010
Cited alongside, same era.
2011
Later among the works it cites.
Neudecker C, Schlarb S, Dogan M, Missier P, Sufi S, Williams A, Wolstencroft K (2011) An experimental workflow development platform for historical document digitisation and analysis. In: Proceedings of the 2011 Workshop on Historical Document Imaging and Processing, Beijing, China, pp 161–168
2011
Later among the works it cites.
Sánchez-Marco C, Boleda G, Padró L (2011) Extending the tool, or how to annotate historical language varieties. In: Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities, Portland, OR, USA, pp 1–9
2011
Later among the works it cites.
de Does J, Depuydt K (2012) Lexicon-supported OCR of eighteenth century Dutch books: a case study. In: Proceedings of the 20th Document Recognition and Retrieval Conference, San Francisco, CA USA, (to appear)
2012
Later among the works it cites.
Erjavec T (2012) The goo300k corpus of historical Slovene. In: Proceedings of the Eight International Conference on Language Resources and Evaluation, European Language Resources Association (ELRA), Istanbul, Turkey
2012
Later among the works it cites.
Kenter T, Erjavec T, Dulmin MZ, Fiser D (2012) Lexicon construction and corpus annotation of historical language with the CoBaLT editor. In: Proceedings of the 6th Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities, Association for Computational Linguistics, Avignon, France, pp 1–6
2012
Later among the works it cites.
Real Academia Española (s.a.) Banco de datos CORDE, corpus diacrónico del español. Online at http://corpus.rae.es/cordenet.html
2012
Later among the works it cites.
Sánchez-Prieto Borja P (2012) Desarrollo y explotación de un corpus de documentos españoles anteriores a 1700 (CODEA). Scriptum Digital 1:5–35
2012
Later among the works it cites.