Fetching the paper…
Reading the bibliography…
We present Ditto, a novel entity matching system based on pre-trained Transformer-based language models.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Machine learning
Tom M Mitchell et al · 1997
Earlier work this paper cites.
Learning to match and cluster large high-dimensional data sets for data integration. In Proc. KDD ’02 . 475–480
William W Cohen and Jacob Richman. 2002 · 2002
Earlier work this paper cites.
Interactive deduplication using active learning. In Proc. KDD ’02 . 269–278
Sunita Sarawagi and Anuradha Bhamidipaty. 2002 · 2002
Earlier work this paper cites.
A comparison of fast blocking methods for record
Linkage Rohan Baxter, Rohan Baxter, Peter Christen, et al · 2003
Earlier work this paper cites.
Adaptive duplicate detection using learnable string similarity measures. In Proc. KDD ’03 . 39–48
Mikhail Bilenko and Raymond J Mooney. 2003 · 2003
Earlier work this paper cites.
TextRank: Bringing order into text. In Proc. EMNLP ’04 . 404–411
Rada Mihalcea and Paul Tarau. 2004 · 2004
Earlier work this paper cites.
Evaluation of entity resolution approaches on real-world match problems
Hanna Köpcke, Andreas Thor, and Erhard Rahm. 2010 · 2010
Earlier work this paper cites.
A survey of indexing techniques for scalable record linkage and deduplication
Peter Christen. 2011 · 2011
Earlier work this paper cites.
Human-powered Sorts and Joins
Adam Marcus Eugene Wu David Karger and Samuel Madden Robert Miller. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al · 2011
Earlier work this paper cites.
Entity matching: How similar is similar
Jiannan Wang, Guoliang Li, Jeffrey Xu Yu, and Jianhua Feng. 2011 · 2011
Earlier work this paper cites.
CrowdER: crowdsourcing entity resolution
Jiannan Wang, Tim Kraska, Michael J Franklin, and Jianhua Feng. 2012 · 2012
Earlier work this paper cites.
Optimal hashing schemes for entity matching. In Proc. WWW ’13 . 295–306
Nilesh Dalvi, Vibhor Rastogi, Anirban Dasgupta, Anish Das Sarma, and Tamas Sarlos. 2013 · 2013
Earlier work this paper cites.
NADEEF/ER: generic and interactive entity resolution. In Proc. SIGMOD ’14 . 1071–1074
Ahmed Elmagarmid, Ihab F Ilyas, Mourad Ouzzani, Jorge-Arnulfo Quiané-Ruiz, Nan Tang, and Si Yin. 2014 · 2014
Earlier work this paper cites.
Corleone: Hands-off crowdsourcing for entity matching. In Proc. SIGMOD ’14 . 601–612
Chaitanya Gokhale, Sanjib Das, AnHai Doan, Jeffrey F Naughton, Narasimhan Rampalli, Jude Shavlik, and Xiaojin Zhu. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation. In Proc. EMNLP ’14 . 1532–1543
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
A clustering-based framework to control block sizes for entity resolution. In Proc. KDD ’15 . 279–288
Jeffrey Fisher, Peter Christen, Qing Wang, and Erhard Rahm. 2015 · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization. In Proc. EMNLP ’15
Alexander M Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Semantic-aware blocking for entity resolution
Qing Wang, Mingyuan Cui, and Huizhi Liang. 2015 · 2015
Earlier work this paper cites.
Magellan: Toward Building Entity Matching Management Systems
Pradap Konda, Sanjib Das, Paul Suganthan G. C., AnHai Doan, Adel Ardalan, Jeffrey R. Ballard, Han Li, Fatemah Panahi, Haojun Zhang, Jeffrey F. Naughton, Shishir Prasad, Ganesh Krishnan, Rohit Deep, and Vijay Raghavendra. 2016 · 2016
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
Active learning for large-scale entity resolution. In CIKM . 1379–1388
Kun Qian, Lucian Popa, and Prithviraj Sen. 2017 · 2017
Cited alongside, same era.
Synthesizing entity matching rules by examples
Rohit Singh, Venkata Vamsikrishna Meduri, Ahmed Elmagarmid, Samuel Madden, Paolo Papotti, Jorge-Arnulfo Quiané-Ruiz, Armando Solar-Lezama, and Nan Tang. 2017 · 2017
Cited alongside, same era.
Attention is all you need. In Proc. NIPS ’17 . 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
The WDC training dataset and gold standard for large-scale product matching. In Companion Proc. WWW ’19 . 381–386
Anna Primpeli, Ralph Peeters, and Christian Bizer. 2019 · 2019
Later among the works it cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019a · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019b · 2019
Later among the works it cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proc. EMNLP-IJCNLP ’19 . 3982–3992
Nils Reimers and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. In Proc. EMC 2 ’19
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Leveraging Knowledge Bases in LSTMs for Improving Machine Reading. In Proc. ACL ’17 . 1436–1446
Bishan Yang and Tom Mitchell. 2017 · 2017
Cited alongside, same era.
Neural natural language inference models enhanced with external knowledge. In Proc. ACL ’18 . 2406–2417
Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, Diana Inkpen, and Si Wei. 2018 · 2018
Cited alongside, same era.
Distributed representations of tuples for entity resolution
Muhammad Ebraheem, Saravanan Thirumuruganathan, Shafiq Joty, Mourad Ouzzani, and Nan Tang. 2018 · 2018
Cited alongside, same era.
Deep learning for entity matching: A design space exploration. In Proc. SIGMOD ’18 . 19–34
Sidharth Mudgal, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra. 2018 · 2018
Cited alongside, same era.
mixup: Beyond empirical risk minimization. In Proc. ICLR ’18
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz. 2018 · 2018
Cited alongside, same era.
To Index or Not to Index: Optimizing Exact Maximum Inner Product Search. In Proc. ICDE ’19 . IEEE, 1250–1261
Firas Abuzaid, Geet Sethi, Peter Bailis, and Matei Zaharia. 2019 · 2019
Cited alongside, same era.
SciBERT: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Cited alongside, same era.
Yu Sun, Shuohuan Wang, Yukun Li, Shikun Feng, Xuyi Chen, Han Zhang, Xin Tian, Danxiang Zhu, Hao Tian, and Hua Wu. 2019 · 2019
Later among the works it cites.
BERT Rediscovers the Classical NLP Pipeline. In Proc. ACL ’19 . 4593–4601
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
KGAT: Knowledge Graph Attention Network for Recommendation. In Proc. KDD ’19 . 950–958
Xiang Wang, Xiangnan He, Yixin Cao, Meng Liu, and Tat-Seng Chua. 2019 · 2019
Later among the works it cites.
EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks. In Proc. EMNLP-IJCNLP ’19 . 6382–6388
Jason Wei and Kai Zou. 2019 · 2019
Later among the works it cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Later among the works it cites.
Unsupervised data augmentation
Qizhe Xie, Zihang Dai, Eduard Hovy, Minh-Thang Luong, and Quoc V Le. 2019 · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding. In Proc. NeurIPS ’19 . 5754–5764
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Auto-EM: End-to-end Fuzzy Entity-Matching using Pre-trained Deep Models and Transfer Learning. In Proc. WWW ’19 . 2413–2424
Chen Zhao and Yeye He. 2019 · 2019
Later among the works it cites.
Entity matching with transformer architectures-a step forward in data integration. In EDBT
Ursin Brunner and Kurt Stockinger. 2020 · 2020
Closest in time.
TAPAS: Weakly Supervised Table Parsing via Pre-training
Jonathan Herzig, Paweł Krzysztof Nowak, Thomas Müller, Francesco Piccinno, and Julian Martin Eisenschlos. 2020 · 2020
Closest in time.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Closest in time.
Deep Entity Matching with Pre-Trained Language Models
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan. 2020 · 2020
Closest in time.
A Comprehensive Benchmark Framework for Active Learning Methods in Entity Matching. In SIGMOD . 1133–1147
Venkata Vamsikrishna Meduri, Lucian Popa, Prithviraj Sen, and Mohamed Sarwat. 2020 · 2020
Closest in time.
Snippext: Semi-supervised Opinion Mining with Augmented Data. In Proc. WWW ’20
Zhengjie Miao, Yuliang Li, Xiaolan Wang, and Wang-Chiew Tan. 2020 · 2020
Closest in time.
TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
Pengcheng Yin, Graham Neubig, Wen-tau Yih, and Sebastian Riedel. 2020 · 2020
Closest in time.