Fetching the paper…
Reading the bibliography…
Transformer-based language models (LMs) pretrained on large text collections are proven to store a wealth of semantic knowledge.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019b · 1907
Earlier work this paper cites.
Distilbert, a distilled version of BERT: Smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Silhouettes: A graphical aid to the interpretation and validation of cluster analysis
Peter J. Rousseeuw. 1987 · 1987
Earlier work this paper cites.
The ATIS Spoken Language Systems Pilot Corpus
Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990 · 1990
Earlier work this paper cites.
Efficient pattern recognition using a new transformation distance
Patrice Y. Simard, Yann LeCun, and John S. Denker. 1992 · 1992
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006 · 2006
Earlier work this paper cites.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2020 · 2007
Earlier work this paper cites.
DialoGLUE: A natural language understanding benchmark for task-oriented dialogue
Shikib Mehri, Mihail Eric, and Dilek Hakkani-Tür. 2020 · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Visualizing non-metric similarities in multiple maps
Laurens van der Maaten and Geoffrey E. Hinton. 2012 · 2012
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Conversational contextual cues: The case of personalization and history for response ranking
Rami Al-Rfou, Marc Pickett, Javier Snaider, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2016 · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Tim Lillicrap, Koray Kavukcuoglu, and Daan Wierstra. 2016 · 2016
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
In defense of the triplet loss for person re-identification
Alexander Hermans, Lucas Beyer, and Bastian Leibe. 2017 · 2017
Earlier work this paper cites.
Semantic specialisation of distributional word vector spaces using monolingual and cross-lingual constraints
Nikola Mrkšić, Ivan Vulić, Diarmuid Ó Séaghdha, Ira Leviant, Roi Reichart, Milica Gašić, Anna Korhonen, and Steve Young. 2017 · 2017
Earlier work this paper cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard S. Zemel. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Universal sentence encoder for English
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Earlier work this paper cites.
Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al. 2018 · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Earlier work this paper cites.
Learning to compare: Relation network for few-shot learning
Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip H. S. Torr, and Timothy M. Hospedales. 2018 · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Earlier work this paper cites.
Interpreting neural networks with nearest neighbors
Eric Wallace, Shi Feng, and Jordan Boyd-Graber. 2018 · 2018
Earlier work this paper cites.
ParaNMT-50M: Pushing the limits of paraphrastic sentence embeddings with millions of machine translations
John Wieting and Kevin Gimpel. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Learning semantic textual similarity from conversations
Yinfei Yang, Steve Yuan, Daniel Cer, Sheng-Yi Kong, Noah Constant, Petr Pilar, Heming Ge, Yun-hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Cited alongside, same era.
Learning cross-lingual sentence representations via a multi-task dual-encoder model
Muthuraman Chidambaram, Yinfei Yang, Daniel Cer, Steve Yuan, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2019 · 2019
Cited alongside, same era.
Cross-lingual language model pretraining
Alexis Conneau and Guillaume Lample. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Training neural response selection for task-oriented dialogue systems
Matthew Henderson, Ivan Vulić, Daniela Gerz, Iñigo Casanueva, Paweł Budzianowski, Sam Coope, Georgios Spithourakis, Tsung-Hsien Wen, Nikola Mrkšić, and Pei-Hao Su. 2019b · 2019
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Contrastive representation distillation
Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020 · 2020
Later among the works it cites.
Probing pretrained language models for lexical semantics
Ivan Vulić, Edoardo Maria Ponti, Robert Litschko, Goran Glavaš, and Anna Korhonen. 2020 · 2020
Later among the works it cites.
A bilingual generative transformer for semantic sentence embedding
John Wieting, Graham Neubig, and Taylor Berg-Kirkpatrick. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An evaluation dataset for intent classification and out-of-scope prediction
Stefan Larson, Anish Mahendran, Joseph J. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K. Kummerfeld, Kevin Leach, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. 2019 · 2019
Cited alongside, same era.
Benchmarking natural language understanding services for building conversational agents
Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser. 2019a · 2019
Cited alongside, same era.
Pretraining methods for dialog context representation learning
Shikib Mehri, Evgeniia Razumovskaia, Tiancheng Zhao, and Maxine Eskenazi. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Improving collaborative metric learning with efficient negative sampling
Viet-Anh Tran, Romain Hennequin, Jimena Royo-Letelier, and Manuel Moussallam. 2019 · 2019
Cited alongside, same era.
SuperGLUE: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a · 2019
Cited alongside, same era.
Multi-similarity loss with general pair weighting for deep metric learning
Xun Wang, Xintong Han, Weilin Huang, Dengke Dong, and Matthew R. Scott. 2019b · 2019
Cited alongside, same era.
Later among the works it cites.
TOD-BERT: Pre-trained natural language understanding for task-oriented dialogue
Chien-Sheng Wu, Steven C.H. Hoi, Richard Socher, and Caiming Xiong. 2020 · 2020
Later among the works it cites.
End-to-end slot alignment and recognition for cross-lingual NLU
Weijia Xu, Batool Haider, and Saab Mansour. 2020 · 2020
Later among the works it cites.
Multilingual universal sentence encoder for semantic retrieval
Yinfei Yang, Daniel Cer, Amin Ahmad, Mandy Guo, Jax Law, Noah Constant, Gustavo Hernandez Abrego, Steve Yuan, Chris Tar, Yun-hsuan Sung, et al. 2020 · 2020
Later among the works it cites.
Discriminative nearest neighbor few-shot intent detection by transferring natural language inference
Jianguo Zhang, Kazuma Hashimoto, Wenhao Liu, Chien-Sheng Wu, Yao Wan, Philip Yu, Richard Socher, and Caiming Xiong. 2020 · 2020
Later among the works it cites.
ProtAugment: Intent detection meta-learning through unsupervised diverse paraphrasing
Thomas Dopierre, Christophe Gravier, and Wilfried Logerais. 2021 · 2021
Closest in time.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021a · 2021
Closest in time.
SimCSE: Simple contrastive learning of sentence embeddings
Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021b · 2021
Closest in time.
Multilingual and cross-lingual intent detection from spoken data
Daniela Gerz, Pei-Hao Su, Razvan Kusztos, Avishek Mondal, Michal Lis, Eshan Singhal, Nikola Mrkšić, Tsung-Hsien Wen, and Ivan Vulić. 2021 · 2021
Closest in time.
Supervised contrastive learning for pre-trained language model fine-tuning
Beliz Gunel, Jingfei Du, Alexis Conneau, and Ves Stoyanov. 2021 · 2021
Closest in time.
ConVEx: Data-efficient and few-shot slot labeling
Matthew Henderson and Ivan Vulić. 2021 · 2021
Closest in time.
Neural data augmentation via example extrapolation
Kenton Lee, Kelvin Guu, Luheng He, Tim Dozat, and Hyung Won Chung. 2021 · 2021
Closest in time.
Evaluating multilingual text encoders for unsupervised cross-lingual retrieval
Robert Litschko, Ivan Vulić, Simone Paolo Ponzetto, and Goran Glavaš. 2021 · 2021
Closest in time.
Self-alignment pre-training for biomedical entity representations
Fangyu Liu, Ehsan Shareghi, Zaiqiao Meng, Marco Basaldella, and Nigel Collier. 2021a · 2021
Closest in time.
Fangyu Liu, Ivan Vulić, Anna Korhonen, and Nigel Collier. 2021b · 2021
Closest in time.
Example-driven intent prediction with observers
Shikib Mehri, Mihail Eric, and Dilek Hakkani-Tür. 2021 · 2021
Closest in time.
Semantically-conditioned negative samples for efficient contrastive learning
James O’Neill and Danushka Bollegala. 2021 · 2021
Closest in time.
Contrastive learning with hard negative samples
Joshua Robinson, Ching-Yao Chuang, Suvrit Sra, and Stefanie Jegelka. 2021 · 2021
Closest in time.
Recent advances in language model fine-tuning
Sebastian Ruder. 2021 · 2021
Closest in time.
A neighbourhood framework for resource-lean content flagging
Sheikh Muhammad Sarwar, Dimitrina Zlatkova, Momchil Hardalov, Yoan Dinkov, Isabelle Augenstein, and Preslav Nakov. 2021 · 2021
Closest in time.
Generating datasets with pretrained language models
Timo Schick and Hinrich Schütze. 2021 · 2021
Closest in time.
LexFit: Lexical fine-tuning of pretrained language models
Ivan Vulić, Edoardo Maria Ponti, Anna Korhonen, and Goran Glavaš. 2021 · 2021
Closest in time.
A closer look at few-shot crosslingual transfer: The choice of shots matters
Mengjie Zhao, Yi Zhu, Ehsan Shareghi, Ivan Vulić, Roi Reichart, Anna Korhonen, and Hinrich Schütze. 2021 · 2021
Closest in time.