2019

An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction

Larson, Stefan, Mahendran, Anish, Peper, Joseph J. et al.

Understand

Task-oriented dialog systems need to know when a query falls outside their range of supported intents, but current text classification corpora only define label sets that cover every example.

  • We introduce a new dataset that includes queries that are out-of-scope---i.e., queries that do not fall into any of the system's supported intents.
  • This poses a new challenge because models cannot assume that every query at inference time belongs to a system-supported intent class.
  • Our dataset also covers 150 intent classes over 10 domains, capturing the breadth that a production task-oriented agent must handle.

Built on

  • Are We Modeling the Task or the Annotator? An Investigation of Annotator Bias in Natural Language Understanding Datasets

    Original

    Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 1908

    Earlier work this paper cites.

  • Learning question classifiers

    Xin Li and Dan Roth. 2002 · 2002

    Earlier work this paper cites.

  • The dialog state tracking challenge

    Jason Williams, Antoine Raux, Deepak Ramachandran, and Alan Black. 2013 · 2013

    Earlier work this paper cites.

  • Glove: Global vectors for word representation

    Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014

    Earlier work this paper cites.

  • Evaluating natural language understanding services for conversational question answering systems

    Daniel Braun, Adrian Hernandez-Mendez, Florian Matthes, and Manfred Langen. 2017 · 2017

    Earlier work this paper cites.

Similar

  • A baseline for detecting misclassified and out-of-distribution examples in neural networks

    Original

    Dan Hendrycks and Kevin Gimpel. 2017 · 2017

    Cited alongside, same era.

  • Bag of tricks for efficient text classification

    Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017 · 2017

    Cited alongside, same era.

  • Universal sentence encoder for English

    Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Brian Strope, and Ray Kurzweil. 2018 · 2018

    Cited alongside, same era.

  • Snips voice platform: an embedded spoken language understanding system for private-by-design voice interfaces

    Original

    Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, Maël Primet, and Joseph Dureau. 2018 · 2018

    Cited alongside, same era.

  • The second dialog state tracking challenge

    Matthew Henderson, Blaise Thomson, and Jason D. Williams. 2014a

    Cited in the paper.

  • The third dialog state tracking challenge

    Matthew Henderson, Blaise Thomson, and Jason D. Williams. 2014b

    Cited in the paper.

Then

  • Data collection for dialogue system: A startup perspective

    Yiping Kang, Yunqi Zhang, Jonathan K. Kummerfeld, Parker Hill, Johann Hauswald, Michael A. Laurenzano, Lingjia Tang, and Jason Mars. 2018 · 2018

    Later among the works it cites.

  • BERT: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019

    Closest in time.

  • Deep unknown intent detection with margin loss

    Ting-En Lin and Hua Xu. 2019 · 2019

    Closest in time.

  • Benchmarking natural language understanding services for building conversational agents

    Original

    Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser. 2019 · 2019

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…