Fetching the paper…
Reading the bibliography…
Why do large pre-trained transformer-based models perform so well across a wide variety of NLP tasks? Recent research suggests the key may lie in multi-headed attention mechanism's ability to learn and represent linguistic information.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Orion i: Interactive graphics
John Alan McDonald. 1988 · 1988
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
Escaping rgbland: Selecting colors for statistical graphics
Achim Zeileis, Kurt Hornik, and Paul Murrell. 2009 · 2009
Earlier work this paper cites.
Resolving object and attribute coreference in opinion mining
Xiaowen Ding and Bing Liu. 2010 · 2010
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
Christopher Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven Bethard, and David McClosky. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Reasoning about entailment with neural attention
Tim Rocktäschel, Edward Grefenstette, Karl Moritz Hermann, Tomás Kociský, and Phil Blunsom. 2016 · 2016
Earlier work this paper cites.
Visual analytics in deep learning: An interrogative survey for the next frontiers
Fred Hohman, Minsuk Kahng, Robert Pienta, and Duen Horng Chau. 2018 · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection
Leland McInnes, John Healy, Nathaniel Saul, and Lukas Grossberger. 2018 · 2018
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
Attention interpretability across NLP tasks
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019 · 2019
Cited alongside, same era.
A multiscale visualization of attention in the transformer model
On identifiability in transformers
Gino Brunner, Yang Liu, Damián Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Later among the works it cites.
exBERT: A Visual Analysis Tool to Explore Learned Representations in Transformer Models
Benjamin Hoover, Hendrik Strobelt, and Sebastian Gehrmann. 2020 · 2020
Later among the works it cites.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020 · 2020
Later among the works it cites.
Weight poisoning attacks on pretrained models
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2020
Later among the works it cites.
A primer in BERTology: What we know about how BERT works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jesse Vig. 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Attention is not not Explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Cited alongside, same era.
A diagnostic study of explainability techniques for text classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 · 2020
Cited alongside, same era.
Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai
Alejandro Barredo Arrieta, Natalia Díaz-Rodríguez, Javier Del Ser, Adrien Bennetot, Siham Tabik, Alberto Barbado, Salvador Garcia, Sergio Gil-Lopez, Daniel Molina, Richard Benjamins, Raja Chatila, and Francisco Herrera. 2020 · 2020
Cited alongside, same era.
The language interpretability tool: Extensible, interactive visualizations and analysis for NLP models
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, and Ann Yuan. 2020 · 2020
Later among the works it cites.
CNN Explainer: Learning Convolutional Neural Networks with Interactive Visualization
Zijie J. Wang, Robert Turko, Omar Shaikh, Haekyu Park, Nilaksh Das, Fred Hohman, Minsuk Kahng, and Duen Horng Chau. 2020 · 2020
Later among the works it cites.
Attention flows: Analyzing and comparing attention mechanisms in language models
J. F. DeRose, J. Wang, and M. Berger. 2021 · 2021
Closest in time.
Putting Humans in the Natural Language Processing Loop: A Survey
Zijie J. Wang, Dongjin Choi, Shenyu Xu, and Diyi Yang. 2021 · 2021
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.