Fetching the paper…
Reading the bibliography…
Advances in language modeling have led to the development of deep attention-based models that are performant across a wide variety of natural language processing (NLP) problems.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Earlier work this paper cites.
Visual graph comparison
K. Andrews, M. Wohlfahrt, and G. Wurzinger · 2009
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Weighted graph comparison techniques for brain connectivity analysis
B. Alper, B. Bach, N. Henry Riche, T. Isenberg, and J.-D. Fekete · 2013
Earlier work this paper cites.
Visual summaries for graph collections
D. Koop, J. Freire, and C. T. Silva · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder–decoder for statistical machine translation
K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Character-aware neural language models
Y. Kim, Y. Jernite, D. Sontag, and A. M. Rush · 2016
Earlier work this paper cites.
Towards better analysis of deep convolutional neural networks
M. Liu, J. Shi, Z. Li, C. Li, J. Zhu, and S. Liu · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Y. Wu, M. Schuster, Z. Chen, Q. V. Le, M. Norouzi, W. Macherey, M. Krikun, Y. Cao, Q. Gao, K. Macherey, et al · 2016
Earlier work this paper cites.
Considerations for visualizing comparison
M. Gleicher · 2017
Earlier work this paper cites.
Deep semantic role labeling: What works and what’s next
L. He, K. Lee, M. Lewis, and L. Zettlemoyer · 2017
Earlier work this paper cites.
Understanding hidden memories of recurrent neural networks
Y. Ming, S. Cao, R. Zhang, Z. Li, Y. Chen, Y. Song, and H. Qu · 2017
Cited alongside, same era.
Semi-supervised sequence tagging with bidirectional language models
M. Peters, W. Ammar, C. Bhagavatula, and R. Power · 2017
Cited alongside, same era.
Deepeyes: Progressive visual analytics for designing deep neural networks
N. Pezzotti, T. Höllt, J. Van Gemert, B. P. Lelieveldt, E. Eisemann, and A. Vilanova · 2017
Cited alongside, same era.
Lstmvis: A tool for visual analysis of hidden state dynamics in recurrent neural networks
H. Strobelt, S. Gehrmann, H. Pfister, and A. M. Rush · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Cnncomparator: Comparative analytics of convolutional neural networks
Summit: Scaling deep learning interpretability by visualizing activation and attribution summarizations
F. Hohman, H. Park, C. Robinson, and D. H. P. Chau · 2019
Later among the works it cites.
exbert: A visual analysis tool to explore learned representations in transformers models
B. Hoover, H. Strobelt, and S. Gehrmann · 2019
Later among the works it cites.
Revealing the dark secrets of bert
O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky · 2019
Later among the works it cites.
Sensebert: Driving some sense into bert
Y. Levine, B. Lenz, O. Dagan, D. Padnos, O. Sharir, S. Shalev-Shwartz, A. Shashua, and Y. Shoham · 2019
Later among the works it cites.
Open sesame: Getting inside bert’s linguistic knowledge
Y. Lina, Y. C. Tana, and R. Frankb · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Zeng, H. Haleem, X. Plantaz, N. Cao, and H. Qu · 2017
Cited alongside, same era.
Visual analytics in deep learning: An interrogative survey for the next frontiers
F. Hohman, M. Kahng, R. Pienta, and D. H. Chau · 2018
Cited alongside, same era.
Nlize: A perturbation-driven visual interrogation tool for analyzing and interpreting natural language inference models
S. Liu, Z. Li, T. Li, V. Srikumar, V. Pascucci, and P.-T. Bremer · 2018
Cited alongside, same era.
Deep contextualized word representations
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Cited alongside, same era.
Seq2seq-vis: A visual debugging tool for sequence-to-sequence models
H. Strobelt, S. Gehrmann, M. Behrisch, A. Perer, H. Pfister, and A. M. Rush · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman · 2018
Cited alongside, same era.
Rnnbow: Visualizing learning via backpropagation gradients in recurrent neural networks
D. Cashman, G. Patterson, A. Mosca, N. Watts, S. Robinson, and R. Chang · 2019
Cited alongside, same era.
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Later among the works it cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
T. McCoy, E. Pavlick, and T. Linzen · 2019
Later among the works it cites.
Deepcompare: Visual and interactive comparison of deep learning model performance
S. Murugesan, S. Malik, F. Du, E. Koh, and T. M. Lai · 2019
Later among the works it cites.
Sanvis: Visual analytics for understanding self-attention networks
C. Park, I. Na, Y. Jo, S. Shin, J. Yoo, B. C. Kwon, J. Zhao, H. Noh, Y. Lee, and J. Choo · 2019
Later among the works it cites.
Visualizing and measuring the geometry of bert
E. Reif, A. Yuan, M. Wattenberg, F. B. Viegas, A. Coenen, A. Pearce, and B. Kim · 2019
Later among the works it cites.
Is attention interpretable?
S. Serrano and N. A. Smith · 2019
Later among the works it cites.
How does bert answer questions? a layer-wise analysis of transformer representations
B. van Aken, B. Winter, A. Löser, and F. A. Gers · 2019
Later among the works it cites.
A multiscale visualization of attention in the transformer model
J. Vig · 2019
Later among the works it cites.
Ernie: Enhanced language representation with informative entities
Z. Zhang, X. Han, Z. Liu, X. Jiang, M. Sun, and Q. Liu · 2019
Later among the works it cites.
On identifiability in transformers
G. Brunner, Y. Liu, D. Pascual, O. Richter, M. Ciaramita, and R. Wattenhofer · 2020
Closest in time.
Spanbert: Improving pre-training by representing and predicting spans
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy · 2020
Closest in time.
Questioning the ai: Informing design practices for explainable ai user experiences
Q. V. Liao, D. Gruen, and S. Miller · 2020
Closest in time.
A primer in bertology: What we know about how bert works
A. Rogers, O. Kovaleva, and A. Rumshisky · 2020
Closest in time.