Fetching the paper…
Reading the bibliography…
We consider the problem of evaluating representations of data for use in solving a downstream task.
A mathematical theory of communication
Shannon, C · 1948
Earlier work this paper cites.
Modeling by shortest data description
Rissanen, J · 1978
Earlier work this paper cites.
A tutorial introduction to the minimum description length principle
Grünwald, P · 2004
Earlier work this paper cites.
Elements of information theory
Cover, T. M. and Thomas, J. A · 2006
Earlier work this paper cites.
Catching up faster by switching sooner: A predictive approach to adaptive stimation with an application to the aic-bic dilemma
Erven, T., Grünwald, P., and Rooij, S · 2012
Earlier work this paper cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Deep learning & convolutional networks
LeCun, Y · 2015
Earlier work this paper cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2016
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
Ettinger, A., Elgohary, A., and Resnik, P · 2016
Earlier work this paper cites.
Learning distributed representations of sentences from unlabelled data
Hill, F., Cho, K., and Korhonen, A · 2016
Earlier work this paper cites.
Does string-based neural MT learn source syntax?
Shi, X., Padhi, I., and Knight, K · 2016
Cited alongside, same era.
The description length of deep learning models
Blier, L. and Ollivier, Y · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
Probing the state of the art: A critical look at visual representation evaluation
Resnick, C., Zhan, Z., and Bruna, J · 2019
Later among the works it cites.
oLMpics–on what language model pre-training captures
Talmor, A., Elazar, Y., Goldberg, Y., and Berant, J · 2019
Later among the works it cites.
Learning and evaluating general linguistic intelligence
Yogatama, D., d’Autume, C. d. M., Connor, J., Kocisky, T., Chrzanowski, M., Kong, L., Lazaridou, A., Ling, W., Yu, L., Dyer, C., et al · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
van den Oord, A., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Zhang, K. and Bowman, S · 2018
Cited alongside, same era.
Learning representations by maximizing mutual information across views
Bachman, P., Hjelm, R. D., and Buchwalter, W · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
Hewitt, J. and Liang, P · 2019
Cited alongside, same era.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Cited alongside, same era.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCand lish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E · 2020
Closest in time.
Learning optimal representations with the decodable information bottleneck
Dubois, Y., Kiela, D., Schwab, D. J., and Vedantam, R · 2020
Closest in time.
Data-efficient image recognition with contrastive predictive coding
Hénaff, O. J., Srinivas, A., Fauw, J., Razavi, A., Doersch, C., Eslami, S., and Oord, A · 2020
Closest in time.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Closest in time.
Formal limitations on the measurement of mutual information
McAllester, D. and Stratos, K · 2020
Closest in time.
Information-theoretic probing for linguistic structure
Pimentel, T., Valvoda, J., Maudslay, R. H., Zmigrod, R., Williams, A., and Cotterell, R · 2020
Closest in time.
Information-theoretic probing with minimum description length
Voita, E. and Titov, I · 2020
Closest in time.
A theory of usable information under computational constraints
Xu, Y., Zhao, S., Song, J., Stewart, R., and Ermon, S · 2020
Closest in time.