Fetching the paper…
Reading the bibliography…
To build intelligent machine learning systems, there are two broad approaches.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 1908
Earlier work this paper cites.
Independent component analysis, a new concept?
P. Comon · 1994
Earlier work this paper cites.
Nonlinear independent component analysis: Existence and uniqueness results
A. Hyvärinen and P. Pajunen · 1999
Earlier work this paper cites.
Independent component analysis: algorithms and applications
A. Hyvärinen and E. Oja · 2000
Earlier work this paper cites.
Causation, prediction, and search
P. Spirtes, C. N. Glymour, and R. Scheines · 2000
Earlier work this paper cites.
Independent component analysis
A. Hyvarinen, J. Karhunen, and E. Oja · 2002
Earlier work this paper cites.
Lectures on Algebraic Statistics , volume 39 of Oberwolfach Seminars
M. Drton, B. Sturmfels, and S. Sullivant · 2009
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
T. Mikolov, W.-t. Yih, and G. Zweig · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Neural word embedding as implicit matrix factorization
O. Levy and Y. Goldberg · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
A. Radford, L. Metz, and S. Chintala · 2015
Earlier work this paper cites.
A latent variable model approach to pmi-based word embeddings
S. Arora, Y. Li, Y. Liang, T. Ma, and A. Risteski · 2016
Earlier work this paper cites.
Deep unsupervised clustering with gaussian mixture variational autoencoders
N. Dilokthanakul, P. A. Mediano, M. Garnelo, M. C. Lee, H. Salimbeni, K. Arulkumaran, and M. Shanahan · 2016
Earlier work this paper cites.
Unsupervised feature extraction by time-contrastive learning and nonlinear ica
A. Hyvarinen and H. Morioka · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations
D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba · 2017
Earlier work this paper cites.
Latent constraints: Learning to generate conditionally from unconditional generative models
J. Engel, M. Hoffman, and A. Roberts · 2017
Earlier work this paper cites.
Skip-gram- zipf+ uniform= vector additivity
A. Gittens, D. Achlioptas, and M. W. Mahoney · 2017
Earlier work this paper cites.
SGDR: stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Elements of causal inference: foundations and learning algorithms
J. Peters, D. Janzing, and B. Schölkopf · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
A. Radford, R. Jozefowicz, and I. Sutskever · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep understanding and improvement
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
A. Conneau, G. Kruszewski, G. Lample, L. Barrault, and M. Baroni · 2018
Earlier work this paper cites.
Towards understanding linear word analogies
K. Ethayarajh, D. Duvenaud, and G. Hirst · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, et al · 2018
Earlier work this paper cites.
Analogies explained: Towards understanding word embeddings
C. Allen and T. Hospedales · 2019
Earlier work this paper cites.
Machine learning interpretability: A survey on methods and metrics
D. V. Carvalho, E. M. Pereira, and J. S. Cardoso · 2019
Earlier work this paper cites.
Nonlinear ica using auxiliary variables and generalized contrastive learning
A. Hyvarinen, H. Sasaki, and R. Turner · 2019
Earlier work this paper cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
F. Locatello, S. Bauer, M. Lucic, G. Raetsch, S. Gelly, B. Schölkopf, and O. Bachem · 2019
Earlier work this paper cites.
Additive compositionality of word vectors
Y. Seonwoo, S. Park, D. Kim, and A. Oh · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
I. Tenney, D. Das, and E. Pavlick · 2019
Earlier work this paper cites.
Identifying through flows for recovering latent representations
S. Li, B. Hooi, and G. H. Lee · 2020
Earlier work this paper cites.
Disentanglement by nonlinear ica with general incompressible-flow networks (gin)
P. Sorrenson, C. Rother, and U. Köthe · 2020
Cited alongside, same era.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, et al · 2021
Cited alongside, same era.
Multi-facet clustering variational autoencoders
F. Falck, H. Zhang, M. Willetts, G. Nicholson, C. Yau, and C. C. Holmes · 2021
Cited alongside, same era.
Independent mechanism analysis, a new concept?
L. Gresele, J. Von Kügelgen, V. Stimper, B. Schölkopf, and M. Besserve · 2021
Cited alongside, same era.
Learning latent causal graphs via mixture oracles
B. Kivva, G. Rajendran, P. Ravikumar, and B. Aragam · 2021
Cited alongside, same era.
Causal structure learning: a combinatorial perspective
C. Squires and C. Uhler · 2022
Later among the works it cites.
Extracting latent steering vectors from pretrained language models
N. Subramani, N. Suresh, and M. E. Peters · 2022
Later among the works it cites.
Intervention target estimation in the presence of latent variables
B. Varici, K. Shanmugam, P. Sattigeri, and A. Tajer · 2022
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Lin, J. Hilton, and O. Evans · 2021
Cited alongside, same era.
Invariant causal representation learning for out-of-distribution generalization
C. Lu, Y. Wu, J. M. Hernández-Lobato, and B. Schölkopf · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Structure learning in polynomial time: Greedy algorithms, bregman information, and exponential families
G. Rajendran, B. Kivva, M. Gao, and B. Aragam · 2021
Cited alongside, same era.
Toward causal representation learning
B. Schölkopf, F. Locatello, S. Bauer, N. R. Ke, N. Kalchbrenner, A. Goyal, and Y. Bengio · 2021
Cited alongside, same era.
Retrieval augmentation reduces hallucination in conversation
K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston · 2021
Cited alongside, same era.
On the identifiability of nonlinear ica: Sparsity and beyond
Y. Zheng, I. Ng, and K. Zhang · 2022
Later among the works it cites.
Learning linear causal representations from interventions under general nonlinear mixing
S. Buchholz, G. Rajendran, E. Rosenfeld, B. Aragam, B. Schölkopf, and P. Ravikumar · 2023
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing · 2023
Later among the works it cites.
Context is environment
S. Gupta, S. Jegelka, D. Lopez-Paz, and K. Ahuja · 2023
Later among the works it cites.
Finding neurons in a haystack: Case studies with sparse probing
W. Gurnee, N. Nanda, M. Pauly, K. Harvey, D. Troitskii, and D. Bertsimas · 2023
Later among the works it cites.
Measuring and manipulating knowledge representations in language models
E. Hernandez, B. Z. Li, and J. Andreas · 2023
Later among the works it cites.
Identifiability of latent-variable and structural-equation models: from linear to nonlinear
A. Hyvärinen, I. Khemakhem, and R. Monti · 2023
Later among the works it cites.
Learning latent causal graphs with unknown interventions
Y. Jiang and B. Aragam · 2023
Later among the works it cites.
Uncovering meanings of embeddings via partial orthogonality
Y. Jiang, B. Aragam, and V. Veitch · 2023
Later among the works it cites.
Partial identifiability for domain adaptation
L. Kong, S. Xie, W. Yao, Y. Zheng, G. Chen, P. Stojanov, V. Akinwande, and K. Zhang · 2023
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
K. Li, O. Patel, F. Viégas, H. Pfister, and M. Wattenberg · 2023
Later among the works it cites.
Biscuit: Causal representation learning from binary interactions
P. Lippe, S. Magliacane, S. Löwe, Y. M. Asano, T. Cohen, and E. Gavves · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
N. Nanda, A. Lee, and M. Wattenberg · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
K. Park, Y. J. Choe, and V. Veitch · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
G. Rajendran, P. Reizinger, W. Brendel, and P. Ravikumar · 2023
Later among the works it cites.
Steering llama 2 via contrastive activation addition
N. Rimsky, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner · 2023
Later among the works it cites.
Bridging the human-ai knowledge gap: Concept discovery and transfer in alphazero
L. Schut, N. Tomasev, T. McGrath, D. Hassabis, U. Paquet, and B. Kim · 2023
Later among the works it cites.
Linear causal disentanglement via interventions
C. Squires, A. Seigal, S. S. Bhate, and C. Uhler · 2023
Later among the works it cites.
Towards the reusability and compositionality of causal representations
D. Talon, P. Lippe, S. James, A. Del Bue, and S. Magliacane · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
C. Tigges, O. J. Hollinsworth, A. Geiger, and N. Nanda · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Later among the works it cites.
Linear spaces of meanings: compositional structures in vision-language models
M. Trager, P. Perera, L. Zancato, A. Achille, P. Bhatia, and S. Soatto · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
A. Turner, L. Thiergart, D. Udell, G. Leech, U. Mini, and M. MacDiarmid · 2023
Later among the works it cites.
Score-based causal representation learning with interventions
B. Varici, E. Acarturk, K. Shanmugam, A. Kumar, and A. Tajer · 2023
Later among the works it cites.
Nonparametric identifiability of causal representations from unknown interventions
J. von Kügelgen, M. Besserve, W. Liang, L. Gresele, A. Kekić, E. Bareinboim, D. M. Blei, and B. Schölkopf · 2023
Later among the works it cites.
Concept algebra for score-based conditional model
Z. Wang, L. Gui, J. Negrea, and V. Veitch · 2023
Later among the works it cites.
Towards best practices of activation patching in language models: Metrics and methods
F. Zhang and N. Nanda · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.-K. Dombrowski, et al · 2023
Later among the works it cites.
On the origins of linear representations in large language models
Y. Jiang, G. Rajendran, P. Ravikumar, B. Aragam, and V. Veitch · 2024
Closest in time.