Fetching the paper…
Reading the bibliography…
We investigate the optimization target of Contrast-Consistent Search (CCS), which aims to recover the internal representations of truth of a large language model.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Hidden factors and hidden topics: Understanding rating dimensions with review text
Julian McAuley and Jure Leskovec · 2013
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions, May 2019
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
arxiv.org/pdf/1902.03129.pdf, 2019
Amirata Ghorbani, James Wexler, James Zou, and Been Kim · 2019
Earlier work this paper cites.
Thread: Circuits, mar 2020
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, and Ludwig Schubert · 2020
Cited alongside, same era.
Unifiedqa: Crossing format boundaries with a single qa system, 2020
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi · 2020
Cited alongside, same era.
SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems, February 2020
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2020
Cited alongside, same era.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Sid Black, Gao Leo, Phil Wang, Connor Leahy, and Stella Biderman · 2021
Cited alongside, same era.
Deberta: Decoding-enhanced bert with disentangled attention, 2021
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2021
Cited alongside, same era.
Discovering Latent Knowledge in Language Models Without Supervision, December 2022
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Later among the works it cites.
Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small, 2022
Kevin Wang, Variengien Re, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2022
Later among the works it cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung · 2023
Closest in time.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda B. Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…