Fetching the paper…
Reading the bibliography…
In many contexts, lying -- the use of verbal falsehoods to deceive -- is harmful.
Risks from Learned Optimization in Advanced Machine Learning Systems
Hubinger, E., C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant (2019, June) · 1906
Earlier work this paper cites.
Dota 2 with Large Scale Deep Reinforcement Learning
OpenAI, C. Berner, G. Brockman, B. Chan, V. Cheung, P. Dębiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. d. O. Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang (2019, December) · 1912
Earlier work this paper cites.
The market for “Lemons”: Quality uncertainty and the market mechanism
Akerlof, G. A. (1970, August) · 1970
Earlier work this paper cites.
The Strategy of Conflict: With a New Preface by the Author
Schelling, T. (1980) · 1980
Earlier work this paper cites.
Structural Inertia and Organizational Change
Hannan, M. T. and J. Freeman (1984) · 1984
Earlier work this paper cites.
Social Capital in the Creation of Human Capital
Coleman, J. S. (1988) · 1988
Earlier work this paper cites.
The Intentional Stance
Dennett, D. C. (1989) · 1989
Earlier work this paper cites.
A Transaction Cost Theory of Politics
North, D. C. (1990, October) · 1990
Earlier work this paper cites.
Path Dependence, Lock-In, and History
Liebowitz, S. J. and S. E. Margolis (1995, April) · 1995
Earlier work this paper cites.
Does social capital have an economic payoff? a cross-country investigation*
Knack, S. and P. Keefer (1997, November) · 1997
Earlier work this paper cites.
The New Chicago School
Lessig, L. (1998, June) · 1998
Earlier work this paper cites.
Bending the rules: Flexible regulation and constraints on agency discretion
Seidenfeld, M. (1999) · 1999
Earlier work this paper cites.
Towards a Human-like Open-Domain Chatbot
Adiwardana, D., M.-T. Luong, D. R. So, J. Hall, N. Fiedel, R. Thoppilan, Z. Yang, A. Kulshreshtha, G. Nemade, Y. Lu, and Q. V. Le (2020, February) · 2001
Earlier work this paper cites.
Regulatory Markets for AI Safety
Clark, J. and G. K. Hadfield (2019, December) · 2001
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J., S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei (2020, January) · 2001
Earlier work this paper cites.
Trust and Growth
Zak, P. J. and S. Knack (2001) · 2001
Earlier work this paper cites.
Economic Analysis of Corruption: A Survey
Aidt, T. S. (2003, November) · 2003
Earlier work this paper cites.
Using TF-IDF to determine word relevance in document queries
Ramos, J. E. (2003) · 2003
Earlier work this paper cites.
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
Brundage, M., S. Avin, J. Wang, H. Belfield, G. Krueger, G. Hadfield, H. Khlaaf, J. Yang, H. Toner, R. Fong, T. Maharaj, P. W. Koh, S. Hooker, J. Leung, A. Trask, E. Bluemke, J. Lebensold, C. O’Keefe, M. Koren, T. Ryffel, J. B. Rubinovitz, T. Besiroglu, F. Carugati, J. Clark, P. Eckersley, S. de Haas, M. Johnson, B. Laurie, A. Ingerman, I. Krawczuk, A. Askell, R. Cammarota, A. Lohn, D. Krueger, C. Stix, P. Henderson, L. Graham, C. Prunkl, B. Martin, E. Seger, N. Zilberman, S. Ó. hÉigeartaigh, F. Kroeger, G. Sastry, R. Kagan, A. Weller, B. Tse, E. Barnes, A. Dafoe, P. Scharre, A. Herbert-Voss, M. Rasser, S. Sodhani, C. Flynn, T. K. Gilbert, L. Dyer, S. Khan, Y. Bengio, and M. Anderljung (2020, April) · 2004
Earlier work this paper cites.
Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?
Jacovi, A. and Y. Goldberg (2020, April) · 2004
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B., B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei (2020, July) · 2005
Earlier work this paper cites.
The Economic Costs of Corruption: A Survey and New Evidence
Dreher, A. and T. Herzfeld (2005) · 2005
Earlier work this paper cites.
A Simple Language Model for Task-Oriented Dialogue
Hosseini-Asl, E., B. McCann, C.-S. Wu, S. Yavuz, and R. Socher (2020, July) · 2005
Earlier work this paper cites.
UnifiedQA: Crossing Format Boundaries With a Single QA System
Khashabi, D., S. Min, T. Khot, A. Sabharwal, O. Tafjord, P. Clark, and H. Hajishirzi (2020, October) · 2005
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Lewis, P., E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, and D. Kiela (2021, April) · 2005
Earlier work this paper cites.
AI Research Considerations for Human Existential Safety (ARCHES)
Critch, A. and D. Krueger (2020, May) · 2006
Earlier work this paper cites.
Emergent Multi-Agent Communication in the Deep Learning Era
Lazaridou, A. and M. Baroni (2020, July) · 2006
Earlier work this paper cites.
The Promise of Prediction Markets
Arrow, K. J., R. Forsythe, M. Gorham, R. Hahn, R. Hanson, J. O. Ledyard, S. Levmore, R. Litan, P. Milgrom, F. D. Nelson, G. R. Neumann, M. Ottaviani, T. C. Schelling, R. J. Shiller, V. L. Smith, E. Snowberg, C. R. Sunstein, P. C. Tetlock, P. E. Tetlock, H. R. Varian, J. Wolfers, and E. Zitzewitz (2008) · 2008
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D., C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt (2021, January) · 2009
Earlier work this paper cites.
Generative Language Modeling for Automated Theorem Proving
Polu, S. and I. Sutskever (2020, September) · 2009
Earlier work this paper cites.
Learning to summarize from human feedback
Stiennon, N., L. Ouyang, J. Wu, D. M. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. Christiano (2020, October) · 2009
Earlier work this paper cites.
Scaling Laws for Autoregressive Generative Modeling
Henighan, T., J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, C. Hallacy, B. Mann, A. Radford, A. Ramesh, N. Ryder, D. M. Ziegler, J. Schulman, D. Amodei, and S. McCandlish (2020, November) · 2010
Earlier work this paper cites.
Imitating Interactive Intelligence
Abramson, J., A. Ahuja, I. Barr, A. Brussee, F. Carnevale, M. Cassin, R. Chhaparia, S. Clark, B. Damoc, A. Dudzik, P. Georgiev, A. Guy, T. Harley, F. Hill, A. Hung, Z. Kenton, J. Landon, T. Lillicrap, K. Mathewson, S. Mokrá, A. Muldal, A. Santoro, N. Savinov, V. Varma, G. Wayne, D. Williams, N. Wong, C. Yan, and R. Zhu (2021, January) · 2012
Earlier work this paper cites.
Thinking Inside the Box: Controlling and Using an Oracle AI
Armstrong, S., A. Sandberg, and N. Bostrom (2012, November) · 2012
Earlier work this paper cites.
Probabilistic topic models
Blei, D. M. (2012, April) · 2012
Earlier work this paper cites.
Open Problems in Cooperative AI
Dafoe, A., E. Hughes, Y. Bachrach, T. Collins, K. R. McKee, J. Z. Leibo, K. Larson, and T. Graepel (2020, December) · 2012
Earlier work this paper cites.
A few useful things to know about machine learning
Domingos, P. (2012, October) · 2012
Earlier work this paper cites.
Imprinting: Toward a Multilevel Theory
Marquis, C. and A. Tilcsik (2013, June) · 2013
Earlier work this paper cites.
A taxonomy of robot deception and its benefits in HRI
Shim, J. and R. C. Arkin (2013) · 2013
Earlier work this paper cites.
Chapter 2 - Trust, Growth, and Well-Being: New Evidence and Policy Implications
Algan, Y. and P. Cahuc (2014, January) · 2014
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N. (2014) · 2014
Cited alongside, same era.
Ethical guidelines for a superintelligence
Davis, E. (2015, March) · 2014
Cited alongside, same era.
The deflationary theory of truth
Stoljar, D. and N. Damnjanovic (2014) · 2014
Cited alongside, same era.
VW Is Said to Cheat on Diesel Emissions; U.S. to Order Big Recall
Davenport, C. and J. Ewing (2015, September) · 2015
Cited alongside, same era.
Research priorities for robust and beneficial artificial intelligence
Russell, S., D. Dewey, and M. Tegmark (2015, December) · 2015
Cited alongside, same era.
Concrete Problems in AI Safety
Amodei, D., C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané (2016, July) · 2016
Cited alongside, same era.
Conversation with Paul Christiano - AI Impacts
Christiano, P., A. Bergal, R. Fernandez, and R. Long (2019, September) · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., M.-W. Chang, K. Lee, and K. Toutanova (2019, May) · 2019
Later among the works it cites.
Machine learning projects for iterated distillation and amplification
Evans, O., W. Saunders, and A. Stuhlmüller (2019) · 2019
Later among the works it cites.
Horizon: Facebook’s Open Source Applied Reinforcement Learning Platform
Gauci, J., E. Conti, Y. Liang, K. Virochsiri, Y. He, Z. Kaden, V. Narayanan, X. Ye, Z. Chen, and S. Fujimoto (2019, September) · 2019
Later among the works it cites.
The Financial Cost of Fraud 2019: The Latest Data from around the World
Gee, J. and M. Button (2019, July) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The Age of Em: Work, Love, and Life When Robots Rule the Earth
Hanson, R. (2016, May) · 2016
Cited alongside, same era.
Honesty, beliefs about honesty, and economic growth in 15 countries
Hugh-Jones, D. (2016, July) · 2016
Cited alongside, same era.
Future of AI 6. Discussion of ’Superintelligence: Paths, Dangers, Strategies’ - inverseprobability.com: Neil Lawrence’s Homepage
Lawrence, N. (2016, May) · 2016
Cited alongside, same era.
Deep Reinforcement Learning for Dialogue Generation
Li, J., W. Monroe, A. Ritter, M. Galley, J. Gao, and D. Jurafsky (2016, September) · 2016
Cited alongside, same era.
The definition of lying and deception
Mahon, J. E. (2016) · 2016
Cited alongside, same era.
Value alignment or misalignment - what will keep systems accountable?
Arnold, T., D. Kasenberg, and M. Scheutz (2017) · 2017
Cited alongside, same era.
Artificial intelligence and its implications for income distribution and unemployment
Korinek, A. and J. E. Stiglitz (2019) · 2019
Later among the works it cites.
Towards Deep Learning Models Resistant to Adversarial Attacks
Madry, A., A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu (2019, September) · 2019
Later among the works it cites.
Categorizing Variants of Goodhart’s Law
Manheim, D. and S. Garrabrant (2019, February) · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever (2019) · 2019
Later among the works it cites.
Technology and Policymakers - Schneier on Security
Schneier, B. (2019, November) · 2019
Later among the works it cites.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Vinyals, O., I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, C. Gulcehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wünsch, K. McKinney, O. Smith, T. Schaul, T. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver (2019, November) · 2019
Later among the works it cites.
Sanity Checks for Saliency Maps
Adebayo, J., J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim (2020, November) · 2020
Later among the works it cites.
The Alignment Problem: Machine Learning and Human Values
Christian, B. (2020) · 2020
Later among the works it cites.
Inaccessible information - AI Alignment
Christiano, P. (2020, June) · 2020
Later among the works it cites.
The correspondence theory of truth
David, M. (2020) · 2020
Later among the works it cites.
The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Gao, L., S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy (2020, December) · 2020
Later among the works it cites.
Zoom In: An Introduction to Circuits
Olah, C., N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter (2020, March) · 2020
Later among the works it cites.
It takes two to lie: One to lie and one to listen
Peskov, D., B. Cheng, A. Elgohary, J. Barrow, C. Danescu-Niculescu-Mizil, and J. Boyd-Graber (2020) · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu (2020) · 2020
Later among the works it cites.
Evaluating arguments one step at a time
Saunders, W., B. Rachbach, O. Evans, Z. Miller, J. Byun, and A. Stuhlmüller (2020) · 2020
Later among the works it cites.
Tackling threats to informed decision-making in democratic societies: Promoting epistemic security in a technologically-advanced world
Seger, E., S. Avin, G. Pearson, M. Briers, S. Ó Heigeartaigh, and H. Bacon (2020, October) · 2020
Later among the works it cites.
A Deep Learning Approach to Antibiotic Discovery
Stokes, J. M., K. Yang, K. Swanson, W. Jin, A. Cubillos-Ruiz, N. M. Donghia, C. R. MacNair, S. French, L. A. Carfrae, Z. Bloom-Ackermann, V. M. Tran, A. Chiappino-Pepe, A. H. Badran, I. W. Andrews, E. J. Chory, G. M. Church, E. D. Brown, T. S. Jaakkola, R. Barzilay, and J. J. Collins (2020, February) · 2020
Later among the works it cites.
Economic growth under transformative AI: A guide to the vast range of possibilities for output growth, wages, and the labor share
Trammell, P. and A. Korinek (2020, October) · 2020
Later among the works it cites.
FEVEROUS: Fact Extraction and VERification Over Unstructured and Structured information
Aly, R., Z. Guo, M. Schlichtkrull, J. Thorne, A. Vlachos, C. Christodoulopoulos, O. Cocarascu, and A. Mittal (2021, September) · 2021
Closest in time.
QNRs: Toward language for intelligent machines
Drexler, K. E. (2021) · 2021
Closest in time.
Proposal for a Regulation of the European Parliament and of the Council, Laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) and amending certain Union Legislative Acts
European Commission, Directorate - General for Communications Networks, C. and Technology (2021, April) · 2021
Closest in time.
Agent incentives: A causal perspective
Everitt, T., R. Carey, E. D. Langlois, P. A. Ortega, and S. Legg (2021) · 2021
Closest in time.
Highly accurate protein structure prediction with AlphaFold
Jumper, J., R. Evans, A. Pritzel, T. Green, M. Figurnov, O. Ronneberger, K. Tunyasuvunakool, R. Bates, A. Žídek, A. Potapenko, A. Bridgland, C. Meyer, S. A. A. Kohl, A. J. Ballard, A. Cowie, B. Romera-Paredes, S. Nikolov, R. Jain, J. Adler, T. Back, S. Petersen, D. Reiman, E. Clancy, M. Zielinski, M. Steinegger, M. Pacholska, T. Berghammer, S. Bodenstein, D. Silver, O. Vinyals, A. W. Senior, K. Kavukcuoglu, P. Kohli, and D. Hassabis (2021, August) · 2021
Closest in time.
Kenton, Z., T. Everitt, L. Weidinger, I. Gabriel, V. Mikulik, and G. Irving (2021, March) · 2021
Closest in time.
GENIE: A Leaderboard for Human-in-the-Loop Evaluation of Text Generation
Khashabi, D., G. Stanovsky, J. Bragg, N. Lourie, J. Kasai, Y. Choi, N. A. Smith, and D. S. Weld (2021, June) · 2021
Closest in time.
TruthfulQA: Measuring How Models Mimic Human Falsehoods
Lin, S., J. Hilton, and O. Evans (2021, September) · 2021
Closest in time.
A unifying review of deep and shallow anomaly detection
Ruff, L., J. R. Kauffmann, R. A. Vandermeulen, G. Montavon, W. Samek, M. Kloft, T. G. Dietterich, and K.-R. Müller (2021) · 2021
Closest in time.
Human-compatible artificial intelligence
Russell, S. (2021) · 2021
Closest in time.
Classical logic
Shapiro, S. and T. Kouri Kissel (2021) · 2021
Closest in time.
Retrieval Augmentation Reduces Hallucination in Conversation
Shuster, K., S. Poff, M. Chen, D. Kiela, and J. Weston (2021, April) · 2021
Closest in time.
Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
Solaiman, I. and C. Dennison (2021, June) · 2021
Closest in time.
Counterfactuals
Starr, W. (2021) · 2021
Closest in time.
CommonsenseQA 2.0: Exposing the limits of AI through gamification
Talmor, A., O. Yoran, R. L. Bras, C. Bhagavatula, Y. Goldberg, Y. Choi, and J. Berant (2021) · 2021
Closest in time.
GPT-J-6B: A 6 billion parameter autoregressive language model
Wang, B. and A. Komatsuzaki (2021) · 2021
Closest in time.
Finetuned Language Models Are Zero-Shot Learners
Wei, J., M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le (2021, September) · 2021
Closest in time.