Fetching the paper…
Reading the bibliography…
At face value, this essay is about understanding a fairly esoteric governance tool called compute thresholds.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 1905
Earlier work this paper cites.
Development Strategy and Planning: The Soviet Experience , pp. 233–278
Alexander Erlich · 1967
Earlier work this paper cites.
Industrial union department v. American Petroleum Institute, June 1980
448 U.S. 607 · 1980
Earlier work this paper cites.
What does it mean to understand language?
Terry Winograd · 1980
Earlier work this paper cites.
Optimal Brain Damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
What every computer scientist should know about floating-point arithmetic
David Goldberg · 1991
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory
James Mcclelland, Bruce Mcnaughton, and Randall O’Reilly · 1995
Earlier work this paper cites.
Moore’s law: past, present and future
R.R. Schaller · 1997
Earlier work this paper cites.
Sparse connection and pruning in large dynamic artificial neural networks, 1997
Nikko Ström · 1997
Earlier work this paper cites.
Zipf’s law and the internet
Lada Adamic and Bernardo Huberman · 2001
Earlier work this paper cites.
Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, Margaret Mitchell, Timnit Gebru, Ben Hutchinson, Jamila Smith-Loud, Daniel Theron, and Parker Barnes · 2001
Earlier work this paper cites.
The early phase of neural network training
Jonathan Frankle, David J. Schwab, and Ari S. Morcos · 2002
Earlier work this paper cites.
The Great Fire of London: In That Apocalyptic Year, 1666
John Peter · 2002
Earlier work this paper cites.
How much information?, 2003
Peter Lyman and Hal R Varian · 2003
Earlier work this paper cites.
The Black Death, 1346-1353: The Complete History
O.J. Benedictow · 2004
Earlier work this paper cites.
The weight of scientific evidence in policy and law
Sheldon Krimsky · 2004
Earlier work this paper cites.
GPU implementation of neural networks
Kyoung-Su Oh and Keechul Jung · 2004
Earlier work this paper cites.
Measuring the algorithmic efficiency of neural networks
Danny Hernandez and Tom B. Brown · 2005
Earlier work this paper cites.
Accelerated 2D Image Processing on GPUs
Bryson R. Payne, Saeid O. Belkasim, G. Scott Owen, Michael C. Weeks, and Ying Zhu · 2005
Earlier work this paper cites.
High performance convolutional neural networks for document processing, 10 2006
Kumar Chellapilla, Sidd Puri, and Patrice Simard · 2006
Earlier work this paper cites.
Power Consumption Variation over Activation Functions
Leon Derczynski · 2006
Earlier work this paper cites.
Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, Debsindhu Bhowmik, and Burkhard Rost · 2007
Earlier work this paper cites.
Historical trends in executive compensation, 1936-2003
Carola Frydman and Raven Molloy · 2007
Earlier work this paper cites.
A new look at screening and diagnosing diabetes mellitus
Christopher D. Saudek, William H. Herman, David B. Sacks, Richard M. Bergenstal, David Edelman, and Mayer B. Davidson · 2007
Earlier work this paper cites.
Limits of viability: definition of the gray zone
Irit Seri and Jonathan Evans · 2008
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models, 2020
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2009
Earlier work this paper cites.
Potential exposure to anti-drug advertising and drug-related attitudes, beliefs, and behaviors among united states youth, 1995-2006
Yvonne M. Terry-McElrath, Sherry Emery, Gery Szczypka, and Lloyd D. Johnston · 2010
Earlier work this paper cites.
Flexible, high performance convolutional neural networks for image classification., 07 2011
Dan Ciresan, Ueli Meier, Jonathan Masci, Luca Maria Gambardella, and Jürgen Schmidhuber · 2011
Earlier work this paper cites.
The growing gap between emerging technologies and the law
Gary E. Marchant · 2011
Earlier work this paper cites.
Benefits and limitations of the precautionary principle
P.F. Ricci and J. Zhang · 2011
Earlier work this paper cites.
Graphics processing unit (gpu) programming strategies and trends in gpu computing
André R. Brodtkorb, Trond R. Hagen, and Martin L. Sætra · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Building high-level features using large scale unsupervised learning, 2012
Quoc V. Le, Marc’Aurelio Ranzato, Rajat Monga, Matthieu Devin, Kai Chen, Greg S. Corrado, Jeff Dean, and Andrew Y. Ng · 2012
Earlier work this paper cites.
Monitoring systemic risk based on dynamic thresholds
Kasper Lund-Jensen · 2012
Earlier work this paper cites.
Critical truths about power laws
Michael P. H. Stumpf and Mason A. Porter · 2012
Earlier work this paper cites.
Deep learning with COTS HPC systems
Adam Coates, Brody Huval, Tao Wang, David Wu, Bryan Catanzaro, and Ng Andrew · 2013
Earlier work this paper cites.
Reference class forecasting: Resolving its challenge to statistical modeling
Robert F. Bordley · 2014
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2014
Earlier work this paper cites.
A pilot study on maternal oral health and birth weight of twins
Ye Shen, Chao Li, Aura Heimonen, Jukka Meurman, Martha Nunn, Donald Miller, Thomas Van Dyke, Prashanti Bollu, Risto Kaaja, and Sok-Ja Janket · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Going deeper with convolutions, 2014
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2014
Earlier work this paper cites.
Deep Learning with Limited Numerical Precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Hidden technical debt in machine learning systems
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Dennison · 2015
Earlier work this paper cites.
An Analysis of Deep Neural Network Models for Practical Applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Earlier work this paper cites.
Compression of Neural Machine Translation Models via Pruning
Abigail See, Minh-Thang Luong, and Christopher D. Manning · 2016
Earlier work this paper cites.
Learning Structured Sparsity in Deep Neural Networks
W. Wen, C. Wu, Y. Wang, Y. Chen, and H. Li · 2016
Earlier work this paper cites.
Critical learning periods in deep neural networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto · 2017
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S. Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, and Simon Lacoste-Julien · 2017
Earlier work this paper cites.
Low birth weight: Case definition & guidelines for data collection, analysis, and presentation of maternal immunization safety data
Claire L. Cutland, Emily M. Lackritz, Tanis Mallett-Moore, Anna Bardají, Ramakrishnan Chandrasekaran, Cagri Lahariya, Muhammad Imran Nisar, Martin D. Tapia, Jay Pathirana, Sumita Kochhar, Francisco M. Muñoz, and Brighton Collaboration Low Birth Weight Working Group · 2017
Earlier work this paper cites.
Learning Sparse Neural Networks through L 0 L_{0} Regularization
Christos Louizos, Max Welling, and Diederik P. Kingma · 2017
Earlier work this paper cites.
Determinants of day–night difference in blood pressure, a comparison with determinants of daytime and night-time blood pressure
Mohammad Musameh, Christopher Nelson, Jonathan Gracey, Jessica Davies, Richard Davies, Denise Francis, Adrian Hughes, Gregory Y H Lip, Helen Mcnamara, Alison Mccarthy, et al · 2017
Earlier work this paper cites.
The History of Microwave Heating
Hua Zhang · 2017
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers, 2018
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, and Blake Hechtman · 2018
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, et al · 2019
Earlier work this paper cites.
The difficulty of training sparse neural networks
Utku Evci, Fabian Pedregosa, Aidan Gomez, and Erich Elsen · 2019
Earlier work this paper cites.
What do compressed deep neural networks forget?, 2019
Sara Hooker, Aaron Courville, Gregory Clark, Yann Dauphin, and Andrea Frome · 2019
Earlier work this paper cites.
Do deep neural networks learn shallow learnable examples first, 2019
Karttikeya Mangalam and Vinay Uday Prabhu · 2019
Earlier work this paper cites.
The bitter lesson, 2019
Richard Sutton · 2019
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Estimating example difficulty using variance of gradients, 2020
Chirag Agarwal and Sara Hooker · 2020
Earlier work this paper cites.
Compressing Neural Machine Translation Models with 4-bit Precision
Alham Fikri Aji and Kenneth Heafield · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale, July 2019
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
Government by algorithm: Artificial intelligence in federal administrative agencies
David Freeman Engstrom, Daniel E. Ho, Catherine M. Sharkey, and Mariano-Florentino Cuellar · 2020
Earlier work this paper cites.
A Study of Gradient Variance in Deep Learning
Fartash Faghri, David Duvenaud, David J. Fleet, and Jimmy Ba · 2020
Earlier work this paper cites.
Characterising Bias in Compressed Models, 2020
Sara Hooker, Nyalleng Moorosi, Gregory Clark, Samy Bengio, and Emily Denton · 2020
Earlier work this paper cites.
Exploring the memorization-generalization continuum in deep learning
Ziheng Jiang, Chiyuan Zhang, Kunal Talwar, and Michael C Mozer · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding, 2020
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2020
Earlier work this paper cites.
Inter- and intraindividual variability in daily resting heart rate and its associations with age, sex, sleep, bmi, and time of year: Retrospective, longitudinal cohort study of 92,457 adults
Giorgio Quer, Pishoy Gouda, Michael Galarnyk, Eric Topol, and Steven Steinhubl · 2020
Earlier work this paper cites.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, 2020
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Green AI
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2020
Earlier work this paper cites.
The Computational Limits of Deep Learning
Neil C. Thompson, Kristjan Greenewald, Keeheon Lee, and Gabriel F. Manso · 2020
Earlier work this paper cites.
Chapter 6: Government roles and related considerations, 2020
US Department of Transportation · 2020
Earlier work this paper cites.
The staircase property: How hierarchical structure can guide deep learning, 2021
Emmanuel Abbe, Enric Boix-Adsera, Matthew Brennan, Guy Bresler, and Dheeraj Nagaraj · 2021
Cited alongside, same era.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
On the Opportunities and Risks of Foundation Models, 2021
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Cited alongside, same era.
Truth, lies, and automation: How language models could change disinformation, May 2021
Ben Buchanan, Andrew Lohn, Micah Musser, and Katerina Sedova · 2021
Cited alongside, same era.
Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay · 2021
Our approach to frontier risk, 2023
OpenAI · 2023
Later among the works it cites.
The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only, 2023
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Later among the works it cites.
Goodtriever: Adaptive toxicity mitigation with retrieval-augmented models
Luiza Pozzobon, Beyza Ermis, Patrick Lewis, and Sara Hooker · 2023
Later among the works it cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis, 2023
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Data and parameter scaling laws for neural machine translation, 2021
Prafulla Dhariwal, Girish Sastry, Mark Chen, Dan I. Moldovan, Alex, Beutel, and Jonathan Deaton · 2021
Cited alongside, same era.
Scaling laws for transfer, 2021
Danny Hernandez, Jared Kaplan, Tom Henighan, and Sam McCandlish · 2021
Cited alongside, same era.
The hardware lottery
Sara Hooker · 2021
Cited alongside, same era.
A distributional approach to controlled text generation, 2021
Muhammad Khalifa, Hady Elsahar, and Marc Dymetman · 2021
Cited alongside, same era.
Carbon Emissions and Large Neural Network Training, 2021
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
Cited alongside, same era.
Deep learning on a data diet: Finding important examples early in training, 2021
Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite · 2021
Cited alongside, same era.
Scaling Language Models: Methods, Analysis & Insights from Training Gopher, 2021
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving · 2021
Cited alongside, same era.
Are emergent abilities of large language models a mirage?, 2023
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning, 2023
Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari S. Morcos · 2023
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models, 2023
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, and Abubakar Abid et al · 2023
Later among the works it cites.
Self-influence guided data reweighting for language model pre-training, 2023
Megh Thakkar, Tolga Bolukbasi, Sriram Ganapathy, Shikhar Vashishth, Sarath Chandar, and Partha Talukdar · 2023
Later among the works it cites.
Executive order on the safe, secure, and trustworthy development and use of artificial intelligence, 2023
The White House · 2023
Later among the works it cites.
D4: Improving llm pretraining via document de-duplication and diversification, 2023
Kushal Tirumala, Daniel Simig, Armen Aghajanyan, and Ari S. Morcos · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment, 2023
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf · 2023
Later among the works it cites.
Attention is all you need, 2023
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
Large language models exhibit significant western cultural bias, study finds, 2023
Kyle Wiggers · 2023
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model, 2023
BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, and Suzana Ilić et al · 2023
Later among the works it cites.
SmoothQuant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Later among the works it cites.
Effective long-context scaling of foundation models, 2023
Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, and Hao Ma · 2023
Later among the works it cites.
Adaptive computation with elastic input sequence, 2023
Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab, Neil Houlsby, Mostafa Dehghani, and Yang You · 2023
Later among the works it cites.
Decentralized training of foundation models in heterogeneous environments, 2023
Binhang Yuan, Yongjun He, Jared Quincy Davis, Tianyi Zhang, Tri Dao, Beidi Chen, Percy Liang, Christopher Re, and Ce Zhang · 2023
Later among the works it cites.
Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning, 2023
Ted Zadouri, Ahmet Üstün, Arash Ahmadian, Beyza Ermiş, Acyr Locatelli, and Sara Hooker · 2023
Later among the works it cites.
Synthetic lies: Understanding ai-generated misinformation and evaluating algorithmic and human solutions
Jiawei Zhou, Yixuan Zhang, Qianni Luo, Andrea G. Parker, and Munmun De Choudhury · 2023
Later among the works it cites.
The multilingual alignment prism: Aligning global and local preferences to reduce harm, 2024
Aakanksha, Arash Ahmadian, Beyza Ermis, Seraphina Goldfarb-Tarrant, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker · 2024
Closest in time.
Back to basics: Revisiting reinforce style optimization for learning from human feedback in llms, 2024
Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, and Sara Hooker · 2024
Closest in time.
Llama 3 model card, 2024
AI@Meta · 2024
Closest in time.
Advanced ai evaluations: May update, 2024
AISI · 2024
Closest in time.
A survey on data selection for language models, 2024
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, Colin Raffel, Shiyu Chang, Tatsunori Hashimoto, and William Yang Wang · 2024
Closest in time.
Aya 23: Open weight releases to further multilingual progress, 2024
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Kelly Marchisio, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Phil Blunsom, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker · 2024
Closest in time.
Benchmark early and red team often: A framework for assessing and managing dual-use hazards of ai foundation models, 2024
Anthony M. Barrett, Krystal Jackson, Evan R. Murphy, Nada Madkour, and Jessica Newman · 2024
Closest in time.
Chinchilla scaling: A replication attempt, 2024
Tamay Besiroglu, Ege Erdil, Matthew Barnett, and Josh You · 2024
Closest in time.
xtrimopglm: Unified 100b-scale pre-trained transformer for deciphering the language of protein, 2024
Bo Chen, Xingyi Cheng, Pan Li, Yangli ao Geng, Jing Gong, Shen Li, Zhilei Bei, Xu Tan, Boyan Wang, Xin Zeng, Chiming Liu, Aohan Zeng, Yuxiao Dong, Jie Tang, and Le Song · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
Critical learning periods: Leveraging early training dynamics for efficient data pruning, 2024
Everlyn Asiko Chimoto, Jay Gala, Orevaoghene Ahia, Julia Kreutzer, Bruce A. Bassett, and Sara Hooker · 2024
Closest in time.
C4ai command r+, 2024
Cohere and Cohere For AI Team · 2024
Closest in time.
Trends in the dollar training cost of machine learning systems, 2023
Ben Cottier · 2024
Closest in time.
Rlhf can speak many languages: Unlocking multilingual preference optimization for llms, 2024
John Dang, Arash Ahmadian, Kelly Marchisio, Julia Kreutzer, Ahmet Üstün, and Sara Hooker · 2024
Closest in time.
A journey of a 1,000 kernels begins with a single step: A retrospective of deep learning on gpus
Michael Davies, Ian McDougall, Selvaraj Anandaraj, Deep Machchhar, Rithik Jain, and Karthikeyan Sankaralingam · 2024
Closest in time.
Climatologically invariant scale invariance seen in distributions of cloud horizontal sizes
Tristan D. DeWitt, Timothy J. Garrett, Katherine N. Rees, Christophe Bois, Scott K. Krueger, and Nicolas Ferlay · 2024
Closest in time.
Key trends and figures in machine learning, 2023
Epoch AI · 2024
Closest in time.
Parameter, compute and data trends in machine learning, 2024
Epoch AI · 2024
Closest in time.
Eu artificial intelligence act, 2024
European Union · 2024
Closest in time.
Llms become more “covertly racist” with human intervention, 2024
Karen Hao · 2024
Closest in time.
Governing through the cloud: The intermediary role of compute providers in ai regulation, 2024
Lennart Heim, Tim Fist, Janet Egan, Sihao Huang, Stephen Zekany, Robert Trager, Michael A Osborne, and Noa Zilberman · 2024
Closest in time.
Trends in machine learning hardware, 2023
Marius Hobbhahn, Lennart Heim, and Gökçe Aydos · 2024
Closest in time.
Predicting emergent abilities with infinite resolution evaluation, 2024
Shengding Hu, Xin Liu, Xu Han, Xinrong Zhang, Chaoqun He, Weilin Zhao, Yankai Lin, Ning Ding, Zebin Ou, Guoyang Zeng, Zhiyuan Liu, and Maosong Sun · 2024
Closest in time.
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Closest in time.
Semantic entropy probes: Robust and cheap hallucination detection in llms, 2024
Jannik Kossen, Jiatong Han, Muhammed Razzak, Lisa Schut, Shreshth Malik, and Yarin Gal · 2024
Closest in time.
Can long-context language models subsume retrieval, rag, sql, and more?, 2024
Jinhyuk Lee, Anthony Chen, Zhuyun Dai, Dheeru Dua, Devendra Singh Sachan, Michael Boratko, Yi Luan, Sébastien M. R. Arnold, Vincent Perot, Siddharth Dalmia, Hexiang Hu, Xudong Lin, Panupong Pasupat, Aida Amini, Jeremy R. Cole, Sebastian Riedel, Iftekhar Naim, Ming-Wei Chang, and Kelvin Guu · 2024
Closest in time.
More agents is all you need, 2024
Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, and Deheng Ye · 2024
Closest in time.
Artificial intelligence law of the people’s republic of china (draft for suggestions from scholars), 2024
Zhang Linghan, Yang Jianjun, Cheng Ying, Zhao Jingwu, Han Xuzhi, Zheng Zhifeng, and Xu Xiaoben · 2024
Closest in time.
Sophia: A scalable stochastic second-order optimizer for language model pre-training, 2024
Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma · 2024
Closest in time.
How does quantization affect multilingual llms?, 2024
Kelly Marchisio, Saurabh Dash, Hongyu Chen, Dennis Aumiller, Ahmet Üstün, Sara Hooker, and Sebastian Ruder · 2024
Closest in time.
Biological sequence models in the context of the ai directives, 2024
Nicole Maug, Aidan O’Gara, and Tamay Besiroglu · 2024
Closest in time.
The Operational Risks of AI in Large-Scale Biological Attacks: Results of a Red-Team Study
Christopher A. Mouton, Caleb Lucas, and Ella Guest · 2024
Closest in time.
The impact of ai on the cyber threat, 2024
NCSC · 2024
Closest in time.
Beyond scaling laws: Understanding transformer performance with associative memory, 2024
Xueyan Niu, Bo Bai, Lei Deng, and Wei Han · 2024
Closest in time.
Risk of poverty, 2024
Swiss Federal Statistical Office · 2024
Closest in time.
H.r.8315 - enhancing national frameworks for overseas restriction of critical exports act or enforce act., 2024
U.S. House Committee on Foreign Affairs · 2024
Closest in time.
Building an early warning system for llm-aided biological threat creation, 2024
OpenAI · 2024
Closest in time.
Teach llms to phish: Stealing private information from language models, 2024
Ashwinee Panda, Christopher A. Choquette-Choo, Zhengming Zhang, Yaoqing Yang, and Prateek Mittal · 2024
Closest in time.
Zhen Qin, Daoyuan Chen, Bingchen Qian, Bolin Ding, Yaliang Li, and Shuiguang Deng · 2024
Closest in time.
Generative ai needs adaptive governance, 2024
Anka Reuel and Trond Arne Undheim · 2024
Closest in time.
Position: Technical research and talent is needed for effective AI governance
Anka Reuel, Lisa Soder, Benjamin Bucknall, and Trond Arne Undheim · 2024
Closest in time.
U.s. eyes curbs on china’s access to ai software behind apps like chatgpt, 2024
Reuters · 2024
Closest in time.
Observational scaling laws and the predictability of language model performance, 2024
Yangjun Ruan, Chris J. Maddison, and Tatsunori Hashimoto · 2024
Closest in time.
Senate bill 1047: Safe and secure innovation for frontier artificial intelligence models act., 2024
California Senate · 2024
Closest in time.
Estimating training compute of deep learning models, 2022b
Jaime Sevilla, Lennart Heim, Marius Hobbhahn, Tamay Besiroglu, Anson Ho, and Pablo Villalobos · 2024
Closest in time.
Llm see, llm do: Guiding data generation to target non-differentiable objectives, 2024
Luísa Shimabucoro, Sebastian Ruder, Julia Kreutzer, Marzieh Fadaee, and Sara Hooker · 2024
Closest in time.
Branch-train-mix: Mixing expert llms into a mixture-of-experts llm, 2024
Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma, Hu Xu, Xi Victoria Lin, Baptiste Rozière, Jacob Kahn, Daniel Li, Wen tau Yih, Jason Weston, and Xian Li · 2024
Closest in time.
Scattered mixture-of-experts implementation, 2024
Shawn Tan, Yikang Shen, Rameswar Panda, and Aaron Courville · 2024
Closest in time.
Gemma, 2024
Gemma Team · 2024
Closest in time.
Trading off compute in training and inference, 2023
Pablo Villalobos and David Atkinson · 2024
Closest in time.
The shift from models to compound ai systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi · 2024
Closest in time.
The nation’s top ai safety lab is decaying from within, scientists say, 2024
Cat Zakrzewski · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model, 2024
Ahmet Üstün, Viraat Aryabumi, Zheng-Xin Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker · 2024
Closest in time.