Fetching the paper…
Reading the bibliography…
Evaluating modern ML models is hard.
Avoiding a tragedy of the commons in the peer review process
D Sculley, Jasper Snoek, and Alex Wiltschko · 1901
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 1905
Earlier work this paper cites.
How to lie with statistics
Darrell Huff · 1954
Earlier work this paper cites.
Supercollider physics
E Eichten, Ian Hinchliffe, K Lane, and Chris Quigg · 1984
Earlier work this paper cites.
How not to lie with statistics: the correct way to summarize benchmark results
Philip J Fleming and John J Wallace · 1986
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vladimir Vapnik · 1991
Earlier work this paper cites.
The Challenger launch decision: Risky technology, culture, and deviance at NASA
Diane Vaughan · 1996
Earlier work this paper cites.
Data set selection
Doudou LaLoudouana, Mambobo Bonouliqui Tarare, Lupano Tecallonou Center, and GUANA Selacie · 2003
Earlier work this paper cites.
Lessons from a failure: Generating tailored smoking cessation letters
Ehud Reiter, Roma Robertson, and Liesl M Osman · 2003
Earlier work this paper cites.
Core Vector Machines: Fast SVM Training on Very Large Data Sets
Ivor W. Tsang, James T. Kwok, and Pak-Ming Cheung · 2005
Earlier work this paper cites.
An empirical comparison of supervised learning algorithms
Rich Caruana and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Comments on the “Core Vector Machines: Fast SVM Training on Very Large Data Sets”
Gaëlle Loosli and Stéphane Canu · 2007
Earlier work this paper cites.
Logical and relational learning
Luc De Raedt · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Learning semantic correspondences with less supervision
Percy Liang, Michael I Jordan, and Dan Klein · 2009
Earlier work this paper cites.
Rethinking embedding coupling in pre-trained language models, 2020
Hyung Won Chung, Thibault Févry, Henry Tsai, Melvin Johnson, and Sebastian Ruder · 2010
Earlier work this paper cites.
Extended-Connectivity Fingerprints
David Rogers and Mathew Hahn · 2010
Earlier work this paper cites.
False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant
Joseph P Simmons, Leif D Nelson, and Uri Simonsohn · 2011
Earlier work this paper cites.
Measuring the prevalence of questionable research practices with incentives for truth telling
Leslie K John, George Loewenstein, and Drazen Prelec · 2012
Earlier work this paper cites.
Open evaluation: a vision for entirely transparent post-publication peer review and rating for science
Nikolaus Kriegeskorte · 2012
Earlier work this paper cites.
Machine learning versus statistical modeling
Anne-Laure Boulesteix and Matthias Schmid · 2014
Earlier work this paper cites.
The statistical crisis in science data-dependent analysis—a ‘garden of forking paths’—explains why many statistically significant comparisons don’t hold up
A Gelman and E Lokem · 2014
Earlier work this paper cites.
Generative Adversarial Nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J. Smola · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
The reusable holdout: Preserving validity in adaptive data analysis
Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Roth · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Baidu Fires Researcher Tied to Contest Disqualification
John Markoff · 2015
Earlier work this paper cites.
Is psychology suffering from a replication crisis? What does “failure to replicate” really mean?
Scott E Maxwell, Michael Y Lau, and George S Howard · 2015
Earlier work this paper cites.
Repeatability in computer systems research
Christian Collberg and Todd A Proebsting · 2016
Earlier work this paper cites.
Increasing transparency through a multiverse analysis
Sara Steegen, Francis Tuerlinckx, Andrew Gelman, and Wolf Vanpaemel · 2016
Earlier work this paper cites.
Degrees of freedom in planning, running, analyzing, and reporting psychological studies: A checklist to avoid p-hacking
Jelte M Wicherts, Coosje LS Veldkamp, Hilde EM Augusteijn, Marjan Bakker, Robbie CM Van Aert, and Marcel ALM Van Assen · 2016
Earlier work this paper cites.
Do reviewers review the code submitted with papers?, 2017
Ulderique Demoitre · 2017
Earlier work this paper cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
You need to understand your corpora! the weathergov example
Ehud Reiter · 2017
Earlier work this paper cites.
The ”Perfect Score” Script
Oleg Trott · 2017
Earlier work this paper cites.
Gpu version gives non-reproducible results even with random seed set
u39kun · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Learning explanatory rules from noisy data
Richard Evans and Edward Grefenstette · 2018
Earlier work this paper cites.
State of the art: Reproducibility in artificial intelligence
Odd Erik Gundersen and Sigbjørn Kjensmo · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Earlier work this paper cites.
Missing data hinder replication of artificial intelligence studies
Matthew Hutson · 2018
Earlier work this paper cites.
Are gans created equal? a large-scale study
Mario Lucic, Karol Kurach, Marcin Michalski, Sylvain Gelly, and Olivier Bousquet · 2018
Earlier work this paper cites.
Sentence encoders on stilts: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman · 2018
Earlier work this paper cites.
Do CIFAR-10 classifiers generalize to CIFAR-10?, 2018
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2018
Earlier work this paper cites.
I tried to reproduce results from a CVPR18 paper, here’s what I found, 2018
tkinter76 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Moleculenet: a benchmark for molecular machine learning
Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay Pande · 2018
Earlier work this paper cites.
Unreproducible Research is Reproducible
Xavier Bouthillier, César Laurent, and Pascal Vincent · 2019
Earlier work this paper cites.
On the Measure of Intelligence, 2019
François Chollet · 2019
Earlier work this paper cites.
Maurizio Ferrari Dacrema, Simone Boglio, Paolo Cremonesi, and Dietmar Jannach · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A Smith · 2019
Earlier work this paper cites.
Detecting computer-generated random responding in questionnaire-based data: A comparison of seven indices
Marc Dupuis, Emanuele Meier, and Félix Cuneo · 2019
Earlier work this paper cites.
Reflecting on neural ODEs
David Duvenaud · 2019
Earlier work this paper cites.
HARK Side of Deep Learning–From Grad Student Descent to Automated Machine Learning
Oguzhan Gencoglu, Mark van Gils, Esin Guldogan, Chamin Morikawa, Mehmet Süzen, Mathias Gruber, Jussi Leinonen, and Heikki Huttunen · 2019
Earlier work this paper cites.
PetFinder.my Contest: 1st Place Winner Disqualified
Mongrel Jedi · 2019
Earlier work this paper cites.
Troubling Trends in Machine Learning Scholarship: Some ML papers suffer from flaws that could mislead the public and stymie future research
Zachary C. Lipton and Jacob Steinhardt · 2019
Earlier work this paper cites.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
A step toward quantifying independently reproducible machine learning research
Edward Raff · 2019
Earlier work this paper cites.
On the difficulty of evaluating baselines: A study on recommender systems
Steffen Rendle, Li Zhang, and Yehuda Koren · 2019
Earlier work this paper cites.
Researcher degrees of freedom in phonetic research
Timo B Roettger · 2019
Earlier work this paper cites.
Critically examining the” neural hype” weak baselines and the additivity of effectiveness gains from neural ranking models
Wei Yang, Kuang Lu, Peilin Yang, and Jimmy Lin · 2019
Earlier work this paper cites.
Do we train on test data? purging CIFAR of near-duplicates
Benedikt Barz and Joachim Denzler · 2020
Earlier work this paper cites.
Time to rethink the publication process in machine learning
Yoshua Bengio · 2020
Earlier work this paper cites.
Experiment tracking with Weights and Biases, 2020
Lukas Biewald · 2020
Earlier work this paper cites.
Replication Issues in AI research
Denny Britz · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Developments in MLflow: A system to accelerate the machine learning lifecycle
Andrew Chen, Andrew Chow, Andrew Davidson, Annanya DCunha, Ali Ghodsi, Stephanie Ann Hong, Andy Konwinski, Clemens Mewald, Simeon Murching, Tomas Nykodym, and Paul Ogilvie · 2020
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
The reproducibility crisis in Machine Learning
Marius Hobbhahn · 2020
Earlier work this paper cites.
I tried a bunch of things: The dangers of unexpected overfitting in classification of brain data
Mahan Hosseini, Michael Powell, John Collins, Chloe Callahan-Flintoft, William Jones, Howard Bowman, and Brad Wyble · 2020
Earlier work this paper cites.
The shape of and solutions to the MTurk quality crisis
Ryan Kennedy, Scott Clifford, Tyler Burleigh, Philip D. Waggoner, Ryan Jewell, and Nicholas J. G. Winter · 2020
Earlier work this paper cites.
Preventing Multiple Comparisons Problems in Data Exploration and Machine Learning
Nikolaos Koulouris · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive NLP tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al · 2020
Earlier work this paper cites.
BERTs of a feather do not generalize together: Large variability in generalization across models with similar test set performance, 2020
R. Thomas McCoy, Junghyun Min, and Tal Linzen · 2020
Earlier work this paper cites.
Got bots? Practical recommendations to protect online survey data from bot attacks
Andie Storozuk, Marilyn Ashley, Véronic Delage, and Erin A Maloney · 2020
Earlier work this paper cites.
HARKing, cherry-picking, p-hacking, fishing expeditions, and data dredging and mining as questionable research practices
Chittaranjan Andrade · 2021
Earlier work this paper cites.
Rip van Winkle’s Razor: A Simple Estimate of Overfit to Test Data, 2021
Sanjeev Arora and Yi Zhang · 2021
Earlier work this paper cites.
The NeurIPS 2021 Consistency Experiment
Alina Beygelzimer, Yann Dauphin, Percy Liang, and Jennifer Wortman Vaughan · 2021
Earlier work this paper cites.
The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail
Samuel R Bowman · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Cited alongside, same era.
AI-assisted peer review
Alessandro Checco, Lorenzo Bracciale, Pierpaolo Loreti, Stephen Pinfield, and Giuseppe Bianchi · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Cited alongside, same era.
Mostafa Dehghani, Anurag Arnab, Lucas Beyer, Ashish Vaswani, and Yi Tay · 2021
Cited alongside, same era.
A troubling analysis of reproducibility and progress in recommender systems research
Maurizio Ferrari Dacrema, Simone Boglio, Paolo Cremonesi, and Dietmar Jannach · 2021
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer · 2023
Later among the works it cites.
Llama 2 model card
Ruan Silva · 2023
Later among the works it cites.
Heap: Hierarchical policies for web actions using llms
P. Sodhi, S.R.K. Branavan, and R. McDonald · 2023
Later among the works it cites.
On the paper “Exploring the MIT Mathematics and EECS Curriculum Using Large Language Models”
Armando Solar-Lezama, Tonio Buonassisi, and Yoon Kim · 2023
Later among the works it cites.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, 2023
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R. Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The multiplicity of analysis strategies jeopardizes replicability: lessons learned across disciplines
Sabine Hoffmann, Felix Schönbrodt, Ralf Elsas, Rory Wilson, Ulrich Strasser, and Anne-Laure Boulesteix · 2021
Cited alongside, same era.
The Privatization of AI Research(-ers): Causes and Potential Consequences – From university-industry interaction to public research brain-drain?, 2021
Roman Jurowetzki, Daniel Hain, Juan Mateos-Garcia, and Konstantinos Stathoulopoulos · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, et al · 2021
Cited alongside, same era.
Why p-values should be interpreted as p-values and not as measures of evidence
Daniel Lakens · 2021
Cited alongside, same era.
Are We Learning Yet? A Meta Review of Evaluation Failures Across Machine Learning
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt · 2021
Cited alongside, same era.
How to avoid machine learning pitfalls: a guide for academic researchers, 2021
Michael A. Lones · 2021
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp · 2021
Cited alongside, same era.
The New XOR Problem
Shawn Tan · 2023
Later among the works it cites.
The Evolved Code Alpaca Dataset
theblackcat102 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
All Jailbreaks seem to be patched
u/wiicrafttech · 2023
Later among the works it cites.
Gemini in reasoning: Unveiling commonsense in multimodal large language models
Yuqing Wang and Yun Zhao · 2023
Later among the works it cites.
Zero-shot information extraction via chatting with ChatGPT
Xiang Wei, Xingyu Cui, Ning Cheng, Xiaobin Wang, Xin Zhang, Shen Huang, Pengjun Xie, Jinan Xu, Yufeng Chen, Meishan Zhang, et al · 2023
Later among the works it cites.
WizardLM: Empowering Large Language Models to Follow Complex Instructions, 2023
Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, and Daxin Jiang · 2023
Later among the works it cites.
Rethinking benchmark and contamination for language models with rephrased samples
Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E Gonzalez, and Ion Stoica · 2023
Later among the works it cites.
Exploring the MIT mathematics and EECS curriculum using large language models
Sarah J Zhang, Samuel Florin, Ariel N Lee, Eamon Niknafs, Andrei Marginean, Annie Wang, Keith Tyser, Zad Chin, Yann Hicke, Nikhil Singh, et al · 2023
Later among the works it cites.
Many-Shot In-Context Learning, 2024
Rishabh Agarwal, Avi Singh, Lei M. Zhang, Bernd Bohnet, Luis Rosias, Stephanie Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D. Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle · 2024
Closest in time.
Inspect: An open-source framework for large language model evaluations
AI Safety Institute · 2024
Closest in time.
Llama 3 Model Card
AI@Meta · 2024
Closest in time.
The Claude 3 Model Family: Opus, Sonnet, Haiku, 2024
Anthropic · 2024
Closest in time.
Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs
Simone Balloccu, Patrícia Schmidtová, Mateusz Lango, and Ondřej Dušek · 2024
Closest in time.
Llama copyright drama: Meta stops disclosing what data it uses to train the company’s giant ai models
Alistair Barr · 2024
Closest in time.
Chinchilla Scaling: A replication attempt, 2024
Tamay Besiroglu, Ege Erdil, Matthew Barnett, and Josh You · 2024
Closest in time.
Lessons from the Trenches on Reproducible Evaluation of Language Models, 2024
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A. Wang, Genta Indra Winata, François Yvon, and Andy Zou · 2024
Closest in time.
Can You Trust An AI Press Release?
Lawrence Chan · 2024
Closest in time.
Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
Nishanth Chandran, Sunayana Sitaram, Divya Gupta, Rahul Sharma, Kashish Mittal, and Manohar Swaminathan · 2024
Closest in time.
Symbolic discovery of optimization algorithms
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, et al · 2024
Closest in time.
Chatbot Arena: An open platform for evaluating LLMs by human preference, 2024
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Closest in time.
ARC-AGI Kaggle, 2024
Francois Chollet · 2024
Closest in time.
Chat GPT DAN (and other Jailbreaks), 2024
Copamine · 2024
Closest in time.
ConStat: Performance-Based Contamination Detection in Large Language Models
Jasper Dekoninck, Mark Niklas Müller, and Martin Vechev · 2024
Closest in time.
Training on the test task confounds evaluation and emergence
Ricardo Dominguez-Olmedo, Florian E Dorner, and Moritz Hardt · 2024
Closest in time.
We call this the [benchmark decoration] phase of LLM development usually happens towards the end of pretrain sometimes mixed with SFT
Yao Fu · 2024
Closest in time.
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, et al · 2024
Closest in time.
Zamba: A Compact 7B SSM hybrid model
Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge · 2024
Closest in time.
Changing Answer Order Can Decrease MMLU Accuracy
V. Gupta, D. Pantoja, C. Ross, A. Williams, and M. Ung · 2024
Closest in time.
Lm contamination index
zentroa HiTZ · 2024
Closest in time.
Revisiting the evidence on thermostatic response to democratic change: degrees of democratic support or researcher degrees of freedom?
Yue Hu, Yuehong Cassandra Tai, and Frederick Solt · 2024
Closest in time.
Common corpus
HuggingFace · 2024
Closest in time.
No train no gain: Revisiting efficient training algorithms for transformer-based language models
Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, and Matt J Kusner · 2024
Closest in time.
S. Kapoor, B. Stroebl, Z.S. Siegel, N. Nadgir, and A. Narayanan · 2024
Closest in time.
Let’s face it… Just because you don’t know what’s in your training data, you can not just call it zero-shot, 2024
Hilda Kuehne · 2024
Closest in time.
ChatGPT D̈AN(̈and other J̈ailbreaks)̈, 2024
Kiho Lee · 2024
Closest in time.
Ten Hard Problems in Artificial Intelligence We Must Get Right
Gavin Leech, Simson Garfinkel, Misha Yagudin, Alexander Briand, and Aleksandr Zhuravlev · 2024
Closest in time.
Task contamination: Language models may not be few-shot anymore
Changmao Li and Jeffrey Flanigan · 2024
Closest in time.
Does style matter? Disentangling style and substance in Chatbot Arena
Tianle Li, Anastasios Angelopoulos, and Wei-Lin Chiang · 2024
Closest in time.
How should you prompt an LM for MMLU?
Percy Liang · 2024
Closest in time.
Weixin Liang, Zachary Izzo, Yaohui Zhang, Haley Lepp, Hancheng Cao, Xuandong Zhao, Lingjiao Chen, Haotian Ye, Sheng Liu, Zhi Huang, et al · 2024
Closest in time.
Rethinking open source generative AI: open washing and the EU AI Act
Andreas Liesenfeld and Mark Dingemanse · 2024
Closest in time.
No need to lift a finger anymore? assessing the quality of code generation by chatgpt
Zhijie Liu, Yutian Tang, Xiapu Luo, Yuming Zhou, and Liang Feng Zhang · 2024
Closest in time.
You need to be spending more money on evals
Kamilė Lukošiūtė · 2024
Closest in time.
Addressing researcher degrees of freedom through minP adjustment
Maximilian M Mandl, Andrea S Becker-Pennrich, Ludwig C Hinske, Sabine Hoffmann, and Anne-Laure Boulesteix · 2024
Closest in time.
On Leakage of Code Generation Evaluation Datasets, 2024
Alexandre Matton, Tom Sherborne, Dennis Aumiller, Elena Tommasone, Milad Alizadeh, Jingyi He, Raymond Ma, Maxime Voisin, Ellen Gilsenan-McMahon, and Matthias Gallé · 2024
Closest in time.
Weak baselines and reporting biases lead to overoptimism in machine learning for fluid-related partial differential equations
Nick McGreivy and Ammar Hakim · 2024
Closest in time.
Ai and the copyright liability overhang: A brief summary of the current state of ai-related copyright litigation
David M. McIntosh, Georgina Jones Suzuki, and Yam Schaal · 2024
Closest in time.
How much should we read into LLM Chatbot Arena leaderboard ranks?
Ashutosh Mehra · 2024
Closest in time.
Gemma: Open Models Based on Gemini Research and Technology
Gemma Team Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, L. Sifre, Morgane Riviere, Mihir Kale, J Christopher Love, Pouya Dehghani Tafti, L’eonard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Am’elie H’eliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Cl’ement Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikula, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Pier Giuseppe Sessa, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vladimir Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Brian Warkentin, Ludovic Peran, Minh Giang, Cl’ement Farabet, Oriol Vinyals, Jeffrey Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy · 2024
Closest in time.
The unconditioned distribution of current open LLMs, 2024
Beren Millidge · 2024
Closest in time.
Project author team stay tuned: I found out that the llama3-V project is stealing a lot of academic work from MiniCPM-Llama3-V 2.5
MiniCPM-V · 2024
Closest in time.
Soliciting Participants for the NeurIPS 2024 Checklist Assistant Study, 2024
Chairs NeurIPS · 2024
Closest in time.
Training on the Benchmark Is Not All You Need
Shiwen Ni, Xiangtao Kong, Chengming Li, Xiping Hu, Ruifeng Xu, Jia Zhu, and Min Yang · 2024
Closest in time.
Shortplav, 2024
niplav · 2024
Closest in time.
ChatGPT — Release Notes, 2024
OpenAI · 2024
Closest in time.
Hello Qwen2
Team Qwen · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.
Reka core, flash, and edge: A series of powerful multimodal language models
Team Reka · 2024
Closest in time.
Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models, 2024
Martin Riddell, Ansong Ni, and Arman Cohan · 2024
Closest in time.
To the Cutoff… and Beyond? A Longitudinal Perspective on LLM data contamination
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley · 2024
Closest in time.
The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates
Giuseppe Russo Latona, Manoel Horta Ribeiro, Tim R Davidson, Veniamin Veselovsky, and Robert West · 2024
Closest in time.
A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Vinija Jain, Samrat Mondal, and Aman Chadha · 2024
Closest in time.
Data-contamination-database
Oscar Sainz, Iker García Ferrero, Eneko Agirre, Jon Ander Campos, Alon Jacovi, Yanai Elazar, and Yoav Goldberg · 2024
Closest in time.
SEAL leaderboards, 2024
Scale AI · 2024
Closest in time.
Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr · 2024
Closest in time.
An Overview of Challenges, Experiments, and Computational Solutions in Peer Review (Extended Version)
Nihar B Shah · 2024
Closest in time.
Reflexion: Language agents with verbal reinforcement learning
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao · 2024
Closest in time.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Closest in time.
Lo-Hi: practical ML drug discovery benchmark
S. Steshin · 2024
Closest in time.
D - If you say in a paper you provide code, it better be code
u/chatterbox272 · 2024
Closest in time.
Created a custom instruction that generates copyright images
u/danneh02 · 2024
Closest in time.
D - Calling out the authors of Trajformer paper for missing code
u/UIPDsmokes · 2024
Closest in time.
Magicoder: Empowering Code Generation with OSS-Instruct, 2024
Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang · 2024
Closest in time.
Good Seed Makes a Good Crop: Discovering Secret Seeds in Text-to-Image Diffusion Models
Katherine Xu, Lingzhi Zhang, and Jianbo Shi · 2024
Closest in time.
Proposal for Renaming of Gemma-7B Model to Gemma-9B
ZeroShot-AI · 2024
Closest in time.
A Careful Examination of Large Language Model Performance on Grade School Arithmetic
Hugh Zhang, Jeff Da, Dean Lee, Vaughn Robinson, Catherine Wu, Will Song, Tiffany Zhao, Pranav Raja, Dylan Slack, Qin Lyu, et al · 2024
Closest in time.
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang, Peiyi Wang, Runxin Xu, Y. Wu, Yukun Li, Huazuo Gao, Shirong Ma, Wangding Zeng, Xiao Bi, Zihui Gu, Hanwei Xu, Damai Dai, Kai Dong, Liyue Zhang, Yishi Piao, Zhibin Gou, Zhenda Xie, Zhewen Hao, Bingxuan Wang, Junxiao Song, Deli Chen, Xin Xie, Kang Guan, Yuxiang You, Aixin Liu, Qiushi Du, Wenjun Gao, Xuan Lu, Qinyu Chen, Yaohui Wang, Chengqi Deng, Jiashi Li, Chenggang Zhao, Chong Ruan, Fuli Luo, and Wenfeng Liang · 2024
Closest in time.