Fetching the paper…
Reading the bibliography…
In this paper, we introduce harmonic loss as an alternative supervisory signal for training neural networks and large language models (LLMs).
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations
Carole H Sudre, Wenqi Li, Tom Vercauteren, Sebastien Ourselin, and M Jorge Cardoso · 2017
Earlier work this paper cites.
Tversky loss function for image segmentation using 3d fully convolutional deep networks
Seyed Sadegh Mohseni Salehi, Deniz Erdogmus, and Ali Gholipour · 2017
Earlier work this paper cites.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2017
Earlier work this paper cites.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2018
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Supervised contrastive learning
Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan · 2020
Earlier work this paper cites.
Towards out-of-distribution generalization: A survey
Jiashuo Liu, Zheyan Shen, Yue He, Xingxuan Zhang, Renzhe Xu, Han Yu, and Peng Cui · 2021
Earlier work this paper cites.
Implicit representations of meaning in neural language models
Belinda Z Li, Maxwell Nye, and Jacob Andreas · 2021
Cited alongside, same era.
Can language models encode perceptual structure without grounding? a case study in color
Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, and Anders Søgaard · 2021
Cited alongside, same era.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra · 2022
Cited alongside, same era.
A comprehensive survey of loss functions in machine learning
Qi Wang, Yue Ma, Kun Zhao, and Yingjie Tian · 2022
Cited alongside, same era.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Later among the works it cites.
Mechanistic interpretability for ai safety–a review
Leonard Bereska and Efstratios Gavves · 2024
Later among the works it cites.
Patch diffusion: Faster and more data-efficient training of diffusion models
Zhendong Wang, Yifan Jiang, Huangjie Zheng, Peihao Wang, Pengcheng He, Zhangyang Wang, Weizhu Chen, Mingyuan Zhou, et al · 2024
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2024
Later among the works it cites.
Monotonic representation of numeric properties in language models
Benjamin Heinzerling and Kentaro Inui · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eric Todd, Millicent L Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau · 2023
Cited alongside, same era.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2023
Cited alongside, same era.
Samuel Marks and Max Tegmark · 2023
Cited alongside, same era.
Topology-aware focal loss for 3d image segmentation
Andac Demir, Elie Massaad, and Bulent Kiziltan · 2023
Cited alongside, same era.
Hybrid wind speed forecasting using iceemdan and transformer model with novel loss function
Bala Saibabu Bommidi, Kiran Teeparthi, and Vishalteja Kosana · 2023
Cited alongside, same era.
Contrastive learning models for sentence representations
Lingling Xu, Haoran Xie, Zongxi Li, Fu Lee Wang, Weiming Wang, and Qing Li · 2023
Cited alongside, same era.
Deepspeed data efficiency: Improving deep learning model quality and training efficiency via efficient data sampling and routing
Conglong Li, Zhewei Yao, Xiaoxia Wu, Minjia Zhang, Connor Holmes, Cheng Li, and Yuxiong He
Cited in the paper.
Omnigrok: Grokking beyond algorithmic data
Ziming Liu, Eric J Michaud, and Max Tegmark
Cited in the paper.
Later among the works it cites.
Opening the ai black box: program synthesis via mechanistic interpretability
Eric J Michaud, Isaac Liao, Vedang Lad, Ziming Liu, Anish Mudide, Chloe Loughridge, Zifan Carl Guo, Tara Rezaei Kheirkhah, Mateja Vukelić, and Max Tegmark · 2024
Later among the works it cites.
Not all language model features are linear
Joshua Engels, Isaac Liao, Eric J Michaud, Wes Gurnee, and Max Tegmark · 2024
Later among the works it cites.
Echocardiographic image segmentation with vision transformers: A comparative analysis of different loss functions
Edoardo Bosco, Giovanni Magenes, and Giulia Matrone · 2024
Later among the works it cites.
Pedro Seber · 2024
Later among the works it cites.
I-con: A unifying framework for representation learning
Shaden Alshammari, John Hershey, Axel Feldmann, William T Freeman, and Mark Hamilton · 2025
Closest in time.