Fetching the paper…
Reading the bibliography…
Many natural language processing (NLP) tasks are naturally imbalanced, as some target categories occur much more frequently than others in the real world.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George F. Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. 2019 · 1907
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Gated recurrent neural network approach for multilabel emotion detection in microblogs
Prabod Rathnayaka, Supun Abeysinghe, Chamod Samarajeewa, Isura Manchanayake, Malaka J. Walpola, Rashmika Nawaratne, Tharindu R. Bandaragoda, and Damminda Alahakoon. 2019 · 1907
Earlier work this paper cites.
LVIS: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollár, and Ross B. Girshick. 2019b · 1908
Earlier work this paper cites.
Classification of imbalanced remote-sensing data by neural networks
L. Bruzzone and S.B. Serpico. 1997 · 1997
Earlier work this paper cites.
Learning from imbalanced data sets: a comparison of various strategies
Nathalie Japkowicz et al. 2000 · 2000
Earlier work this paper cites.
A mixture-of-experts framework for text classification
Andrew Estabrooks and Nathalie Japkowicz. 2001 · 2001
Earlier work this paper cites.
Smote: Synthetic minority over-sampling technique
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. 2002 · 2002
Earlier work this paper cites.
A boosted maximum entropy model for learning text chunking
Seong-Bae Park and Byoung-Tak Zhang. 2002 · 2002
Earlier work this paper cites.
Latent dirichlet allocation
David M. Blei, Andrew Y. Ng, and Michael I. Jordan. 2003 · 2003
Earlier work this paper cites.
Neighbor-weighted k-nearest neighbor for unbalanced text corpus
Songbo Tan. 2005 · 2005
Earlier work this paper cites.
Multilabel neural networks with applications to functional genomics and text categorization
Min-Ling Zhang and Zhi-Hua Zhou. 2006 · 2006
Earlier work this paper cites.
A survey of active learning for text classification using deep neural networks
Christopher Schröder and Andreas Niekler. 2020 · 2008
Earlier work this paper cites.
Reducing class imbalance during active learning for named entity annotation
Katrin Tomanek and Udo Hahn. 2009 · 2009
Earlier work this paper cites.
Haiyang Yu, Ningyu Zhang, Shumin Deng, Zonggang Yuan, Yantao Jia, and Huajun Chen. 2020 · 2009
Earlier work this paper cites.
On data augmentation for extreme multi-label classification
Danqing Zhang, Tao Li, Haiyang Zhang, and Bing Yin. 2020 · 2009
Earlier work this paper cites.
The balanced accuracy and its posterior distribution
Kay Henning Brodersen, Cheng Soon Ong, Klaas Enno Stephan, and Joachim M. Buhmann. 2010 · 2010
Earlier work this paper cites.
Imbalanced sentiment classification
Shoushan Li, Guodong Zhou, Zhongqing Wang, Sophia Yat Mei Lee, and Rangyang Wang. 2011 · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Adapt-and-adjust: Overcoming the long-tail problem of multilingual speech recognition
Genta Indra Winata, Guangsen Wang, Caiming Xiong, and Steven C. H. Hoi. 2020 · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
Addressing class imbalance for improved recognition of implicit discourse relations
Junyi Jessy Li and Ani Nenkova. 2014 · 2014
Earlier work this paper cites.
A SICK cure for the evaluation of compositional distributional semantic models
Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014 · 2014
Earlier work this paper cites.
Addressing imbalance in multilabel classification: Measures and random resampling algorithms
Francisco Charte, Antonio J. Rivera, María J. del Jesus, and Francisco Herrera. 2015 · 2015
Earlier work this paper cites.
Classifying relations by ranking with convolutional neural networks
Cícero dos Santos, Bing Xiang, and Bowen Zhou. 2015 · 2015
Earlier work this paper cites.
Addressing class imbalance in grammatical error detection with evaluation metric optimization
Anoop Kunchukuttan and Pushpak Bhattacharyya. 2015 · 2015
Earlier work this paper cites.
V-net: Fully convolutional neural networks for volumetric medical image segmentation
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. 2016 · 2016
Earlier work this paper cites.
Relay backpropagation for effective learning of deep convolutional neural networks
Li Shen, Zhouchen Lin, and Qingming Huang. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on Twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Ranking convolutional recurrent neural networks for purchase stage identification on imbalanced Twitter data
Heike Adel, Francine Chen, and Yan-Ying Chen. 2017 · 2017
Earlier work this paper cites.
Focal loss for dense object detection
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollar. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Learning deep latent space for multi-label classification
Chih-Kuan Yeh, Wei-Chieh Wu, Wei-Jen Ko, and Yu-Chiang Frank Wang. 2017 · 2017
Earlier work this paper cites.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and Devon Hjelm. 2018 · 2018
Earlier work this paper cites.
A systematic study of the class imbalance problem in convolutional neural networks
Mateusz Buda, Atsuto Maki, and Maciej A. Mazurowski. 2018 · 2018
Earlier work this paper cites.
Smote for learning from imbalanced data: Progress and challenges, marking the 15-year anniversary
Alberto Fernández, Salvador Garcia, Francisco Herrera, and Nitesh V Chawla. 2018 · 2018
Earlier work this paper cites.
Hierarchical relation extraction with coarse-to-fine grained attention
Xu Han, Pengfei Yu, Zhiyuan Liu, Maosong Sun, and Peng Li. 2018 · 2018
Cited alongside, same era.
Boosted cascaded convnets for multilabel classification of thoracic diseases in chest radiographs
Pulkit Kumar, Monika Grewal, and Muktabh Mayank Srivastava. 2018 · 2018
Cited alongside, same era.
Explainable prediction of medical codes from clinical text
James Mullenbach, Sarah Wiegreffe, Jon Duke, Jimeng Sun, and Jacob Eisenstein. 2018 · 2018
Cited alongside, same era.
Dynamic sampling in convolutional neural networks for imbalanced data classification
Samira Pouyanfar, Yudong Tao, Anup Mohan, Haiman Tian, Ahmed S. Kaseb, Kent Gauen, Ryan Dailey, Sarah Aghajanzadeh, Yung-Hsiang Lu, Shu-Ching Chen, and Mei-Ling Shyu. 2018 · 2018
Cited alongside, same era.
Scoring and classifying implicit positive interpretations: A challenge of class imbalance
Chantal van Son, Roser Morante, Lora Aroyo, and Piek Vossen. 2018 · 2018
Cited alongside, same era.
Distribution-balanced loss for multi-label classification in long-tailed datasets
Tong Wu, Qingqiu Huang, Ziwei Liu, Yu Wang, and Dahua Lin. 2020 · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020 · 2020
Later among the works it cites.
HSCNN: A hybrid-Siamese convolutional neural network for extremely imbalanced multi-label text classification
Wenshuo Yang, Jiyi Li, Fumiyo Fukumoto, and Yanming Ye. 2020 · 2020
Later among the works it cites.
Drill: Dynamic representations for imbalanced lifelong learning
Kyra Ahrens, Fares Abawi, and Stefan Wermter. 2021 · 2021
Later among the works it cites.
Handling extreme class imbalance in technical logbook datasets
Farhad Akhbardeh, Cecilia Ovesdotter Alm, Marcos Zampieri, and Travis Desell. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning imbalanced datasets with label-distribution-aware margin loss
Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Synthetic propaganda embeddings to train a linear projection
Adam Ek and Mehdi Ghanimifard. 2019 · 2019
Cited alongside, same era.
Incorporating label dependencies in multilabel stance detection
William Ferreira and Andreas Vlachos. 2019 · 2019
Cited alongside, same era.
Deep learning and thresholding with class-imbalanced big data
Justin M. Johnson and Taghi M. Khoshgoftaar. 2019a · 2019
Cited alongside, same era.
Cost-sensitive regularization for label confusion-aware event detection
Hongyu Lin, Yaojie Lu, Xianpei Han, and Le Sun. 2019 · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Biplob Biswas, Thai-Hoang Pham, and Ping Zhang. 2021 · 2021
Later among the works it cites.
Supercharging imbalanced data learning with energy-based contrastive representation transfer
Junya Chen, Zidi Xiu, Benjamin Goldstein, Ricardo Henao, Lawrence Carin, and Chenyang Tao. 2021 · 2021
Later among the works it cites.
A survey of data augmentation approaches for NLP
Steven Y. Feng, Varun Gangal, Jason Wei, Sarath Chandar, Soroush Vosoughi, Teruko Mitamura, and Eduard Hovy. 2021 · 2021
Later among the works it cites.
Applying occam’s razor to transformer-based dependency parsing: What works, what doesn’t, and what is really necessary
Stefan Grünewald, Annemarie Friedrich, and Jonas Kuhn. 2021 · 2021
Later among the works it cites.
A survey on recent approaches for natural language processing in low-resource scenarios
Michael A. Hedderich, Lukas Lange, Heike Adel, Jannik Strötgen, and Dietrich Klakow. 2021 · 2021
Later among the works it cites.
Balancing methods for multi-label text classification with long-tailed class distribution
Yi Huang, Buse Giledereli, Abdullatif Köksal, Arzucan Özgür, and Elif Ozkirimli. 2021 · 2021
Later among the works it cites.
Sequential targeting: A continual learning approach for data imbalance in text classification
Joel Jang, Yoonjeon Kim, Kyoungho Choi, and Sungho Suh. 2021 · 2021
Later among the works it cites.
Textcut: A multi-region replacement data augmentation approach for text imbalance classification
Wanrong Jiang, Ya Chen, Hao Fu, and Guiquan Liu. 2021 · 2021
Later among the works it cites.
To share or not to share: Predicting sets of sources for model transfer learning
Lukas Lange, Jannik Strötgen, Heike Adel, and Dietrich Klakow. 2021 · 2021
Later among the works it cites.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021 · 2021
Later among the works it cites.
Uncovering main causalities for long-tailed information extraction
Guoshun Nan, Jiaqi Zeng, Rui Qiao, Zhijiang Guo, and Wei Lu. 2021 · 2021
Later among the works it cites.
Swiss-judgment-prediction: A multilingual legal judgment prediction benchmark
Joel Niklaus, Ilias Chalkidis, and Matthias Stürmer. 2021 · 2021
Later among the works it cites.
Supertagging the long tail with tree-structured decoding of complex categories
Jakob Prange, Nathan Schneider, and Vivek Srikumar. 2021 · 2021
Later among the works it cites.
A multi-task approach to neural multi-label hierarchical patent classification using transformers
Subhash Chandra Pujari, Annemarie Friedrich, and Jannik Strötgen. 2021 · 2021
Later among the works it cites.
Multitask semi-supervised learning for class-imbalanced discourse classification
Alexander Spangher, Jonathan May, Sz-Rung Shiang, and Lingjia Deng. 2021 · 2021
Later among the works it cites.
Not all negatives are equal: Label-aware contrastive loss for fine-grained text classification
Varsha Suresh and Desmond Ong. 2021 · 2021
Later among the works it cites.
Re-embedding difficult samples via mutual information constrained semantically oversampling for imbalanced text classification
Jiachen Tian, Shizhan Chen, Xiaowang Zhang, Zhiyong Feng, Deyi Xiong, Shaojuan Wu, and Chunliu Dou. 2021 · 2021
Later among the works it cites.
Multi-task learning in argument mining for persuasive online discussions
Nhat Tran and Diane Litman. 2021 · 2021
Later among the works it cites.
Good-enough example extrapolation
Jason Wei. 2021 · 2021
Later among the works it cites.
Discovering topics in long-tailed corpora with causal intervention
Xiaobao Wu, Chunping Li, and Yishu Miao. 2021 · 2021
Later among the works it cites.
Delving into deep imbalanced regression
Yuzhe Yang, Kaiwen Zha, Ying-Cong Chen, Hao Wang, and Dina Katabi. 2021 · 2021
Later among the works it cites.
Multi-label sentiment analysis on 100 languages with dynamic weighting for label imbalance
Selim F. Yilmaz, E. Batuhan Kaynak, Aykut Koç, Hamdi Dibeklioğlu, and Suleyman Serdar Kozat. 2021 · 2021
Later among the works it cites.
An improved baseline for sentence-level relation extraction
Wenxuan Zhou and Muhao Chen. 2021 · 2021
Later among the works it cites.
Alleviating asr long-tailed problem by decoupling the learning of representation and classification
Keqi Deng, Gaofeng Cheng, Runyan Yang, and Yonghong Yan. 2022 · 2022
Closest in time.
Why only micro-f1? class weighting of measures for relation classification
David Harbecke, Yuxuan Chen, Leonhard Hennig, and Christoph Alt. 2022 · 2022
Closest in time.
Segcn-dcr: A syntax-enhanced event detection framework with decoupled classification rebalance
Bo Hu, Yun Liu, Naiyue Chen, Lifu Wang, Ning Liu, and Xing Cao. 2022 · 2022
Closest in time.
Self-supervised learning is more robust to dataset imbalance
Hong Liu, Jeff Z. HaoChen, Adrien Gaidon, and Tengyu Ma. 2022 · 2022
Closest in time.
True few-shot learning with Prompts—A real-world perspective
Timo Schick and Hinrich Schütze. 2022 · 2022
Closest in time.
Memorisation versus generalisation in pre-trained language models
Michael Tänzer, Sebastian Ruder, and Marek Rei. 2022 · 2022
Closest in time.
Reprint: a randomized extrapolation based on principal components for data augmentation
Jiale Wei, Qiyuan Chen, Pai Peng, Benjamin Guedj, and Le Li. 2022 · 2022
Closest in time.
Adam: An attentional data augmentation method for extreme multi-label text classification
Jiaxin Zhang, Jie Liu, Shaowei Chen, Shaoxin Lin, Bingquan Wang, and Shanpeng Wang. 2022 · 2022
Closest in time.
Fairness-aware class imbalanced learning
Shivashankar Subramanian, Afshin Rahimi, Timothy Baldwin, Trevor Cohn, and Lea Frermann. 2021 · 2051
Closest in time.