Fetching the paper…
Reading the bibliography…
The 'macro F1' metric is frequently used to evaluate binary, multi-class and multi-label classification problems.
A systematic analysis of performance measures for classification tasks
Marina Sokolova and Guy Lapalme · 2009
Earlier work this paper cites.
A comparative analysis of classification methods to multi-label tasks in different application domains
A Santos, A Canuto, and Antonino Feitosa Neto · 2011
Earlier work this paper cites.
Optimal thresholding of classifiers to maximize f1 measure
Zachary C Lipton, Charles Elkan, and Balakrishnan Naryanaswamy · 2014
Earlier work this paper cites.
Semeval-2015 task 10: Sentiment analysis in twitter
Sara Rosenthal, Preslav Nakov, Svetlana Kiritchenko, Saif Mohammad, Alan Ritter, and Veselin Stoyanov · 2015
Cited alongside, same era.
A unified view of multi-label performance measures
Xi-Zhu Wu and Zhi-Hua Zhou · 2017
Cited alongside, same era.
Neural-davidsonian semantic proto-role labeling
Rachel Rudinger, Adam Teichert, Ryan Culkin, Sheng Zhang, and Benjamin Van Durme · 2018
Later among the works it cites.
An argument-marker model for syntax-agnostic proto-role labeling
Juri Opitz and Anette Frank · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…