Machine learning has been proven to be susceptible to carefully crafted samples, known as adversarial examples.
The generation of these adversarial examples helps to make the models more robust and gives us an insight into the underlying decision-making of these models.
Over the years, researchers have successfully attacked image classifiers in both, white and black-box settings.
However, these methods are not directly applicable to texts as text data is discrete.
TextDecepter: Hard Label Black Box Attack on Text Classifiers · Around
Built on
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
W. Medhat, A. Hassan, and H. Korashy, “Sentiment analysis algorithms and applications: A survey,” Ain Shams engineering journal , vol. 5, no. 4, pp. 1093–1113, 2014
F. Hill, R. Reichart, and A. Korhonen, “Simlex-999: Evaluating semantic models with (genuine) similarity estimation,” Computational Linguistics , vol. 41, no. 4, pp. 665–695, 2015
2015
Earlier work this paper cites.
Similar
C. Nobata, J. Tetreault, A. Thomas, Y. Mehdad, and Y. Chang, “Abusive language detection in online user content,” in Proceedings of the 25th international conference on world wide web , 2016, pp. 145–153
J. Gao, J. Lanchantin, M. L. Soffa, and Y. Qi, “Black-box generation of adversarial text sequences to evade deep learning classifiers,” in 2018 IEEE Security and Privacy Workshops (SPW) . IEEE, 2018, pp. 50–56