2021

Multilingual training for Software Engineering

Ahmed, Toufique, Devanbu, Premkumar

Understand

Well-trained machine-learning models, which leverage large amounts of open-source software data, have now become an interesting approach to automating many software engineering tasks.

  • Several SE tasks have all been subject to this approach, with performance gradually improving over the past several years with better models and training methods.
  • More, and more diverse, clean, labeled data is better for training; but constructing good-quality datasets is time-consuming and challenging.
  • Ways of augmenting the volume and diversity of clean, labeled data generally have wide applicability.

Reading the bibliography…