Home |

Classifying software changes: Clean or buggy?

This paper introduces a new technique for predicting latent software bugs, called change classification. Change classification uses a machine learning classifier to determine whether a new software change is more similar to prior buggy changes or clean changes. In this manner, change classification predicts the existence of bugs in software changes. The classifier is trained using features (in the machine learning sense) extracted from the revision history of a software project stored in its software configuration management repository. The trained classifier can classify changes as buggy or clean, with a 78 percent accuracy and a 60 percent buggy change recall on average. Change classification has several desirable qualities: 1) The prediction granularity is small (a change to a single file), 2) predictions do not require semantic information about the source code, 3) the technique works for a broad array of project types and programming languages, and 4) predictions can be made immediately upon the completion of a change. Contributions of this paper include a description of the change classification approach, techniques for extracting features from the source code and change histories, a characterization of the performance of change classification across 12 open source projects, and an evaluation of the predictive power of different groups of features.

Topics

Featured Report

Future Skills Centre’s Impact Report: Building a Resilient Workforce in Canada

Latest Reports

Workers Perspectives on Alberta’s Coal Workers Transition Program

Mapping Food Sustainability in Canada

Indigenous Community Health Workers in Canada

Latest News

AI Isn’t Replacing Jobs. It’s Rewriting Them, Experts Tell Canadian Workforce Webinar

Canada’s AI boom is about to collide with a major labour shortage

Future Skills Centre invests to turn PIAAC skills data into Canadian workforce insights

Classifying software changes: Clean or buggy?

Join our mailing list!