Publication:
Automatic Software Categorization Using Ensemble Methods and Bytecode Analysis

No Thumbnail Available

Date

2017-09

Authors

Çatal, Çağatay
Tugul, Serkan
Akpınar, Başar

Journal Title

Journal ISSN

Volume Title

Publisher

World Scientific Publ Co Pte Ltd, 5 Toh Tuck Link, Singapore 596224, Singapore

Research Projects

Organizational Units

Journal Issue

Abstract

Software repositories consist of thousands of applications and the manual categorization of these applications into domain categories is very expensive and time-consuming. In this study, we investigate the use of an ensemble of classifiers approach to solve the automatic software categorization problem when the source code is not available. Therefore, we used three data sets (package level/class level/method level) that belong to 745 closed-source Java applications from the Sharejar repository. We applied the Vote algorithm, AdaBoost, and Bagging ensemble methods and the base classifiers were Support Vector Machines, Naive Bayes, J48, IBk, and Random Forests. The best performance was achieved when the Vote algorithm was used. The base classifiers of the Vote algorithm were AdaBoost with J48, AdaBoost with Random Forest, and Random Forest algorithms. We showed that the Vote approach with method attributes provides the best performance for automatic software categorization; these results demonstrate that the proposed approach can effectively categorize applications into domain categories in the absence of source code.

Description

Keywords

Software categorization, machine learning, software repository, bytecode

Citation