Semantic code clone detection using hybrid intermediate representations and BiLSTM networks.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41557716.
- Also identified by DOI 10.1371/journal.pone.0340971 and PMC identifier 12818651.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Semantic code clone detection plays an essential role in software maintenance and quality assurance, as it helps uncover fragments of code that express the same logic even when their syntax has been altered or deliberately obfuscated. In this study, we propose a framework that combines hybrid representation learning with deep bidirectional LSTM networks. The model is applied to two intermediate forms of Java programs-Baf and Jimple-extracted through the Soot framework, which together provide both syntactic structure and semantic detail. This design allows the method to cope with difficult obfuscation strategies such as polymorphism and metamorphism. In our experiments, the framework showed strong and stable performance. Training accuracy reached about 98%, while validation accuracy stayed above 95%, with good generalization across the different clone categories described in the Twilight-Zone taxonomy. When compared with other recurrent models, the BiLSTM consistently performed better, especially when combined with multiple intermediate representations and attention mechanisms. On the BigCloneBench dataset, the approach matched or exceeded the results of state-of-the-art tools, achieving recall and F1-scores of up to 97% on challenging clone types. These findings confirm the practical applicability of hybrid intermediate representations for semantic clone detection and suggest promising directions for future research using transformer-based models and large-scale deployment.
Medical subject headings
- Software
- Semantics