Icing on the Cake: Automatic Code Summarization at Ericsson
Giriprasad Sridhara, Sujoy Roychowdhury, Sumit Soman, Ranjani H G, Ricardo Britto
TL;DR
The paper tackles the challenge of automatic Java method summarization to aid software maintenance in Ericsson. It systematically compares the SOTA ASAP baseline, which uses static analysis and exemplar prompts, against four lightweight prompting strategies that operate solely on the method body, including a concise WordRestrict prompt. Across Ericsson and open-source Java projects, and using multiple LLMs, the simpler prompts achieve equal or better performance on eight similarity metrics, with robust behavior under method-name masking. This work suggests a practical path to faster, more robust code summarization in commercial environments without heavy reliance on static analysis or exemplar corpora, and demonstrates generalizability through replication on Guava and Elasticsearch.
Abstract
This paper presents our findings on the automatic summarization of Java methods within Ericsson, a global telecommunications company. We evaluate the performance of an approach called Automatic Semantic Augmentation of Prompts (ASAP), which uses a Large Language Model (LLM) to generate leading summary comments for Java methods. ASAP enhances the $LLM's$ prompt context by integrating static program analysis and information retrieval techniques to identify similar exemplar methods along with their developer-written Javadocs, and serves as the baseline in our study. In contrast, we explore and compare the performance of four simpler approaches that do not require static program analysis, information retrieval, or the presence of exemplars as in the ASAP method. Our methods rely solely on the Java method body as input, making them lightweight and more suitable for rapid deployment in commercial software development environments. We conducted experiments on an Ericsson software project and replicated the study using two widely-used open-source Java projects, Guava and Elasticsearch, to ensure the reliability of our results. Performance was measured across eight metrics that capture various aspects of similarity. Notably, one of our simpler approaches performed as well as or better than the ASAP method on both the Ericsson project and the open-source projects. Additionally, we performed an ablation study to examine the impact of method names on Javadoc summary generation across our four proposed approaches and the ASAP method. By masking the method names and observing the generated summaries, we found that our approaches were statistically significantly less influenced by the absence of method names compared to the baseline. This suggests that our methods are more robust to variations in method names and may derive summaries more comprehensively from the method body than the ASAP approach.
