Pemanfaatan Local Large Language Model dalam Software Effort Estimation Menggunakan Metode LOOCV

Authors

DOI:

https://doi.org/10.24843/MITE.2026.v25.01.p13

Keywords:

Software Effort Estimation; Large Language Model; Phi-3; Akurasi Estimasi; LOOCV

Abstract

 Software Effort Estimation (SEE) in the early stages of a project is generally confronted with minimal technical specifications and high ambiguity in requirement documents. Large Language Models (LLMs) offer the potential to automate information extraction from unstructured natural language text. This exploratory study evaluates the standalone performance of a local LLM (Phi-3 Mini architecture) in predicting man-days effort using a within-company historical dataset from a local digital agency. The experiment was conducted using the Leave-One-Out Cross-Validation (LOOCV) method to ensure objective evaluation within a limited data population. Model optimization utilized Chain-of-Thought (CoT) prompt engineering combined with a methodological strategy of few-shot selection based on boundary distribution (Min, Median, Max) to serve as cognitive anchors. Qualitative analysis proves that the LLM possesses high semantic intelligence in evaluating relative project complexity. However, quantitative evaluation using MMRE, PRED(25), and Coefficient of Determination (R-squared) indicates fundamental limitations in standalone mathematical regression. The model achieved an MMRE of 0.6963 and a PRED(25) of 0.25, performing below the Mean Baseline (MMRE 0.44) with a negative R-squared of -2.038. These results confirm that while LLMs are proficient in understanding context, they struggle with precise numerical interpolation. Consequently, this study recommends a hybrid estimation framework for future research, positioning the LLM purely as a qualitative feature extractor for conventional Machine Learning algorithms to produce robust and precise estimation accuracy.

Downloads

Download data is not yet available.

References

[1] M. Abdin, et al., "Phi-3 technical report: A highly capable language model locally on your phone," arXiv preprint arXiv:2404.14219, 2024.

[2] C. N. Coelho Jr, et al., "Effort and size estimation in software projects with large language model-based intelligent interfaces," arXiv preprint arXiv:2402.07158, 2024.

[3] M. Jørgensen, "A review of studies on expert estimation of software development effort," Journal of Systems and Software, vol. 70, no. 1-2, pp. 37-60, 2004.

[4] Y. Mahmood, N. Kama, A. Azmi, A. S. Khan, and M. Ali, "Software effort estimation accuracy prediction of machine learning techniques: A systematic performance evaluation," arXiv preprint arXiv:2101.10658, 2021.

[5] Ł. Radliński and J. Swacha, "Large language models for early-stage software project estimation: A systematic mapping study," Applied Sciences, vol. 15, no. 24, p. 13099, 2025.

[6] S. A. Salihu, K. B. Saliu, and O. A. Owoyemi, "A systematic literature review of machine learning and AutoML in software effort estimation," International Journal of Advanced Computer Science and Applications, vol. 15, no. 8, 2024.

[7] N. Tran, T. Tran, and N. Nguyen, "Leveraging AI for enhanced software effort estimation: A comprehensive study and framework proposal," arXiv preprint arXiv:2402.05484, 2024.

[8] E. Kocaguneli, T. Menzies, and J. W. Keung, "On the value of ensemble effort estimation," IEEE Transactions on Software Engineering, vol. 38, no. 3, pp. 1403-1416, 2012.

[9] J. Wei, et al., "Chain-of-thought prompting elicits reasoning in large language models," Advances in Neural Information Processing Systems, vol. 35, pp. 24824-24837, 2022.

[10] T. Brown, et al., "Language models are few-shot learners," Advances in Neural Information Processing Systems, vol. 33, pp. 1877-1901, 2020.

[11] T. Foss, E. Stensrud, B. Kitchenham, and I. Myrtveit, "A simulation study of the model evaluation criterion MMRE," IEEE Transactions on Software Engineering, vol. 29, no. 11, pp. 985-995, Nov. 2003.

[12] X. Hou, et al., "Large language models for software engineering: A systematic literature review," ACM Transactions on Software Engineering and Methodology, vol. 33, no. 5, pp. 1-41, 2023.

[13] A. Fan, et al., "Large language models for software engineering: Survey and open problems," arXiv preprint arXiv:2310.03533, 2023.

[14] B. W. Boehm, Software engineering economics. Englewood Cliffs, NJ, USA: Prentice-Hall, 1981.

[15] B. A. Kitchenham, E. Mendes, and G. H. Travassos, "Cross versus within-company cost estimation studies: A systematic review," IEEE Transactions on Software Engineering, vol. 28, no. 5, pp. 316-329, 2002.

[16] J. Wen, S. Li, Z. Lin, Y. Hu, and C. Huang, "Systematic literature review of machine learning based software development effort estimation models," Information and Software Technology, vol. 54, no. 1, pp. 41-59, 2012.

[17] I. P. S. Handika, I. A. Giriantari, and A. Dharma, "Perbandingan Metode Extreme Learning Machine dan Particle Swarm Optimization Extreme Learning Machine untuk Peramalan Jumlah Penjualan Barang," Majalah Ilmiah Teknologi Elektro, vol. 15, no. 1, pp. 84-90, 2016.

[18] N. L. K. T. W. M.U, I. M. O. Widyantara, and N. M. A. E. Dewi W, "Deteksi Tipe Modulasi Digital Pada Automatic Modulation Recognition Menggunakan Support Vector Machine dan Conjugate Gradient Polak Ribiere-Backpropagation," Majalah Ilmiah Teknologi Elektro, vol. 18, no. 2, pp. 281-286, 2019.

[19] T. Kojima, A. Gu, M. Reid, Y. Matsuo, and S. Iwasawa, "Large language models are zero-shot reasoners," Advances in Neural Information Processing Systems, vol. 35, pp. 22199-22213, 2022.

[20] I. Myrtveit, E. Stensrud, and M. Shepperd, "Reliability and validity in comparative studies of software prediction models," IEEE Transactions on Software Engineering, vol. 31, no. 5, pp. 380-391, 2005.

Downloads

Published

2026-07-31