From Alert to Patch: MCP and APM in AI Agent-Assisted Bug Resolution
DOI:
https://doi.org/10.37497/opsbrazil.38Keywords:
Artificial Intelligence, Software Engineering, Intelligent Agents, Model Context Protocol, Application Performance Monitoring, Observability, Failure Diagnosis, Automated Bug ResolutionResumo
The increasing complexity of distributed digital systems has amplified the challenges associated with software failure diagnosis and resolution. In modern environments, essential information for incident investigation is distributed across multiple sources, including alerts, traces, logs, metrics, deployment records, technical documentation, and source code. Although artificial intelligence-based agents have demonstrated the ability to analyze code, execute tests, and propose modifications, a scientific gap remains regarding how structured access to operational context influences the quality and efficiency of these activities.
This paper presents an experimental design to investigate whether read-only access to observability data through the Model Context Protocol (MCP) can assist artificial intelligence agents in diagnosing failures and producing verifiable patches. The study proposes a comparison among three experimental conditions: investigation based on static information, investigation supported by manual Application Performance Monitoring (APM) tools, and investigation assisted by agents connected to MCP-enabled tools.
The research does not present results from the controlled experiment and does not assume prior improvements in time, accuracy, or quality. However, it includes motivation based on anonymized industrial observations conducted with an MCP-assisted agent and a commercial APM platform. The objective is to establish a reproducible protocol capable of evaluating different dimensions of the software bug resolution process, including time to identify the root cause, quality of produced patches, evidence used, need for human intervention, and risks associated with the use of intelligent agents.
The expected contribution is to provide a scientific methodology for evaluating how context architectures can influence the performance of artificial intelligence agents in complex software engineering tasks, promoting greater transparency, traceability, and reproducibility.
Downloads
Referências
Beyer, B.; Jones, C.; Petoff, J.; Murphy, N. R. Site Reliability Engineering: How Google Runs Production Systems. O’Reilly Media, 2016.
Beyer, B.; Murphy, N. R.; Rensin, D. K.; Kawahara, K.; Thorne, S. The Site Reliability Workbo-ok: Practical Ways to Implement SRE. O’Reilly Media, 2018.
Google Cloud. DORA Research: Accelerate State of DevOps Report. Google Cloud, multiple years.
OpenTelemetry Authors. OpenTelemetry Documentation and Specification. Cloud Native Compu-ting Foundation.
Zhang, C.; Yang, J.; et al. Large Language Models for Software Engineering: A Systematic Lite-rature Review. ACM/IEEE Software Engineering Research Literature, 2023.
Chen, M.; Tworek, J.; Jun, H.; et al. Evaluating Large Language Models Trained on Code. arXiv, 2021.
Pearl, J. Causality: Models, Reasoning, and Inference. Cambridge University Press, 2009.
Sculley, D.; et al. Hidden Technical Debt in Machine Learning Systems. Advances in Neural In-formation Processing Systems, 2015.
Downloads
Postado
Categorias
Licença
Copyright (c) 2026 Open Science Brazil

Este trabalho está licenciado sob uma licença Creative Commons Attribution 4.0 International License.






