ResearchPod Summary
Systematic Literature Reviews (SLRs) are foundational to software engineering research, providing a rigorous, transparent, and reproducible method for synthesizing evidence. As Generative AI (GenAI) and Large Language Models (LLMs) gain popularity, researchers are increasingly attempting to use these tools to automate various stages of the SLR process. However, the authors argue that the current methodology for evaluating and using these tools is often naive, risking the integrity of the scientific record. This paper aims to establish a set of preliminary process recommendations, termed GUEST (GenAI Use and Evaluation in SLR Tasks), to guide researchers in both conducting and evaluating GenAI-assisted reviews.
The authors identify several inherent limitations of GenAI that conflict with the requirements of a formal SLR. These include:
The GUEST framework emphasizes that GenAI should act as a 'guest'—a helpful assistant that operates under the constant supervision of human researchers. The authors stress that SLR tasks require validity, traceability, and verifiability. Consequently, any use of GenAI must be fully documented, including the specific model versions, prompts used, and the human validation steps taken to verify the AI's output. The authors argue that researchers must treat GenAI outputs as expert opinions rather than definitive systematic evidence, requiring rigorous human-led verification for every step of the review process.
Without standardized guidelines, the software engineering community risks adopting flawed practices that could undermine the reliability of evidence synthesis. By establishing these preliminary recommendations, the authors provide a necessary framework to ensure that as AI-assisted research matures, it remains grounded in the principles of scientific integrity, transparency, and accountability.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.