Yufeng Wang, Yuan Xu, Anastasia Nikolova, Yuxuan Wang, Jianyu Wang, Chongyang Wang, Xin Tong
5 min
Advances in large language models (LLMs) are profoundly reshaping the field of human-robot interaction (HRI). While prior work has highlighted the technical potential of LLMs, few studies have systematically examined their human-centered impact (e.g., human-oriented understanding, user modeling, and levels of autonomy), making it difficult to consolidate emerging challenges in LLM-driven HRI systems. Therefore, we conducted a systematic literature search following the PRISMA guideline, identifying 86 articles that met our inclusion criteria. Our findings reveal that: (1) LLMs are transforming the fundamentals of HRI by reshaping how robots sense context, generate socially grounded interactions, and maintain continuous alignment with human needs in embodied settings; and (2) current research is largely exploratory, with different studies focusing on different facets of LLM-driven HRI, resulting in wide-ranging choices of experimental setups, study methods, and evaluation metrics. Finally, we identify key design considerations and challenges, offering a coherent overview and guidelines for future research at the intersection of LLMs and HRI.
Human-Robot Interaction (HRI) has long struggled to bridge the gap between technical task execution and the unpredictable nature of human social behavior. The emergence of Large Language Models (LLMs) has provided a new cognitive foundation for robots, enabling them to move beyond pre-scripted responses. This systematic review of 86 empirical studies reveals that LLMs are not merely adding conversational capabilities; they are fundamentally reconfiguring the interaction lifecycle.
The authors propose a new framework to categorize how LLMs are being integrated into robotics:
This review highlights that we are at a critical juncture in robotics. While the technical potential of LLMs is immense, the field is currently fragmented. By synthesizing diverse approaches—from humanoid social robots to functional industrial arms—this work provides a roadmap for researchers to move toward more robust, human-centered systems. It underscores that the future of HRI lies in balancing the generative power of LLMs with the physical constraints and ethical responsibilities of embodied agents.
Alex: Which raises the measurement problem directly. How are researchers actually evaluating success when the behavior they're measuring is this fluid?
Sam: That's where the methodological tension gets sharpest. Traditional robotics leaned on hard metrics—task completion time, positional accuracy. With LLM-driven systems, you're now evaluating things like perceived intelligence, anthropomorphism, and conversational fluidity alongside response latency. The authors note that inter-rater reliability varies significantly across these categories. Morphological classification is tractable. Measuring the quality of social interaction is much noisier—constructs like cognitive load or relational experience resist standardization in ways that make cross-study comparison genuinely difficult.
Alex: So a study on emotional repair in a healthcare setting isn't really comparable to one on conversational fluidity in a domestic robot, even if both claim to be measuring "interaction quality."
Sam: Exactly. And that's the core argument for the taxonomy. By decoupling perception from execution and adding alignment as a distinct, measurable layer, you create a structure where methodology—whether it's a lab experiment or a longitudinal field deployment—can be matched to appropriate evaluation metrics. The authors explicitly invoke PRISMA 2020 as the model: bring that same reporting discipline to this generative era of HRI. [[RP_SECTION:methodological-limitations-and-scope|Methodological limitations and scope]]
Alex: It's essentially asking the field to stop treating each study as its own universe and start building cumulative evidence. What are the limits of the review itself, though? The authors must have made scope decisions that constrain what the taxonomy covers.
Sam: They did, and they're transparent about it. The most consequential exclusion is Late-Breaking Reports—by restricting to fully peer-reviewed archival papers, they've prioritized methodological stability over coverage of the most experimental work currently running in labs. The taxonomy is well-grounded, but it's a snapshot of established science rather than a live feed of the frontier. For a systematic review, that's a defensible trade-off, but a researcher designing a study at the edge of the field should treat it as a floor, not a ceiling. [[RP_SECTION:future-of-longitudinal-deployment|Future of longitudinal deployment]]
Alex: So the practical upshot for someone designing their next study—what does this review actually tell them to do differently?
Sam: Stop treating HRI as a series of one-off demonstrations. The authors' argument is that longitudinal deployments—where the system evolves alongside its users—are where the field needs to go. Behavioral repair shouldn't be treated as a failure mode to be minimized; it should be a standard, expected part of the interaction lifecycle. The shift in mindset is from optimizing initial performance to designing systems with the capacity to coexist and adapt over time. The technology to do that is largely here. What's missing is the scientific infrastructure to validate it rigorously—and that's what this taxonomy is trying to provide.
Alex: Less about the robot getting it right on day one, and more about whether it can still be trusted on day three hundred.
Sam: That's the right frame. And until the field builds that validation infrastructure, the gap between an impressive prototype and a deployable system will remain wider than the demo suggests. Thanks for listening to ResearchPod.