Jasmin Han, Janardan Devkota, Joseph Waring, Amanda Luken, Felix Naughton, Roger Vilardaga, Jonathan Bricker, Carl Latkin, Meghan Moran, Yiqun Chen, Johannes Thrul
6 min
Abstract
Perceived message effectiveness (PME) by potential intervention end-users is important for selecting and optimizing personalized smoking cessation intervention messages for mobile health (mHealth) platform delivery. This study evaluates whether large language models (LLMs) can accurately predict PME for smoking cessation messages. We evaluated multiple models for predicting PME across three domains: content quality, coping support, and quitting support. The dataset comprised 3010 message ratings (5-point Likert scale) from 301 young adult smokers. We compared (1) supervised learning models trained on labeled data, (2) zero and few-shot LLMs prompted without task-specific fine-tuning, and (3) LLM-based digital twins that incorporate individual characteristics and prior PME histories to generate personalized predictions. Model performance was assessed on three held-out messages per participant using accuracy, Cohen's kappa, and F1. LLM-based digital twins outperformed zero and few-shot LLMs (12 percentage points on average) and supervised baselines (13 percentage points), achieving accuracies of 0.49 (content), 0.45 (coping), and 0.49 (quitting), with directional accuracies of 0.75, 0.66, and 0.70 on a simplified 3-point scale. Digital twin predictions showed greater dispersion across rating categories, indicating improved sensitivity to individual differences. Integrating personal profiles with LLMs captures person-specific differences in PME and outperforms supervised and zero and few-shot approaches. Improved PME prediction may enable more tailored intervention content in mHealth. LLM-based digital twins show potential for supporting personalization of mobile smoking cessation and other health behavior change interventions.
Sam: Yes. This shifts from group averages to individual forecasts, which the paper suggests speeds up tailoring texts for apps during cravings. One clear measure was directional accuracy—getting the up-or-down trend right—which improved notably.
Alex: Huh. So instead of slow surveys to guess what works, these twins simulate it fast for each user.
Sam: Exactly. The evidence points to this as a practical step for mobile health tools.
Alex: That ties back to the different therapy styles—like practical steps versus acceptance techniques. How do those play into the messages?
Sam: The messages came from two established approaches. One type gives specific actions to distract from cravings, like doing a quick task to shift focus—that's from cognitive behavioral therapy, or CBT. The other type says it's okay to notice the craving without fighting it, just stay in the moment and let it pass—that's acceptance and commitment therapy, or ACT.
Alex: So half push action and distraction, the other half acceptance. And the people rating them—who were they? How did they split ratings to test predictions fairly?
Sam: The study looked at 301 young adults, ages 18 to 30, who smoked regularly and wanted to quit soon. Each rated 10 messages—five from each approach. They used a within-person setup: seven messages and ratings as history to inform the AI, three new ones held out to check accuracy. This mimics real life, where you have limited past data per user.
Alex: So the twins use that exact split to learn from one person's quirks across CBT and ACT, predicting the holdouts better than simpler baselines?
Sam: The baselines were standard machine learning setups that predict ratings solely from traits like age or smoking habits, without seeing message examples. The digital twins extend few-shot prompting by using that person's own seven ratings as examples, plus their full profile—like building a custom guide from their exact past reactions. This captures personal quirks, such as favoring one message style.
Alex: Huh. So the individual history turns generic prompting into something tailored. How did they check if the twins could actually pick the best messages from a bigger pool?
Sam: They tested by pretending to build a message library for an app. From all the rated ones, they picked the top few that the AI thought would score highest for a person, based on their twin. The digital twin picks consistently got higher real ratings than random selection. The gap to the human-chosen best was small on that five-point scale.
Alex: Huh. So not perfect, but close enough to beat guessing for apps. What limits did they flag?
Sam: The data came only from 301 young adults in an online panel, so results may not hold for older smokers or diverse groups—the paper calls for broader samples. Exact-match accuracy topped at about forty-five to forty-nine percent, likely due to natural wiggles in human ratings; directional accuracy fared better at sixty-six to seventy-five percent, which matters more for picking winners. PME predicts perceived help but not proven quit rates—future work must link to actual behavior changes.
Alex: So the noise caps exact hits, but direction guides practical choices well enough. Balanced view—not flawless, but a step forward.
Sam: Exactly. Overall, the evidence positions these twins as a meaningful tool for personalizing smoking cessation apps, shifting from group trends to individual patterns.
Alex: That's a solid synthesis—thanks for breaking it down, Sam. This look at digital twins for anti-smoking messages shows clear potential in thoughtful personalization. Thanks for listening to ResearchPod.