ResearchPod Summary
Talk2AI is a groundbreaking dataset capturing 3,080 conversations (30,800 turns) between 770 Italian adults and large language models (LLMs) over four weekly sessions. Collected in Spring 2025, it focuses on how AI can persuade humans on real-world topics like climate change, math anxiety, and health misinformation. Unlike prior studies that measure attitude change in one-off interactions, Talk2AI tracks shifts longitudinally—revealing how repeated chats with AI influence beliefs, opinions, and behaviors over time.
Participants were randomly assigned to one of four LLMs: GPT-4o, Claude Sonnet 3.7, DeepSeek-chat V3, or Mistral Large. Each engaged in 10-turn discussions per topic per session, followed by feedback on opinion change, conviction stability, perceived AI humanness, and behavioral intentions. Rich metadata—including sociodemographics and psychometric profiles—links to every conversation, enabling analysis of who is most swayed by AI and why.
Most AI persuasion research treats change as a 'single endpoint' after brief exposure, missing the dynamics of ongoing dialogue. Talk2AI's within-subject, four-wave design (weekly sessions) captures temporal evolution: How does conviction waver? Does trust in AI grow? Do effects compound or fade?
This setup disentangles AI's role as a 'continuous persuasive agent,' influenced by model architecture, user traits, and context. For instance, early sessions might build rapport (boosting perceived humanness), while later ones test sustained influence on attitudes.
Core Components:
Processing: Data cleaned, translated (EN, ES, NL, DE, PT, FR), and structured for ML/NLP analysis. Figure 1 in the paper outlines the workflow: recruitment → profiling → interact → feedback → repeat.
Alex: Welcome to another episode of ResearchPod.
Sam: Today we're discussing the Talk2AI dataset. It's a collection of over 3,000 conversations between 770 Italian adults and AI language models, gathered over four weekly sessions in spring 2025.
Alex: So this is about people chatting with AI on topics like climate change or math anxiety... and tracking how those talks might shift their views over time?
Sam: Exactly. Most studies so far look at what happens in just one quick conversation—like a snapshot that misses the full picture. But in real life, people talk to AI repeatedly, and opinions can evolve gradually over weeks. This dataset captures that longer process by having the same people debate fixed topics each week with the same AI model, whether GPT-4o, Claude Sonnet 3.7, DeepSeek-chat V3, or Mistral Large.
Alex: Right, so instead of one-and-done experiments, it's the same folks coming back week after week... as their own comparison?
Sam: Yes, researchers call this a within-subject design—meaning each person acts as their own baseline, so you can spot real changes in their beliefs without mixing up different people's starting points. After every session, they reported on things like how steady their convictions felt, if their opinions shifted, how human-like the AI seemed, and even intentions to act, like donating to a related cause.
Alex: Okay, that addresses a gap then—the time factor. But what made them focus on these specific topics and models?
Sam: The topics—climate change, math anxiety, health misinformation—are ones where beliefs matter in daily life and can be debated without forcing preset positions. Participants were matched to one model randomly, letting comparisons across architectures while controlling for individual differences. This setup, with quality checks to filter incomplete data, creates a solid base for studying how AI dialogue influences attitudes over a month, not just instantly.
Alex: Huh. So it's not just about whether AI persuades... but how repeated exposure interacts with who you are.
Sam: Precisely. The dataset's strength lies in tying 30,800 conversation turns to psychometric profiles—like need for cognition or Big Five traits—enabling analysis of those moderating effects over time.
AI is everywhere in info-seeking and debate. Talk2AI quantifies its persuasive power at scale, addressing gaps like:
Findings preview: LLMs induce measurable shifts, but effects vary by model (e.g., Claude vs. Mistral) and user (e.g., open personalities more receptive). This scales to digital ecosystems—think social media bots or chat advisors shaping public opinion.
At ~400MB, Talk2AI is public-ready for cognitive science, HCI, and AI safety. It proves AI isn't just a tool—it's a dialogue partner reshaping minds, one turn at a time.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.
Alex: So tying those traits to the chats over time sounds useful... but how did they actually run these sessions week after week without people dropping out or giving sloppy answers?
Sam: They built a web platform that combined surveys with a live chat window, starting with basic info like age, job, family size, and education at sign-up. Each week, before chatting, participants filled out short questionnaires on their current mindset—things like how much they enjoy thinking deeply about ideas, or how comfortable they feel with AI tools.
Alex: Okay, so repeat those mindset checks each time to track if they stay steady... then the chat itself. What kept the talks focused and deep?
Sam: Users got randomly assigned one AI model and one topic—like climate change or math worries—that stayed the same all four weeks, for ten back-and-forth turns each session. To spark real debate, the first user message had to be at least 50 words, and pop-up hints suggested ways to push back, like asking for proof or giving counterpoints. Hidden instructions told the AI to spot flaws in user arguments early on, then keep replies short, without letting topics drift. After ten turns, they rated conviction strength in their views, any opinion shift, how human the AI felt, and did a quick choice: split 100 euros between themselves and a related charity—the less they kept, the more persuaded they seemed.
Alex: That charity split as a persuasion check is clever... but with thousands signing up, how did they end up with reliable data from just 770 people?
Sam: They ran a strict cleaning process on the raw files—over 2,600 sign-ups at first. They tossed sessions missing the full 20 messages from glitches, cut users who gave identical answers across questions, and kept only those finishing all four weeks completely. This left 770 solid participants, with data split into JSON files for demographics, psychometrics and ratings, chats, plus a ready-to-analyze CSV linking traits to feedback scores.
Alex: Huh... so the filters ensured each person's full journey was tracked consistently, traits to persuasion shifts.
Sam: Yes. That rigor lets researchers link steady personality measures to how beliefs evolved over weeks, a meaningful advance for studying repeated AI influence.
Alex: So with those clean profiles linked to chats... what exactly did they ask right after each conversation to measure persuasion?
Sam: After the ten turns, participants answered five quick questions. The first asked how convinced they felt of their starting arguments, on a scale from one to 100. The second gauged opinion shift on the topic from zero meaning no change to 100 for a full flip. Third was how human-like the AI chat felt, one to 100. Fourth imagined splitting 100 euros between self and a related charity. Last was an open write-up of at least 50 words on their final thoughts.
Alex: Right, so beyond describing the data... what could this mean for building better AI coaches?
Sam: The paper suggests it could help create AI that tailors persuasion to a person's traits—like using thinking style to adjust arguments on climate action or misinformation. For instance, train models to predict if someone high in mental diligence shifts views faster, scaling talks for real therapy or public campaigns. But results depend on future tests; this dataset provides the foundation.
Alex: That sounds practical... though fixed ten-turn chats might cut off deeper talks, right? And starting in Italian could skew translations for other languages.
Sam: Exactly—those are noted limits. Shorter sessions keep data clean but miss free-flowing debates; Italian originals work for that language but need care for multilingual tools, as machine translations might lose nuances. Still, the within-subject tracking over weeks offers a clear step forward in understanding repeated AI influence.
Alex: Well put. That's our look at the Talk2AI dataset—a solid base for linking who you are to how chats reshape beliefs long-term. Thanks for listening to ResearchPod.