AI agents that communicate on behalf of individuals need to capture how each person actually communicates, yet current approaches either require costly per-person fine-tuning, produce generic outputs from shallow persona descriptions, or optimize preferences without modeling communication style. We present ASPECT (Automated Social Psychometric Evaluation of Communication Traits), a pipeline that directs LLMs to assess constructs from a validated communication scale against behavioral evidence from workplace data, without per-person training. In a case study with 20 participants (1,840 paired item ratings, 600 scenario evaluations), ASPECT-generated profiles achieved moderate alignment with self-assessments, and ASPECT-generated responses were preferred over generic and self-report baselines on aggregate, with substantial variation across individuals and scenarios. During the profile review phase, linked evidence helped participants identify mischaracterizations, recalibrate their own self-ratings, and negotiate context-appropriate representations. We discuss implications for building inspectable, individually scoped communication profiles that let individuals control how agents represent them at work.
Alex: Welcome to another episode of ResearchPod. Today, we're looking at a practical issue with AI: when these systems handle emails or meetings for you while you're away, they often sound too bland and corporate, not like your own voice.
Sam: Right. This paper introduces ASPECT, a method that uses a large language model—a computer program trained on vast amounts of text to understand and generate human-like writing—to analyze your actual emails and chats. It builds a profile of your unique communication style, like how direct or chatty you are, without needing special training just for you. The central puzzle is how to make AI speak *as* you, capturing those personal patterns scalably and transparently.
Alex: So this is basically tackling why current AI agents sound generic, even with some personalization? And the core problem is figuring out your private style from everyday work data without heavy customization?
Sam: Exactly. Existing ways either give the AI a short description that leads to shallow copies, or they fine-tune the model on your data alone—which takes lots of time and computer power per person, and hides how it decides things. ASPECT changes that by guiding the AI with a tested framework from psychology: sets of questions that reliably measure communication traits, like being straightforward or elaborate. The AI pulls specific examples from your messages to score those traits, linking each score back to quotes you can check.
Alex: Okay, that makes sense for avoiding the 'corporate robot' feel. But how does pulling those examples actually capture something as varied as how someone talks to a boss versus a teammate?
Sam: It scans your data piece by piece, focusing on one trait at a time—like spotting moments of brevity in status updates. This creates a structured summary that's inspectable: you see the evidence and can adjust it. The study with 20 people showed these profiles led to AI replies preferred over plain or self-described ones, though results varied by person and situation.
Alex: What makes this profile different from just asking someone to describe themselves?
Sam: People often describe themselves inconsistently or overlook habits they don't notice. The system starts by having them fill out the same set of questions they use for measuring communication traits, rating statements like how direct they are in emails. Then it shows their ratings side-by-side with the AI's scores from analyzing their actual messages, highlighting big differences and linking each score to 2 to 5 short quotes pulled straight from their chats or meetings as proof.
Alex: So users get to see the evidence right there, like a report card with notes from their own words? That sounds like it lets them spot mistakes or tweak things.
Sam: Yes. They can flag wrong quotes or adjust scores, turning it into a back-and-forth where they negotiate what fits. After that review, the study tested it in fake work scenarios—like replying to a teammate about a deadline delay, customized with their real team names but same setup for everyone. Three AI replies per scenario: one generic with no personal info, one using their self-description, and one from the profile with evidence snippets.
Alex: Same structure across people, but personalized details so it feels real? How did the profiled one stack up?
Sam: In those tests with 20 participants, the replies using the profiled style got ranked highest for matching how they'd actually respond—better than self-reports or generic ones. Users rated them a clear step up on a 1-to-5 scale for sounding like themselves. The evidence links helped explain why, and while the AI sometimes leaned too positive, the relative patterns held steady.
Alex: Huh. So it's not perfect, but the transparency lets people fix the biases before using it.
Sam: Precisely. This auditing makes the profile usable in practice, shifting from black-box guesses to something collaborative.
Alex: Before that step, how close was the AI's initial guess to what people rated about themselves?
Sam: The study compared the AI's scores to participants' own ratings on 92 statements about communication habits, like whether they talk a lot or stay concise. Overall, exact matches happened in about one in four cases, with typical differences of roughly one point on a five-point scale. More importantly, the AI captured each person's relative pattern—the traits where they score high or low compared to their other traits—with a moderate rank match of about 0.39 on average.
Alex: A quarter exact, one-point average gap, and decent pattern matching. That suggests it's getting the shape right even if absolute numbers drift—maybe because some traits show up clearly in words, others less so?
Sam: Yes. It worked better for obvious traits signaled directly in text, like anger through sharp words or humor via jokes, with smaller gaps there. Subtler ones, depending more on hidden intent or context, had larger mismatches. Still, across six main categories of style—like how expressive or precise someone is—the relative shapes held up reasonably.
Alex: During auditing, did people mostly reject the AI, or did it change their minds sometimes?
Sam: Researchers reviewed comments from all 23 trait areas per person, sorting them into types like full agreement, rejecting the AI while sticking to self-view, or meeting in the middle. In about 17 percent of those reviews, people shifted their own ratings—either fully toward the AI after seeing the quotes, or compromising. This showed mismatches often came from people underestimating habits they forgot, or rating an ideal self instead of their actual workplace behavior.
Alex: Huh—so the evidence prompts real reflection, like realizing you structure messages more than you thought. That bidirectional tweak makes the profile more reliable for AI use.
Sam: Exactly. Auditing reveals self-biases and context nuances, turning initial estimates into collaborative, evidence-backed representations that preserve key patterns while allowing calibration. The paper notes this process strengthens the whole system for realistic scenarios.
Alex: Did those refined profiles actually make the AI's replies better in practice—like in simulated work situations?
Sam: Yes, the study tested that directly. They created ten workplace scenarios for each of the twenty people, like responding to a deadline slip or planning a team update, using real names from their work but the same basic setup. For each, the AI generated three replies: one plain and neutral, one based on the person's self-ratings alone, and one using the profiled style with its evidence links. On average across all six hundred judgments, the profiled replies ranked highest, coming first about 43 percent of the time—clearly ahead of the others.
Alex: Ranked highest overall, but results varied by person. What explained those differences—did some scenarios just suit generic replies better?
Sam: Individual preferences dominated the patterns. About half the participants strongly favored profiled replies, seeing them as more true to their usual way of organizing thoughts or balancing firmness with support. But four preferred generic ones, often because profiled versions added too much enthusiasm that felt off for casual coordination tasks, like quick schedule changes. Even self-reports won once, for those wanting a safer, low-key tone. The study clusters people this way, showing no one-size-fits-all.
Alex: So in structured spots—like defending an approach—the profile shone because it drew from real history. But what about the mismatches during auditing—did the paper pinpoint common reasons the AI got traits wrong at first?
Sam: They cataloged five main sources from participant feedback. First, data gaps: the logs miss off-record chats or rare events, so traits from those stay hidden. Second, mistaking situation for style—like reading a manager's structured talk as always dominant, when it's just the role. Third, LLMs literalize tone: sarcasm or jokes get scored as real aggression or commands. Fourth, fuzzy definitions: people and the scale disagree on what counts as argumentative versus helpful pushback. Fifth, evidence mishandling: over-weighting one example or mixing up who said what in meetings.
Alex: That typology makes the limits clear and fixable through review. Overall, it points to keeping users central, picking the right measurement tools, and grounding everything in actual behavior.
Sam: Precisely. The work stresses data for evidence, validated scales like CSI for structure, and human oversight to add missing context—building profiles that users trust for AI stand-ins. The study highlights how profiles respect boundaries, like distinguishing your everyday work style from other sides of yourself, since the data comes only from professional chats. Participants often wanted to dial back certain traits for specific situations, avoiding a full average that might expose too much.
Alex: Scoped like different modes for boss emails versus peer check-ins? But weren't there spots where the AI still got things systematically off, even after review?
Sam: Yes, the reviews uncovered predictable over-ratings, especially on preciseness—scoring people higher by about 1.7 points on average—likely because work messages are polished and structured by default, masking natural messiness. Traits hard to spot in text, like quiet thinking before speaking, also skewed low when evidence was thin. The paper notes real constraints: data from one organization and just text chats over 90 days, mostly tech workers okay with AI, so it may not stretch to other jobs or casual talks. Prompting the language model kept things simple and checkable but couldn't fully fix biases like occasional mix-ups in who said what. No long-term tests either—just initial sessions with made-up scenarios.
Alex: Fair points; it's a starting snapshot, not a full picture. Still, for quick audits in controlled spots like missing a meeting, those portable profiles seem ready to deploy without the creepy fine-tune feel.
Sam: Agreed. Users hit a threshold where slightly off personalization feels worse than generic, like a bad mask versus none, so signaling limits or hitting high match matters. Beyond agents, it sparked self-reflection from seeing their own patterns, and could steady biased self-ratings with evidence as anchor. Overall, this provides a solid base: evidence-linked profiles that users refine for trustworthy stand-ins, with clear paths to broaden data and contexts.
Alex: Makes sense—prioritizing reviewable, context-smart reps over perfect copies. A meaningful step for AI that stands in without distorting. Thanks, Sam; that's a clear look at how to make these tools more personal yet reliable.
Sam: My pleasure, Alex. It underscores building with transparency and user control at the core.