Ruoxi Shang, Dan Marshall, Edward Cutrell, Denae Ford
10 min
Abstract
AI agents that communicate on behalf of individuals need to capture how each person actually communicates, yet current approaches either require costly per-person fine-tuning, produce generic outputs from shallow persona descriptions, or optimize preferences without modeling communication style. We present ASPECT (Automated Social Psychometric Evaluation of Communication Traits), a pipeline that directs LLMs to assess constructs from a validated communication scale against behavioral evidence from workplace data, without per-person training. In a case study with 20 participants (1,840 paired item ratings, 600 scenario evaluations), ASPECT-generated profiles achieved moderate alignment with self-assessments, and ASPECT-generated responses were preferred over generic and self-report baselines on aggregate, with substantial variation across individuals and scenarios. During the profile review phase, linked evidence helped participants identify mischaracterizations, recalibrate their own self-ratings, and negotiate context-appropriate representations. We discuss implications for building inspectable, individually scoped communication profiles that let individuals control how agents represent them at work.
Sam: Yes. They can flag wrong quotes or adjust scores, turning it into a back-and-forth where they negotiate what fits. After that review, the study tested it in fake work scenarios—like replying to a teammate about a deadline delay, customized with their real team names but same setup for everyone. Three AI replies per scenario: one generic with no personal info, one using their self-description, and one from the profile with evidence snippets.
Alex: Same structure across people, but personalized details so it feels real? How did the profiled one stack up?
Sam: In those tests with 20 participants, the replies using the profiled style got ranked highest for matching how they'd actually respond—better than self-reports or generic ones. Users rated them a clear step up on a 1-to-5 scale for sounding like themselves. The evidence links helped explain why, and while the AI sometimes leaned too positive, the relative patterns held steady.
Alex: Huh. So it's not perfect, but the transparency lets people fix the biases before using it.
Sam: Precisely. This auditing makes the profile usable in practice, shifting from black-box guesses to something collaborative.
Alex: Before that step, how close was the AI's initial guess to what people rated about themselves?
Sam: The study compared the AI's scores to participants' own ratings on 92 statements about communication habits, like whether they talk a lot or stay concise. Overall, exact matches happened in about one in four cases, with typical differences of roughly one point on a five-point scale. More importantly, the AI captured each person's relative pattern—the traits where they score high or low compared to their other traits—with a moderate rank match of about 0.39 on average.
Alex: A quarter exact, one-point average gap, and decent pattern matching. That suggests it's getting the shape right even if absolute numbers drift—maybe because some traits show up clearly in words, others less so?
Sam: Yes. It worked better for obvious traits signaled directly in text, like anger through sharp words or humor via jokes, with smaller gaps there. Subtler ones, depending more on hidden intent or context, had larger mismatches. Still, across six main categories of style—like how expressive or precise someone is—the relative shapes held up reasonably.
Alex: During auditing, did people mostly reject the AI, or did it change their minds sometimes?
Sam: Researchers reviewed comments from all 23 trait areas per person, sorting them into types like full agreement, rejecting the AI while sticking to self-view, or meeting in the middle. In about 17 percent of those reviews, people shifted their own ratings—either fully toward the AI after seeing the quotes, or compromising. This showed mismatches often came from people underestimating habits they forgot, or rating an ideal self instead of their actual workplace behavior.
Alex: Huh—so the evidence prompts real reflection, like realizing you structure messages more than you thought. That bidirectional tweak makes the profile more reliable for AI use.
Sam: Exactly. Auditing reveals self-biases and context nuances, turning initial estimates into collaborative, evidence-backed representations that preserve key patterns while allowing calibration. The paper notes this process strengthens the whole system for realistic scenarios.
Alex: Did those refined profiles actually make the AI's replies better in practice—like in simulated work situations?
Sam: Yes, the study tested that directly. They created ten workplace scenarios for each of the twenty people, like responding to a deadline slip or planning a team update, using real names from their work but the same basic setup. For each, the AI generated three replies: one plain and neutral, one based on the person's self-ratings alone, and one using the profiled style with its evidence links. On average across all six hundred judgments, the profiled replies ranked highest, coming first about 43 percent of the time—clearly ahead of the others.
Alex: Ranked highest overall, but results varied by person. What explained those differences—did some scenarios just suit generic replies better?
Sam: Individual preferences dominated the patterns. About half the participants strongly favored profiled replies, seeing them as more true to their usual way of organizing thoughts or balancing firmness with support. But four preferred generic ones, often because profiled versions added too much enthusiasm that felt off for casual coordination tasks, like quick schedule changes. Even self-reports won once, for those wanting a safer, low-key tone. The study clusters people this way, showing no one-size-fits-all.
Alex: So in structured spots—like defending an approach—the profile shone because it drew from real history. But what about the mismatches during auditing—did the paper pinpoint common reasons the AI got traits wrong at first?
Sam: They cataloged five main sources from participant feedback. First, data gaps: the logs miss off-record chats or rare events, so traits from those stay hidden. Second, mistaking situation for style—like reading a manager's structured talk as always dominant, when it's just the role. Third, LLMs literalize tone: sarcasm or jokes get scored as real aggression or commands. Fourth, fuzzy definitions: people and the scale disagree on what counts as argumentative versus helpful pushback. Fifth, evidence mishandling: over-weighting one example or mixing up who said what in meetings.
Alex: That typology makes the limits clear and fixable through review. Overall, it points to keeping users central, picking the right measurement tools, and grounding everything in actual behavior.
Sam: Precisely. The work stresses data for evidence, validated scales like CSI for structure, and human oversight to add missing context—building profiles that users trust for AI stand-ins. The study highlights how profiles respect boundaries, like distinguishing your everyday work style from other sides of yourself, since the data comes only from professional chats. Participants often wanted to dial back certain traits for specific situations, avoiding a full average that might expose too much.
Alex: Scoped like different modes for boss emails versus peer check-ins? But weren't there spots where the AI still got things systematically off, even after review?
Sam: Yes, the reviews uncovered predictable over-ratings, especially on preciseness—scoring people higher by about 1.7 points on average—likely because work messages are polished and structured by default, masking natural messiness. Traits hard to spot in text, like quiet thinking before speaking, also skewed low when evidence was thin. The paper notes real constraints: data from one organization and just text chats over 90 days, mostly tech workers okay with AI, so it may not stretch to other jobs or casual talks. Prompting the language model kept things simple and checkable but couldn't fully fix biases like occasional mix-ups in who said what. No long-term tests either—just initial sessions with made-up scenarios.
Alex: Fair points; it's a starting snapshot, not a full picture. Still, for quick audits in controlled spots like missing a meeting, those portable profiles seem ready to deploy without the creepy fine-tune feel.
Sam: Agreed. Users hit a threshold where slightly off personalization feels worse than generic, like a bad mask versus none, so signaling limits or hitting high match matters. Beyond agents, it sparked self-reflection from seeing their own patterns, and could steady biased self-ratings with evidence as anchor. Overall, this provides a solid base: evidence-linked profiles that users refine for trustworthy stand-ins, with clear paths to broaden data and contexts.
Alex: Makes sense—prioritizing reviewable, context-smart reps over perfect copies. A meaningful step for AI that stands in without distorting. Thanks, Sam; that's a clear look at how to make these tools more personal yet reliable.
Sam: My pleasure, Alex. It underscores building with transparency and user control at the core.