James Brand, Ayelet Israeli, Donald Ngwe
4 min
Companies often struggle to efficiently filter early-stage product ideas, as traditional human-centric market research is frequently slow and costly. This paper explores the potential of using Large Language Models (LLMs) as synthetic customers to simulate market responses. By presenting structured product configurations to these models, researchers can estimate willingness-to-pay (WTP) and compare product alternatives before ever engaging a human respondent.
The authors utilized a programmatic approach to prompt LLMs with choice-based questions, mimicking traditional conjoint analysis. Across various product categories like toothpaste and consumer electronics, the researchers found that LLMs could generate preference rankings that aligned with human-derived results. This capability allows firms to broaden the top of their innovation funnel, enabling the rapid testing of dozens or even hundreds of concepts at a fraction of the time and cost required for traditional studies.
While "off-the-shelf" LLMs provide a useful baseline, they are prone to overestimating interest in novel or unusual features. The authors demonstrate that fine-tuning models with a company’s proprietary historical survey data significantly improves accuracy. This process helps the model learn specific consumer trade-offs, preventing the exaggerated enthusiasm often seen in base models. However, this advantage is domain-specific; fine-tuning a model on one product category does not necessarily improve its performance in a different, even if related, category.
Synthetic customers should be viewed as a tool for augmentation rather than a replacement for human research. LLMs currently lack the emotional intelligence and demographic nuance required for final validation or deep segmentation. Instead, they are best deployed as a high-speed filter to prioritize promising directions. Firms that successfully integrate these synthetic insights with traditional human research can achieve a significant competitive advantage by innovating faster and reducing the risk of costly product failures.
Alex: [thoughtful] Then you wouldn't use these for final validation or segment-level targeting. Where do they fit in a product team's workflow? The authors talk about an innovation funnel. [[RP_SECTION:innovation-funnel-application|Innovation Funnel Application]]
Sam: [building momentum] That's where they see the value. Traditional conjoint analysis is slow and expensive, so you test only a handful of configurations. Synthetic agents let you screen far more variations in hours rather than weeks. The team uses them to filter out obvious failures and shortlist the few concepts worth the cost of a human study.
Alex: [checking understanding] So the cost asymmetry does the work. A false positive at the top of the funnel is cheap, because a real study still stands behind it.
Sam: [precise] Yes, and it lets you explore ideas you'd never have risked a formal survey on. The risk is that you automate your own biases. If your historical surveys were flawed, the synthetic agents will reproduce that flaw at scale.
Alex: [reflective] Could these models adapt to dynamic market shifts without manual retraining?
Sam: [sitting back, broader perspective] Not currently. The models are static, with no real-time access to shifting trends unless you update the training data yourself. The authors are clear that this is an augmentation tool. A human researcher still has to interpret the synthetic signal, account for its limitations, and validate with real people.
Alex: [concluding] So it's an efficiency tool, but only if you already have the data infrastructure to support it. [[RP_SECTION:data-infrastructure-requirements|Data Infrastructure Requirements]]
Sam: [quiet conviction] That's the takeaway. The advantage doesn't come from the language model, which everyone can access. It comes from the quality of the proprietary historical data you fine-tune on. Without it, you get a generic, potentially biased average. With it, you get a usable lens for early-stage exploration, within the limits we've covered.
Alex: If you want the figures and the method choices we skipped, you can generate a deep dive of this paper. The paper has the rest either way.
Sam: Thanks for listening.