AI in UX Research: What Changes When the Researcher Is Also the Model
Ask a model to simulate a user, then ask the same model to analyze what that simulated user said, and you haven’t run a study. You’ve held a mirror up to a mirror and called the reflection data.Photo by Microsoft Copilot on UnsplashNearly half of UX researchers now name synthetic…
Ask a model to simulate a user, then ask the same model to analyze what that simulated user said, and you haven’t run a study. You’ve held a mirror up to a mirror and called the reflection data.Photo by Microsoft Copilot on UnsplashNearly half of UX researchers now name synthetic users — AI-generated stand-ins for real people — among the biggest trends shaping the field this year. That’s worth taking seriously. It’s also worth asking exactly what happens, epistemically, when the researcher and the model doing the researching become hard to tell apart.The Problem UX Research Has Not Fully Operationalized YetMost advice on “AI in UX research” treats AI as a faster typewriter—transcribing this, summarizing that, clustering these themes. That use is real and mostly fine.What that framing skips is the newer, riskier move: using AI not just to process research but to generate it—synthetic personas standing in for participants, agents that design and interpret their own studies. When the tool producing the “data” and the tool making sense of it share the same training biases, consistency stops being evidence of anything except that the model is being consistent with itself.This guidance is for researchers, UX leads, and product teams deciding how far to let AI into the research pipeline — not to reject the tool, but to know exactly which jobs it’s actually qualified for.Three Different Jobs, One Word“AI in research” hides at least three very different roles, and they carry three very different risk levels.There’s AI as an assistant—transcribing, coding, and clustering data a human already collected from real people. There’s AI as a proxy—a synthetic persona standing in for a participant who was never recruited. And increasingly there’s AI as researcher — an agent that designs a study, runs it, and draws its own conclusions with a human mostly watching from the sidelines.Conflating these three is the main source of trouble. A tool that’s genuinely useful in the first role gets quietly promoted into the second and third without anyone deciding that on purpose.Design implication: Before evaluating any AI research tool, ask which of the three jobs it’s actually doing—not which one its marketing page implies.The Assistant Role: Where AI Actually Earns Its KeepUsing AI to process data that real people already generated is the lowest-risk, highest-value use in the entire category.Computer-assisted qualitative analysis isn’t new — researchers have used software to code transcripts and cluster themes for decades. What’s new is that a language model can do it faster, in more flexible categories, without a rigid coding scheme set up in advance.The person still designed the study. The person still recruited real participants. The person still makes the interpretive call about what a theme means. AI is doing the mechanical middle, not the judgment at either end.AI in research. Source: ChatGPTDesign implication: This is the version of “AI in research” worth adopting without much hesitation—the human designs and interprets; the model accelerates the labor in between.The Proxy Problem: When a Model Plays the UserA language model predicting a user's words differs from a user’s nervous system responding.When you ask a model how a specific kind of user would feel about a checkout flow, it isn’t consulting a database of that person’s felt experience. It’s pattern-matching to whatever text about people like that appeared together in its training data—and that text skews in specific, measurable directions.A 2023 Stanford-led study, “Whose Opinions Do Language Models Reflect?" built a benchmark from a major Pew Research survey and found that models’ simulated opinions lined up far better with some demographic groups than others—closer to the more educated, more liberal-leaning, English-speaking end of the spectrum and further from everyone else. A more recent study that had a model simulate public opinion across six countries using the World Values Survey found dramatically higher accuracy for the U.S. than for countries like Japan or South Africa and a consistent skew toward matching male, white, older, and more highly educated respondents over other groups. A separate study modeling climate-change opinions found the model systematically underestimated how worried Black American respondents actually were, compared to real survey data.None of this makes synthetic personas useless. It makes them a specific kind of instrument, with a specific and now well-documented blind spot.From: Performance and biases of Large Language Models in public opinion simulationDesign implication: treat a synthetic user’s response as a hypothesis about what a real person might say—never as a finding about what one did.Ask a model what a user would feel, and it will answer — fluently, confidently, from a population that doesn’t exist.The Circularity Trap: Same Model, Two JobsThe sharpest risk isn’t using AI to simulate users or using AI to analyze data. It allows the same model to perform both tasks.If a model generates the synthetic interview responses and codes them thematically and draws the conclusion, there’s no independent signal left to catch a blind spot with. Any bias baked into that model shows up identically on both sides of the “study”—the input and the interpretation were never actually separate.It’s the equivalent of calibrating a scale using itself. The number will come back perfectly consistent, every time, and there’s no way to know from inside the loop whether it’s also correctly zeroed.Qualitative researchers have a name for a milder version of this problem—reflexivity, the risk that a human researcher’s own assumptions shape what they notice in the data. The human version at least comes with the possibility of self-awareness catching it mid-study. A closed model loop doesn’t notice anything. It just produces the next plausible sentence.Design implication: If the same model (or model family) is generating your “data” and interpreting it, that’s not a study—it's one system talking to itself. Include an independent check, ideally by a real person, before the finding leaves the room.A hall of mirrors is still just one mirror. However many times the image bounces, it never becomes a window.Agentic AI as “The Researcher”The newest and least tested version of this problem is agentic AI running the whole pipeline — design, execution, and interpretation — with a human mostly watching.Gartner predicts that up to 40% of enterprise applications will include task-specific AI agents by 2026, up from less than 5% in 2025. The appeal is obvious: an agent that plans a study, runs it across simulated or real participants, and hands back conclusions removes a significant amount of manual labor.The risk compounds rather than adds to it. Now the same system that decided how to run the study also gets to decide what counts as a meaningful pattern in the results—with whatever the model’s training predisposes it to notice, amplified rather than caught.Current best practice among teams actually using synthetic and agentic tools well is notably modest: use them for roughly the first 80% of exploratory work, label every output as AI-generated, and validate anything high-stakes with real people before it ships.Design implication: agentic AI can own the mechanical scale of a study. It should not also own the judgment call about what the study means—that's the one part still worth a human bottleneck.What Doesn’t Change: The Nervous System Still Belongs to a PersonNo model has a thumb that becomes tired, an amygdala that fires in fifty milliseconds, or a working-memory limit of three or four items.Everything this kind of research actually measures—a longer eye-tracking fixation, a frontal-alpha shift toward avoidance, and a hesitation that shows up in reaction time but never in what someone says out loud—depends on a real nervous system producing a real, involuntary signal. A language model can describe that fatigue or that hesitation convincingly because it’s seen thousands of descriptions of it. It cannot generate the constraint underneath the description, because there’s no body attached to the text.That’s not a knock on the technology. It’s a reminder of what the technology fundamentally is: a very effective predictor of plausible language, not a nervous system standing in for one.Design implication: anywhere your research question depends on an actual felt, physiological, or edge-case human response, a synthetic stand-in isn’t a shortcut—it's a different question wearing the same clothes as the one you meant to ask.Where This Gets PersonalThe whole reason the “polite user problem” I’ve written about before matters is that it’s proof a really nervous system can say one thing and show another—a participant telling me a screen is “fine” while their frontal alpha asymmetry says otherwise. That gap is precisely what a synthetic user has no way to produce. There’s no nervous system underneath the text for the truth to leak out of.A recent campus project made the point more concretely than I expected. We were building a reward system to nudge hostel students toward waste segregation, and the easy path was to let a model imagine a persona—a plausible-sounding “eco-conscious student” with plausible-sounding habits. Instead, we ran a real, if small, survey first: 14 hostel students, all first-years, all under six months on campus. The persona that came out the other side—Akarsh, 20, tech-savvy, a heavy Blinkit and Zepto user who spends enough at the tuck shop and laundry counter that an eco-coin would actually matter to him—isn't just a name attached to a paragraph that an LLM generated from general priors about college students. He’s a synthesis of real answers: nine of fourteen students discard six or more reusable items a week; ten of fourteen say unused items pile up in their room “fairly often” or “all the time"; and no single dominant excuse— “no time,” “no awareness,” and “no motivation or benefit”—came back almost perfectly tied.That tie is the detail an imagined persona would have been unlikely to invent, and it’s the detail that actually shapes the design: it’s why a rewards mechanic, not another awareness poster, was the right call. It’s also worth saying plainly that fourteen responses from one cohort is a pilot, not a finding—Akarsh is a well-grounded hypothesis about who we’re designing for, not a substitute for testing the eco-coin system on real students once it exists.I show students a synthetic persona’s answer and a real participant’s answer, unlabeled. They’re usually less confident than they expect to be — which is itself the point.Persona profile: Akarsh, a 20-year-old student, alongside survey data demonstrating the specific, real-world barriers that informed his profile.The AI-in-Research Litmus TestBefore you trust the output, run any AI-assisted research plan through these checks:Name the job. Is this AI acting as an assistant, proxy, or researcher—and did the team choose that deliberately?Trace the nervous system. Is a real person’s behavior or physiology producing this data, or is a model predicting what one might produce?Check for circularity. Is the same model (or family) generating the data and interpreting it? If yes, what independent verification exists before the finding ships?Treat synthetic output as a hypothesis, not a finding. Would this conclusion survive being run past a real, demographically matched participant?Keep a human in the judgment call. Is AI scaling a mechanical task or quietly making the interpretive decision a person should be making?The TakeawayThree things worth remembering:First, “AI in research” hides three different jobs—assistant, proxy, and researcher—and each one carries a different level of risk.Second, the sharpest danger isn’t AI simulating a user or AI analyzing data. It’s the same model doing both, with no independent signal left to catch what it missed.Third, no model has a nervous system. It can describe fatigue, hesitation, and fifty-millisecond judgments convincingly—it cannot generate the constraint those descriptions come from.Where has AI genuinely earned a place in your research process, and where have you caught it quietly standing in for a judgment call that should have stayed yours? Tell me in the comments.This story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!AI in UX Research: What Changes When the Researcher Is Also the Model was originally published in Generative AI on Medium, where people are continuing the conversation by highlighting and responding to this story.Source: Generative AI Pub — Published — Category: Image AI