sözaltı news Journal
Journal
EN AZ
AI agents aren’t ready to replace humans in behavioral research

AI agents aren’t ready to replace humans in behavioral research

sciencenews.org 03.09.2026 21:30 1 views
A new study finds that digital twins don’t yet replicate the views of the individuals they are modeled after.

Replacing human subjects with AI surrogates, or digital twins, is on some social scientists’ wish lists. Human subjects are expensive, tire easily and can suffer psychological distress from some studies. But those wishes may take more time to be realized, a study appearing September 2 in Science Advances suggests.

AI twins designed to mimic a given individual’s behavior instead seem to distort their surrogate’s views, creating a “funhouse mirror” effect, the researchers note. Those respondents answered 500-plus questions about characteristics including age, ethnicity, income, education, religious practices, political preferences, personality traits, spending habits, mathematical abilities and vocabulary skills. Many of the questions come from scales that are used commonly in psychology, economic and business research.

Respondents also completed various online tests designed to assess thought patterns and biases. Because research testing the value of digital twins remains limited, the goal of that project was to create an open-source dataset for others to use, Toubia says. Across 19 social science experiments, the team evaluated everything from how individuals and their twins respond to people who donate to both Republican and Democratic party candidates to what they “think” about algorithmic hiring.

These digital twins performed better than chance, the team found, but they were wrong on average about a quarter of the time. They performed roughly on par with chatbots that received demographic information alone. The digital twins did, however, better capture real variation in people’s responses compared with the LLMs that knew only demographic info.

For example, one person may rate themselves as a 2 on a scale of self-control while another rates themselves as a 4. The LLM with more limited info might report a 3 for each person, washing out any differences. But the digital twin might report a 3 and 5.

Though still wrong, the predictions give a better sense of potential differences across the group. Toubia and colleagues attribute the digital twins’ poor performance overall to several key distortions: Twins’ responses tended to be more homogenous than people’s responses and often skewed to demographic stereotypes. And their accuracy increased with more affluent and educated participants.

Extract — continue reading at the source.

Read full story