Talk:Synthetic data
Add topic| This article is rated C-class on Wikipedia's content assessment scale. It is of interest to the following WikiProjects: | |||||||||||
| |||||||||||
| This article is based on material taken from the Free On-line Dictionary of Computing prior to 1 November 2008 and incorporated under the "relicensing" terms of the GFDL, version 1.3 or later. |
COI edit request: limitations of LLM-generated synthetic survey samples
[edit]| This edit request by an editor with a conflict of interest has now been answered. Set |answered=no to reactivate the request if necessary. |
I have a disclosed conflict of interest because I am working on behalf of Verasight. I am requesting editor review rather than editing the article directly.
Proposed addition to the “Scientific research” section:
In survey research, large-language-model-generated synthetic responses may approximate some aggregate results while failing to reproduce subgroup variation or relationships between variables. The American Association for Public Opinion Research recommends distinguishing synthetic responses from human responses and validating them against human benchmarks.[1] An experimental comparison by Morris, Leff, and Enns found that adding administrative and attitudinal information did not consistently improve synthetic survey estimates.[2]
I welcome independent editors’ judgment regarding wording, placement, weight, and source suitability. SurveyDataNotes (talk) 16:09, 13 July 2026 (UTC)
- Your company's reports do not meet WP:RS, not on this article, not on other articles you have proposed them, probably not on any article. - MrOllie (talk) 16:12, 13 July 2026 (UTC)
References
- ↑ Responsible AI Integration in Survey Research (PDF) (Report). American Association for Public Opinion Research. 2026.
- ↑ Morris, G. Elliott; Leff, Benjamin; Enns, Peter K. (September 26, 2025). "The Limits of Synthetic Samples in Survey Research". Verasight.
Edit request: research on synthetic survey respondents
[edit]| This edit request by an editor with a conflict of interest has now been answered. Set |answered=no to reactivate the request if necessary. |
Disclosure: I am working on behalf of Verasight, which published the source below.
Please consider adding the following to the “Scientific research” section:
“Large language models have also been tested as synthetic survey respondents. In a 2026 company-conducted study comparing AI-generated and human survey responses across 16 questions, the synthetic responses had a mean absolute error of 6.6 percentage points at the national topline, with higher errors among demographic and partisan subgroups.[1]”
The wording identifies the research as company-conducted and limits the claim to the reported experiment. I am requesting independent editorial review rather than editing the article directly. SurveyDataNotes (talk) 19:24, 1 September 2026 (UTC) SurveyDataNotes (talk) 19:24, 1 September 2026 (UTC)
Not done as in other recent requests. WeyerStudentOfAgrippa (talk) 21:26, 6 September 2026 (UTC)
References
- ↑ Morris, G. Elliott; Leff, Benjamin; Enns, Peter K. (2 June 2026). "Can AI "digital twins" replace human respondents?". Verasight. Retrieved 1 September 2026.