Talk:Synthetic data
Add topic| This article is rated C-class on Wikipedia's content assessment scale. It is of interest to the following WikiProjects: | |||||||||||
| |||||||||||
| This article is based on material taken from the Free On-line Dictionary of Computing prior to 1 November 2008 and incorporated under the "relicensing" terms of the GFDL, version 1.3 or later. |
COI edit request: limitations of LLM-generated synthetic survey samples
[edit]| This edit request by an editor with a conflict of interest has now been answered. |
I have a disclosed conflict of interest because I am working on behalf of Verasight. I am requesting editor review rather than editing the article directly.
Proposed addition to the “Scientific research” section:
In survey research, large-language-model-generated synthetic responses may approximate some aggregate results while failing to reproduce subgroup variation or relationships between variables. The American Association for Public Opinion Research recommends distinguishing synthetic responses from human responses and validating them against human benchmarks.[1] An experimental comparison by Morris, Leff, and Enns found that adding administrative and attitudinal information did not consistently improve synthetic survey estimates.[2]
I welcome independent editors’ judgment regarding wording, placement, weight, and source suitability. SurveyDataNotes (talk) 16:09, 13 July 2026 (UTC)
- Your company's reports do not meet WP:RS, not on this article, not on other articles you have proposed them, probably not on any article. - MrOllie (talk) 16:12, 13 July 2026 (UTC)
- ↑ Responsible AI Integration in Survey Research (PDF) (Report). American Association for Public Opinion Research. 2026.
- ↑ Morris, G. Elliott; Leff, Benjamin; Enns, Peter K. (September 26, 2025). "The Limits of Synthetic Samples in Survey Research". Verasight.