Can Large Language Models Become Reliable Silicon Respondents?
A Fine-tuning Experiment Based on CFPS Response Distributions
Lin Zeteng1Ye Zhenye2 Wang Renhe2*
(1. Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, Guangdong, China;
2. School of Politics and Public Administration, South China Normal University, Guangzhou 510006, Guangdong, China)
Abstract: With the expanding use of generative artificial intelligence in social science research,whether large language models can serve as reliable "silicon respondents" has become an important methodological question for computational social science and survey research. Based on individual-level panel data from the China Family Panel Studies (CFPS) from 2010 to 2022, this study examineswhether large language models can predict future response distributions for two subjective variables:life satisfaction and self-rated health. Methodologically, we fine-tune large language models using supervised fine-tuning with Low-Rank Adaptation under a Kullback-Leibler divergence constraint(FT-KL), and compare this approach with three zero-shot prompting strategies: question-answer prompting, first-person biographical prompting, and third-person portrayal prompting. The resultsshow that FT-KL consistently outperforms all prompting-based baselines across different model sizes,task settings, and subjective variables. Models calibrated with real survey response distributionsgenerate predictions that are closer to actual CFPS distributions in terms of JSD, KL, L1, and EMD.These findings suggest that large language models do not naturally become reliable surveyrespondents through demographic role prompting alone. However, when constrained by real surveydistributions, they can generate statistically more realistic silicon samples within clearly defined taskboundaries. This study provides empirical evidence for using calibrated large language models as anauxiliary tool for questionnaire pretesting, policy pilot evaluation, and evidence-informed decision-making.Distribution; Computational Social Science
Keywords: Large Language Models; Silicon Respondents; China Family Panel Studies; Response
[原文下載][Download PDF]
WeChat