The probabilistic encoder: methodological considerations – On the validation of data produced by LLMs
Title: The probabilistic encoder: Methodological reflections on the validation of data produced by LLMs
Guest : Guillaume Sauvé (Université du Québec à Montréal and Université de Montréal)
Guest profile: His research lies at the intersection of comparative politics, political sociology and intellectual history, with a particular focus on Russia during the late Soviet and post-Soviet eras. Drawing on archival research and interviews, he also incorporates methods from computational social science.
Discussant: Jérémy Robine (University of Paris 8)
Abstract: Large language models are increasingly being used to process large corpora in the social sciences. However, their probabilistic nature poses a problem: the same text and the same query can produce different outputs. How can we ensure the reliability and validity of the results? Drawing on a project involving the semi-automated analysis of petitions in the Russian press, I propose to consider the LLM as a probabilistic coder performing interpretative tasks. I will discuss the limitations of validation based primarily on a human reference coder and present an alternative approach combining repeated coding, measurement of coding stability, targeted human audits and the explicit retention of ambiguous cases. The challenge is to treat the model’s uncertainty as information to be analysed, rather than simply as a flaw.
Date and times : Thursday 24 September, 10.00–12.00
Venue : Condorcet Campus – North Research Building (14, cr des Humanités, 93300 Aubervilliers) – Room 0.009