Publication

  • Cecilia Y. Sui
    Forthcoming at PSRM

    Political science research increasingly uses text classification to gauge subtle concepts like toxicity or anger. Traditional methods treat words in isolation, missing the contextual dynamics where meaning resides. Recent transformer-based language models overcome such limitations, but remain underutilized in published articles especially among applied scholars, partly due to misconceptions about the underlying mechanisms and the data and computational requirements. To bridge this gap, this article offers an accessible explication of the pretrain-finetune paradigm focusing on the mechanisms behind transformer-based language models and its potential to improve text analysis in political science by harnessing robust models trained on extensive datasets while requiring relatively small amounts of labeled data and computational resources from researchers. I employ the approach to identify toxic language in conversation threads following U.S. Senators' tweets and provide an accessible online tutorial to support nontechnical scholars in adopting this methodology.

Working Papers

  • Should Candidates Campaign Like Influencers? Evidence From 16,572 Instagram Reels and Two Generative AI Video Experiments
    Cecilia Y. Sui
    Draft available upon request
    Job Market Paper

    Political candidates are increasingly urged to communicate like social media influencers on the theory that platform-native intimacy builds parasocial bonds and support. However, evidence that candidates follow the advice or benefit from it remains thin. Voters judge candidates as prospective officeholders and informality that signals authenticity in influencers may signal incompetence in candidates. I test these claims using 16,572 Instagram Reels from 200 accounts of state legislative candidates and two pre-registered experiments using AI-generated audiovisual stimuli that vary visual style and relational script while holding a candidate's face, voice, and message fixed. Candidates adopt the style selectively, with widespread use of engagement-predicting informal visuals and rare use of relational cues. The experiments show that neither dimension increases perceived closeness while both impose competence costs. Platforms thus reward what voters penalize. More broadly, the study contributes to an enduring debate of political communication and self-presentation, the need to appear fit to govern while remaining relatable to the governed, and introduces a reusable design for audiovisual experiments.

  • One Language is Enough: Using Transfer Learning to Detect Populist Rhetoric Across 25 Languages Without Translation
    Cecilia Y. Sui, Soyeon Jeon, Christopher Lucas, Jacob Montgomery, and Margit Tavits
    Draft available upon request
    Under Review

    Social scientists increasingly analyze text data across multiple languages, yet current methods typically rely on costly translation or require labeled training data for multiple languages. We demonstrate how multilingual language models enable cross-lingual text analysis without translation. Using a new dataset of populist rhetoric labeled in 25 languages, we find that fine-tuning a multilingual model on just one language matches the performance of models trained on 25 languages, eliminating the need for translation or language-specific training data. We validate this approach through out-of-sample testing, comparison to expert surveys, and an analysis of campaign rhetoric in Switzerland. Our results provide evidence that multilingual language models offer a practical, efficient solution for global comparative research, dramatically reducing the data requirements for cross-lingual political text analysis while still producing accurate, valid measures.

Work in Progress

  • Seeing Inside the Black Box: Feature Attribution for Political Text in the Age of Language Models
    Cecilia Y. Sui