Abstract
Large language models (LLMs) are now widely used in applications that depend on closed-ended decisions, including automated surveys, policy screening, and decision-support tools. In such contexts, these models are typically expected to produce consistent binary or ternary responses (for example, Yes, No, or Neither) when presented with questions that are semantically equivalent. However, recent studies show that LLM outputs can be influenced by relatively minor changes in prompt wording, raising concerns about the reliability of their decisions under paraphrasing. In this paper, we conduct a systematic analysis of paraphrase robustness across five widely used LLMs. To support this evaluation, we develop a controlled dataset consisting of 200 opinion-based questions drawn from multiple domains, each accompanied by five human-validated paraphrases. All models are evaluated under deterministic inference settings and constrained to a fixed Yes/No/Neither response format. We assess model behavior using a set of complementary metrics that capture the stability of each evaluated model. DeepSeek Reasoner and Gemini 2.0 Flash show the highest stability when responding to paraphrased inputs, whereas Claude 3.7 Sonnet exhibits strong internal consistency but produces judgments that differ more frequently from those of other models. By contrast, GPT-3.5 Turbo and LLaMA 3 70B display greater sensitivity to surface-level variations in prompt phrasing. Overall, these findings suggest that robustness to paraphrasing is driven more by alignment strategies and reasoning design choices than by model size alone.
| Original language | English |
|---|---|
| Title of host publication | WASSA 2026 - 15th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, Proceedings of the Workshop |
| Publisher | Association for Computational Linguistics (ACL) |
| Pages | 52-59 |
| Number of pages | 8 |
| ISBN (Electronic) | 9798891763784 |
| DOIs | |
| State | Published - 2026 |
| Event | 15th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, WASSA 2026 - Rabat, Morocco Duration: 29 Mar 2026 → … |
Publication series
| Name | WASSA 2026 - 15th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, Proceedings of the Workshop |
|---|
Conference
| Conference | 15th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, WASSA 2026 |
|---|---|
| Country/Territory | Morocco |
| City | Rabat |
| Period | 29/03/26 → … |
Bibliographical note
Publisher Copyright:© 2026 Association for Computational Linguistics.
ASJC Scopus subject areas
- Language and Linguistics
- Linguistics and Language
Fingerprint
Dive into the research topics of 'Measuring LLMs’ Sensitivity to Paraphrased Opinion Prompt'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver