# ssrn-persona-prompting-ai-accuracy-2025-12-07

## Veille

Wharton study (Generative AI Labs): expert personas don't improve LLM factual accuracy - GPQA Diamond and MMLU-Pro benchmarks - SSRN

## Titre Article

Playing Pretend: Expert Personas Don't Improve Factual Accuracy

## Date

2025-12-07

## URL

https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5879722

## Keywords

AI prompting, personas, LLM accuracy, AI benchmarking, GPQA Diamond, MMLU-Pro, AI performance, prompt engineering, expert personas, AI evaluation

## Authors

Savir Basil, Ina Shapiro, Dan Shapiro, Ethan Mollick, Lilach Mollick, Lennart Meincke (Generative AI Labs, The Wharton School, University of Pennsylvania)

## Ton

Profile: academic research paper (SSRN working paper, Wharton's Generative AI Labs), scientific and empirical register, high technical level (statistics, benchmarks).
Style: rigorous experimental approach — detailed protocol (25 independent responses per question and per model-prompt pair), 95% confidence intervals, results reported by benchmark and by model. Measured and cautious tone, typical of "null results": the authors refute a widespread practice without over-generalizing and flag their limitations. Authority grounded in the Wharton affiliation and Ethan Mollick's reputation. Target audience: prompt engineering practitioners, researchers, enterprise AI teams.

## Pense-betes

- Question tested: does assigning an expert persona to an LLM improve its accuracy on difficult factual questions? Answer: generally no
- 6 models tested: GPT-4o, GPT-4o-mini, o3-mini, o4-mini, Gemini 2.0 Flash, Gemini 2.5 Flash
- 2 benchmarks: GPQA Diamond (198 doctoral-level questions) and MMLU-Pro (300 professional-level questions)
- Protocol: baseline (no persona), expert personas (physics, mathematics, economics, biology, chemistry, engineering, law, history), "low-knowledge" personas (Layperson, Young Child, Toddler); 25 independent responses per question and per model-prompt pair
- Main result: most persona conditions are statistically indistinguishable from the baseline
- Low-knowledge personas often degrade accuracy; the "Toddler" persona is harmful in 4/6 models
- Exception: Gemini 2.0 Flash shows modest improvements with all 5 expert personas on MMLU-Pro (notably Engineering and Chemistry)
- Aligning the expert persona with the question's domain provides no consistent improvement
- Observed failure mode: Gemini Flash models sometimes refuse to answer with an out-of-domain expert persona; over-specialization leads to underuse of actual knowledge
- Practical implication: favor task-specific instructions over personas; personas remain useful for tone/style, not for accuracy
- Limitations: subset of models, academic benchmarks, limited number of personas, no interaction with other prompting techniques tested

## RésuméDe400mots

This study from Wharton's Generative AI Labs examines whether assigning expert personas to AI models improves their performance on difficult objective multiple-choice questions. The researchers tested six models (GPT-4o, GPT-4o-mini, o3-mini, o4-mini, Gemini 2.0 Flash, Gemini 2.5 Flash) on two demanding benchmarks: GPQA Diamond (198 doctoral-level questions) and MMLU-Pro (300 professional-level questions).

The protocol compares three conditions: a baseline with no persona, expert personas (expert in physics, mathematics, economics, biology, chemistry, engineering, law, history), and "low-knowledge" personas (Layperson, Young Child, Toddler — "a 4-year-old who believes the moon is made of cheese"). Each model-prompt pair is evaluated over 25 independent responses per question (4,950 runs per pair on GPQA, 7,500 on MMLU-Pro), with 95% confidence intervals.

The results are essentially null: most persona conditions produce performance statistically indistinguishable from the baseline. On GPQA Diamond, no expert or low-knowledge persona reliably improves performance; the sole exception is a small gain from the "Young Child" prompt on Gemini 2.5 Flash (RD = 0.098). On MMLU-Pro, no expert persona delivers a statistically significant improvement for 5 of the 6 models, and nine significant negative differences are observed. Low-knowledge personas often degrade accuracy: the "Toddler" persona reduces performance in 4 of 6 models and proves significantly worse than "Layperson" in 5 of 6 models.

The notable exception is Gemini 2.0 Flash, which shows modest positive differences with all five expert personas on MMLU-Pro, particularly in engineering and chemistry. Additionally, aligning the expert persona with the question's domain provides no consistent benefit. The researchers identify failure modes: the Gemini Flash models sometimes refuse to answer when assigned an out-of-domain expert persona, and overly narrow role instructions lead the models to underuse their actual knowledge.

The practical implications are significant: the widespread practice of persona prompting is likely ineffective for improving factual accuracy. Organizations will derive more value from task-specific instructions, and should test multiple prompt variants for their concrete problems. Personas may nonetheless retain other uses, such as modulating tone or presentation style. The study's limitations (a limited number of models and personas, academic benchmarks) open avenues for future research.

## GrapheDeConnaissance

- Generative AI Labs —publie→ étude personas et précision IA (DOCUMENT, 0.98)
- Ethan Mollick —publie→ étude personas et précision IA (DOCUMENT, 0.95)
- Lilach Mollick —publie→ étude personas et précision IA (DOCUMENT, 0.95)
- étude personas et précision IA —affirme_que→ les personas experts n'améliorent pas la précision factuelle des LLM (AFFIRMATION, 0.95)
- personas faible connaissance —réduit→ performance des LLM (CONCEPT, 0.92)
- Gemini 2.0 Flash —mesure→ amélioration modeste avec personas experts (MMLU-Pro) (MESURE, 0.85)
- persona Toddler —réduit→ performance dans 4/6 modèles (CONCEPT, 0.9)
- étude personas et précision IA —utilise→ GPQA Diamond (TECHNOLOGIE, 0.95)
- étude personas et précision IA —utilise→ MMLU-Pro (TECHNOLOGIE, 0.95)
- étude personas et précision IA —affirme_que→ la correspondance domaine-persona n'améliore pas la performance de manière consistante (AFFIRMATION, 0.88)
- refus de répondre avec personas hors domaine —observé_dans→ modèles Gemini Flash (TECHNOLOGIE, 0.82)

---
Canonical: https://www.thekb.eu/en/fiches/ssrn-persona-prompting-ai-accuracy-2025-12-07/
