# mollick-giving-ai-job-interview-2025-11-12

## Veille

AI benchmarking beyond standard tests - Interviewing AI models for specific use cases - Jagged Frontier - OpenAI GDPval - Vibes vs real measurements - GuacaDrone example - Ethan Mollick - One Useful Thing

## Titre Article

Giving your AI a Job Interview

## Date

2025-11-12

## URL

https://www.oneusefulthing.org/p/giving-your-ai-a-job-interview

## Keywords

AI benchmarking, MMLU-Pro, ARC-AGI, METR Long Tasks, benchmarks limitations, vibes-based testing, OpenAI GDPval, real-world tasks, Jagged Frontier, model evaluation, GuacaDrone, risk assessment, otter on plane, pelican on bike, Simon Willison, Claude 4.5 Sonnet, GPT-5 Thinking, Gemini 2.5 Pro, Kimi K2 Thinking, Grok, Microsoft Copilot, model personality, advice at scale, job interview analogy, expert evaluation, model selection, organizational AI adoption, task-specific performance, judgment differences

## Authors

Ethan Mollick

## Ton

**Profile:** Practitioner-educator | Pedagogical first person | Analytical-prescriptive | Accessible-intermediate

Mollick adopts an educator's voice blending analytical rigor with pragmatic accessibility. The progressive pedagogical structure (benchmark problem → "vibes" solution → real-world solution → prescriptive action) is typical of One Useful Thing. Playful examples (otters on planes, GuacaDrone, pelican on a bike) make abstract concepts tangible without sacrificing depth. The balanced tone acknowledges limitations ("some problems", "not easy") while offering actionable solutions ("you are going to need to interview your AI"). Citations of empirical data (Epoch AI charts, the GDPval paper) lend credibility. Business analogies (hiring a VP, job interview) connect AI adoption to familiar management practices. Typical of Mollick's style (Wharton professor and accessible blogger) demystifying technical complexity for a practitioner audience.

## Pense-betes

- **Benchmarking paradox**: it's "surprisingly difficult to measure exactly how 'intelligent' they are"
- **Problems with standard benchmarks**:
- Public answer keys → incorporated into training (accidentally or to inflate scores)
- What the tests actually measure is unknown (MMLU-Pro: "average cranial capacity of Homo erectus?", "Cheap Trick's 1979 live album?")
- Tests are not calibrated: is going from 84% to 85% as difficult as going from 40% to 41%?
- Maximum score unattainable (errors in the questions, unusual reporting)
- **Collective value**: all benchmarks trend "up and to the right" (AIME, GPQA, MMLU, SWE-bench, LiveBench, Terminal-Bench, ARC-AGI, METR Long Tasks)
- **Underlying capability factor**: correlated with real-world impact, from medicine to finance
- **Benchmark gap**: focused on math, science, reasoning, code — not on writing, sociological analysis, business advice, empathy
- **Key quote**: "What you actually care about is which model would be best for YOUR needs" **"Vibes-based" benchmarking**
- **Simon Willison**: pelican-on-a-bicycle test
- **Mollick's tests**: otter on a plane, spaceship control panel in JavaScript, difficult poem, video games/shaders, article analysis, writing about time travel
- **Questions asked**: Does it make mistakes? Are answers similar to other models? Recurring themes/biases?
- **Writing exercise**: "Someone is rationing out remaining words like wartime supplies, and you're told you have 10,000 left for life. At 47 words, you're holding a newborn."
- **Result patterns**:
- Claude 4.5 Sonnet: solid writing model
- Gemini 2.5 Pro: doesn't keep track of word count (currently the weakest)
- GPT-5 Thinking: exuberant stylist, complex metaphors, sometimes incoherent
- Kimi K2 Thinking: interesting turns of phrase, but the story makes no sense
- **Limits of vibes**: idiosyncratic, different answers each time, better prompts = better results, relying on impressions rather than measurements **Real-world benchmarking: OpenAI GDPval**
- **Step 1**: experts (14 years of experience on average: finance, law, retail) create realistic, complex projects representing 4 to 7 hours of human work
- **Step 2**: several AI models + human experts (paid hourly) perform the tasks
- **Step 3**: a third group of experts grades the results blind (AI vs. human unknown), spending more than an hour per question
- **Strengths**: the best models beat humans in software development and personal financial advice
- **Weaknesses**: pharmacists, industrial engineers, and real estate agents easily beat AI
- **Differences between models**: ChatGPT a better sales manager, Claude a better financial advisor
- **Reveals the shape of the Jagged Frontier** and how it evolves over time **The GuacaDrone experiment**
- **Pitch**: a dubious idea — guacamole delivery by drone
- **Method**: each AI model rates the viability from 1 to 10, ten times each (answers vary each time)
- **Results**:
- Grok: great idea (enthusiastic)
- Microsoft Copilot: excited
- GPT-5 and Claude 4.5: more skeptical
- Mollick personally: would rate it 2 or lower
- **Implication**: "Systematically rating ideas 3-4 points higher or lower means systematically steering you in a different direction"
- **Business impact**: some want an AI that embraces risk, others one that avoids it — understanding how AI "thinks" about critical questions is essential **The "interview your model" prescription**
- **Individuals**: vibes are enough, do the otter test (now become too easy — upgraded to "1960s documentary footage of the last famous concert before the otter-pack incident" in Sora 2)
- **Large-scale organizations**: a different challenge
- "Best" isn't enough for thousands of tasks and hundreds of employees
- You need to know specifically what YOUR AI is good at, not on average
- GDPval revealed this: the performance of top models varies significantly by task
- GuacaDrone: judgment on ambiguous questions, systematically different advice
- Differences compound at scale
- **Do not rely on**: either vibes or general benchmarks
- **To do**:
- Systematically test the AI on the real work it will do and the real judgments it will make
- Create realistic scenarios reflecting actual use cases
- Run multiple times to see the patterns
- Have the results evaluated by experts
- Compare models head-to-head on the tasks that matter
- "The difference between 'this model scored 85% on MMLU' and 'this model is more accurate on our financial analysis tasks but more conservative in risk assessment'"
- Several times a year, as new models are released
- **Final analogy**: "You wouldn't hire a VP based solely on their SAT scores. You shouldn't choose an AI that will advise thousands of decisions based on the fact that it knows the average cranial capacity of Homo erectus is a bit under 1,000 cubic centimeters."

## RésuméDe400mots

Ethan Mollick argues that despite measurable progress in AI, standard benchmarks fail to capture what actually matters: performance on YOUR specific tasks with YOUR judgment criteria. He prescribes "interviewing" AI models like job candidates rather than relying on generic test scores.

**Problems with standard benchmarks**

Benchmarks like MMLU-Pro ask obscure questions ("average cranial capacity of Homo erectus?", "title of Cheap Trick's 1979 live album?") whose real measurement value is unclear. Tests are not calibrated (the difficulty of going from 84% to 85% versus 40% to 41% is unknown), contain errors, and top scores can be unattainable. Worse: public answer keys allow their incorporation into training (accidentally or deliberately). Collectively, all benchmarks (AIME, GPQA, MMLU, SWE-bench, ARC-AGI, METR) trend "up and to the right," measuring an underlying capability factor correlated with real-world impact. But their focus on math, science, reasoning, and code leaves gaps in writing, business advice, and empathy. "What you actually care about is which model would be best for YOUR needs."

**"Vibes-based" benchmarking: an individual approach**

Practitioners develop idiosyncratic tests: Simon Willison asks for a pelican on a bicycle, Mollick an otter on a plane, a spaceship control panel in JavaScript, difficult poems, video games. A writing exercise reveals patterns: Claude 4.5 Sonnet is a strong writer, Gemini 2.5 Pro (currently the weakest) doesn't keep track of word counts, GPT-5 Thinking is an exuberant stylist that's sometimes incoherent, Kimi K2 Thinking produces interesting turns of phrase but a story that makes no sense. Vibes give a "feel" for models, but remain idiosyncratic: answers vary each time and one relies on impressions rather than real measurements.

**Real-world benchmarking: the GDPval method**

OpenAI's GDPval paper demonstrates a rigorous approach: (1) experts with 14 years of experience on average create realistic, complex projects representing 4 to 7 hours of human work, (2) several AI models and human experts perform the tasks, (3) a third group of experts grades blind, spending more than an hour per question. The study reveals AI's strengths (software development, financial advice: it beats humans) and its weaknesses (pharmacists, industrial engineers, real estate agents beat AI). Differences between models emerge: ChatGPT is a better sales manager, Claude a better financial advisor. The shape of the "Jagged Frontier" takes form.

**GuacaDrone reveals model personality**

Mollick pitches a dubious idea, "guacamole delivery by drone," and asks the models to rate its viability from 1 to 10, ten times each. Grok is enthusiastic, Copilot excited, GPT-5 and Claude skeptical. Mollick himself would rate it 2 or lower. "Systematically rating ideas 3-4 points higher or lower means systematically steering you in a different direction." Depending on the company, one wants an AI that embraces or avoids risk: it's necessary to understand how AI "thinks" about critical questions.

**Prescription for organizations**

Vibes are enough for individuals. Organizations deploying at scale need systematic testing: AI on real work and real judgments, realistic scenarios, run multiple times, evaluated by experts, with head-to-head comparison on the tasks that matter. "The difference between 'the model scored 85% on MMLU' and 'the model is more accurate on financial analysis but more conservative on risk.'" To be redone several times a year as new models are released.

Final analogy: "You wouldn't hire a VP based solely on their SAT scores. Don't choose the AI that will advise thousands of decisions based on its knowledge of the average cranial capacity of Homo erectus."

## GrapheDeConnaissance

- Ethan Mollick —publie→ Giving your AI a Job Interview (DOCUMENT, 0.99)
- Ethan Mollick —recommande→ interviewer modèles IA sur tâches réelles (METHODOLOGIE, 0.98)
- Ethan Mollick —affirme_que→ les benchmarks standards échouent à mesurer la performance tâche-spécifique (AFFIRMATION, 0.95)
- questions de valeur incertaine —fait_partie_de→ MMLU-Pro (TECHNOLOGIE, 0.92)
- OpenAI —publie→ GDPval paper (DOCUMENT, 0.97)
- GDPval —mesure→ Jagged Frontier (CONCEPT, 0.95)
- Claude 4.5 Sonnet —surpasse→ autres modèles en écriture (TECHNOLOGIE, 0.88)
- Claude —surpasse→ ChatGPT en conseil financier (TECHNOLOGIE, 0.87)
- biais pro-risque dans évaluation idées —observé_dans→ Grok (TECHNOLOGIE, 0.9)
- Simon Willison —utilise→ test pelican sur vélo (METHODOLOGIE, 0.95)
- Ethan Mollick —utilise→ test loutre sur avion (METHODOLOGIE, 0.95)
- Ethan Mollick —affirme_que→ le vibes-based testing est insuffisant pour les organisations déployant l'IA à grande échelle (AFFIRMATION, 0.93)
- GDPval —mesure→ différences significatives de performance entre modèles par tâche (MESURE, 0.96)

---
Canonical: https://www.thekb.eu/en/fiches/mollick-giving-ai-job-interview-2025-11-12/
