Anthropic - Measuring political bias in Claude - Even-handedness 94-95% - Paired Prompts method - Open-source evaluation - Character training - Comparison of 6 models - Neutrality system prompt - GitHub
AI benchmarking beyond standard tests - Interviewing AI models for specific use cases - Jagged Frontier - OpenAI GDPval - Vibes vs real measurements - GuacaDrone example - Ethan Mollick - One Useful Thing