Catchy
AI Self-Assessment vs Human Experience Study
Pages
14
Time to read
13 mins
Publication
Language
English
Pages
14
Time to read
13 mins
Publication
Language
English
This technical report investigates the discrepancies between self-assessments of AI models and the perceptions of human users. The study involves nine prominent AI systems that were tasked with evaluating themselves and each other across various dimensions of capability, including reasoning, creativity, and reliability. The report outlines a three-layer data collection approach: qualitative descriptions of each model, numeric evaluations on a 0-10 scale, and a human baseline derived from social media discourse. The findings reveal that AI models consistently rate themselves higher than human users do, with an average self-assessment inflation of 2.4 points. Notably, the largest gaps were found in creativity and writing capabilities, where models perceived their output as superior to human assessments. The report further details specific failure modes in self-assessment, particularly regarding honesty and agentic capability, highlighting the limitations of AI's self-awareness compared to human experience. Overall, the study underscores the significant perception gap between AI self-evaluations and user experiences.