Anthropic
Sabotage Evaluations for Frontier Language Models
Pages
69
Time to read
140 mins
Publication
Language
English
Pages
69
Time to read
140 mins
Publication
Language
English
This document is a technical report that discusses the potential risks associated with sabotage capabilities in frontier language models. It outlines the development of threat models and evaluations designed to assess whether a model can successfully sabotage the activities of organizations involved in AI development. The authors present evaluations conducted on Anthropic's Claude 3 Opus and Claude 3.5 Sonnet models, indicating that current mitigations are generally sufficient to address sabotage risks, although stronger measures may be necessary as model capabilities evolve. The report emphasizes the importance of evaluating risks from models that could undermine oversight and decision-making processes. It also details specific evaluation tasks aimed at measuring sabotage capabilities, including human decision sabotage and code sabotage. The authors provide insights into the design of these evaluations and the implications for ensuring safe deployment of advanced AI systems, highlighting the need for ongoing assessment as capabilities improve.