OpenAI
GPT-5.4 Thinking System Safety Evaluations
Pages
38
Time to read
60 mins
Publication
Language
English
Pages
38
Time to read
60 mins
Publication
Language
English
This white paper presents the safety evaluations for the GPT-5.4 Thinking model, detailing the training data and methodologies as well as baseline evaluations against disallowed content. The document describes the model's training on diverse datasets, emphasizing the rigorous filtering processes used to ensure data quality and reduce risks associated with harmful content. The evaluations include benchmarks for disallowed content across various categories, highlighting the model's performance in areas such as nonviolent illicit behavior and self-harm. Additionally, the paper outlines dynamic multi-turn evaluations that reflect real user interactions, aiming to identify potential issues that may arise in extended conversations. The findings indicate that GPT-5.4 Thinking generally performs on par with its predecessor while showing significant improvements in certain evaluation areas. This document aims to provide transparency in evaluating safety measures and performance standards for this reasoning model.