International Computing Centre
Safety Evaluation of Institutional LLM-RAG Deployment
Pages
36
Time to read
76 mins
Publication
Language
English
Pages
36
Time to read
76 mins
Publication
Language
English
This document is a technical report detailing a safety evaluation framework for institutional Retrieval-Augmented Generation (RAG) deployments, specifically focusing on the Apertus model. Conducted by the United Nations International Computing Centre (UNICC) and the Machine Learning and Optimization Laboratory of EPFL, the evaluation presents a four-layer safety audit framework that assesses various safety properties including jailbreak resistance, out-of-scope refusal, hallucination robustness, and politically framed sycophancy. The report outlines the methodology used to benchmark the Apertus deployment against an identically harnessed GPT-5 RAG comparator. Findings indicate significant gaps in performance, particularly in out-of-scope refusal and factual correctness. The document also contributes a reusable safety assessment methodology and mitigation recommendations aimed at enhancing safety practices for organizations deploying retrieval-augmented AI systems. The objective is to advance governance approaches for such systems in high-trust institutional settings.