John Snow Labs
Automated De-Identification of Clinical Text Datasets
Pages
13
Time to read
34 mins
Publication
Language
English
Pages
13
Time to read
34 mins
Publication
Language
English
This research article presents findings on automated de-identification of large real-world clinical text datasets, specifically focusing on a system that has successfully de-identified over one billion clinical notes. The paper outlines the challenges faced in achieving high accuracy in de-identification without manual review and describes a hybrid context-based model architecture that outperforms existing models, including those from major cloud service providers. The system achieves over 98% coverage of sensitive data across multiple languages and employs a method for data obfuscation that maintains clinical consistency. The article also discusses the importance of making unstructured clinical data available for secondary uses while ensuring compliance with privacy regulations. It details the architecture of the proposed system, including its scalable NLP pipeline and the processes involved in identifying and obfuscating protected health information (PHI). Overall, the paper contributes to the field of natural language processing in healthcare by addressing practical challenges in automated de-identification.