Aid4Mail
Podesta Corpus as AI Classification Benchmark Methodology Note
Pages
14
Time to read
24 mins
Publication
Language
English
Pages
14
Time to read
24 mins
Publication
Language
English
This methodology note documents an experiment evaluating the WikiLeaks Podesta release as a benchmark for AI-based email classification in eDiscovery and digital forensics. The experiment involved analyzing a specific slice of the corpus, consisting of 1,083 emails from March 2016, to assess the reliability of the corpus for measuring classifier accuracy on subjective themes. The study outlines two main operations: theme identification and prompt classification, executed independently across seven model instances. The findings indicate significant variability in model outputs, with 45% of flagged emails being identified by only one model. The note concludes that while the Podesta corpus has attractive properties for AI evaluation, its subjective nature and lack of external ground truth limit its effectiveness as a standalone benchmark. Recommendations for corpus design are provided, suggesting the use of the Podesta corpus as background data paired with a separately constructed responsive set based on verifiable forensic categories.