Ziff Davis
Analysis of High-Authority Web Publisher Content in LLM Training
Pages
39
Time to read
47 mins
Language
English
Pages
39
Time to read
47 mins
Language
English
This technical report analyzes the predominant use of high-authority commercial web publisher content in training leading large language models (LLMs). The report outlines the methodology and datasets utilized in training these models, specifically focusing on four key datasets: Common Crawl, the Colossal Clean Crawled Corpus (C4), OpenWebText, and OpenWebText2. Each dataset is described in terms of its composition, processing methods, and relevance to the training of breakthrough commercial LLMs. The report highlights that major LLM companies have prioritized high-quality content from commercial publishers, which has implications for intellectual property rights and technological progress. It also discusses recent legislative proposals aimed at increasing transparency regarding the training data used by model developers. The findings are intended to inform public discourse on the intersection of AI technology and copyright law, emphasizing the significance of understanding the sources of training data for LLMs.