Oxford Lets OpenAI Train AI Models on Bodleian Texts, Guardian Reports
Oxford has allowed OpenAI to use old texts from its Bodleian Library to train its AI models, according to internal papers seen by the Guardian. Ethan Penny and Dan Milmo broke the story on Saturday.
The papers reveal that texts OpenAI scanned at the library were integrated into the firm’s training data. Oxford publicly announced its deal with OpenAI in March 2025, stating that OpenAI’s tools would assist in scanning rare texts for wider access. However, it did not initially disclose that the texts would be used for AI training.
By June 2025, the Bodleian Library had provided OpenAI with 125,000 scans of old PhD theses, dating back to the 19th and 20th centuries, from European and US universities.
Internal meeting notes, obtained via a freedom of information request, show some Oxford staff expressed concerns:
- Reputational risk: Fear of potential harm to Oxford’s name.
- Energy use: Worries about the environmental impact of AI.
Oxford, however, defended the deal, stating the scans were small in scale, out of copyright, and not exclusive to OpenAI. The library intends to make the scans publicly available online in the coming months.
OpenAI stressed the importance of diverse historical and cultural perspectives in AI training: "With more than a billion people using this technology in everyday life, it’s important it reflects different cultures, histories, and perspectives."
This deal comes as AI firms acquire printed books for slop-free training data, as web-based text is becoming saturated with AI-generated content.
The Bodleian Library’s books remain intact under this agreement, contrasting with reports of AI-focused book scanning practices that involve destruction of rare volumes.