Rethinking how AI reads history
Dr Sebastian Ahnert has been awarded funding through the 2025 Schmidt Sciences Humanities and AI Virtual Institute (HAVI) programme.
Schmidt Sciences has awarded $11 million to 23 research teams worldwide working at the intersection of artificial intelligence and the humanities, with Dr Ahnert’s project receiving around half a million dollars.
Working with Stanford University (lead) and Milan Polytechnic, the project will develop a new machine learning framework designed to handle the complexity and ambiguity found in historical archives. These sources often contain multiple layers of meaning beyond text – including handwriting styles, layout, annotations, ownership traces and historical revisions – and current AI models struggle to interpret them reliably.
The approach, called SETS, uses mathematical set theory to represent information as evolving graph structures rather than through large neural networks. The aim is to build a system that can retain nuance, ambiguity and multiple interpretations, and that reflects how humanities scholars read and question archival material. The framework will be applied to case studies exploring travel, identity and property rights in pre-modern Europe, Western Asia and Africa, drawing on dense multilingual records that are currently difficult to analyse using existing AI approaches.
The project funding covers two years and includes support for a postdoctoral researcher based in Cambridge.
“It has mostly been a side project of mine for the last few years, but it’s slowly becoming my main focus,” Ahnert said. “I am delighted it has been funded, as it is an unusual project. It represents a completely different way of looking at machine learning. It felt like a long shot, but it goes to show why it’s important to apply anyway.”
“There’s more and more criticism of the limitations of current models and so…the timing is perfect for a project like this.”
Dr Sebastian Ahnert
The King’s College, Cambridge, Fellow added: “The dominant framework is definitely the neural networks currently in use and that’s seen as the way forward, but more and more people are starting to talk about the power demands of training a model that just absorbs more and more data, but doesn’t necessarily produce better results, and the bias and other issues with this. There’s more and more criticism of its limitations and so I think the timing is perfect for a project like this to be funded.
“In theory, we are proposing a simpler method and a very different approach to machine learning.”
The framework aims to rethink how machine learning models are built and applied. Rather than relying on systems that absorb vast amounts of data and infer meaning statistically, SETS is designed to be transparent and interpretable. The team believes this will make it better suited to working with abstract and context-dependent material, where meaning depends on how information relates and changes rather than on a fixed text.
Early testing has already shown promise. While the humanities case studies present the biggest challenge because of the level of nuance involved, Ahnert is also planning to apply the approach to other fields, including biochemical research, where patterns are more structured and less open to interpretation.