Generative AI Doesn’t Replace Theory. It Makes Theory More Important Than Ever.
Everyone is excited about using large language models for humanities research.
For the first time we can ask models to identify much higher level humanistic constructs than we have in the past. This is what I call the second wave of scale. The first wave was about quantitative scaling — observing large amounts of cultural artifacts instead of anecdotal evidence. The second wave is about scaling our concepts up from more rudimentary features like words and “topics.”
The problem though is that the easier AI makes interpretation, the easier it becomes to mistake plausible answers for meaningful scholarship.
The question is not simply Can AI do this?
But: How do we know we’re studying the thing we think we’re studying?
My latest article on a theory-first approach to AI-assisted humanities research tries to address this problem.
AI has changed what “method” means
Traditional computational humanities research mostly worried about algorithms.
Today our measurement instrument is often a prompt.
Every prompt contains assumptions.
Every benchmark contains assumptions.
Every fine-tuned model contains assumptions.
Instead of pretending these assumptions don’t exist, we should make them explicit.
Three questions every AI project should answer
- What exactly are you trying to study? Are you studying the AI itself? Or are you using AI as an instrument to study literature, history, or culture? It is essential to distinguish between AI as object or research instrument.
- Are you measuring the right thing? Here the concept of construct validity can be very valuable. Stress-testing your prompts can help figure out if you are capturing what you think you are in your generative workflow and also guard against “LLM hacking,” i.e. when results depend on particular model configurations not real-world effects.
- “Validation” is more complicated than you think. Different approaches to validation are tied to different research goals. Are you capturing central tendencies or human diversity? What does theory say with respect to your question? Second, what effect is your model’s training and biases having on your results? How might you take your research to the next level where model tuning is part of the theoretical framework too?

Conclusion
The future of humanities AI isn’t about finding the perfect “killer” model (or AGI, ugh).
It’s about building better research design.
Theory helps us decide what to ask, how to measure it, how to validate it, and how to interpret the results.
Generative AI may be transforming research, but the foundations of good scholarship remain remarkably familiar.
