Text mining
Extract reproducible patterns from a defined corpus
Text mining uses computational representations and algorithms to extract patterns or classify content in a corpus. It includes supervised classification, dictionary approaches and information extraction. Choices about document boundaries, tokenisation, language and labels affect the construct being measured. Quantitative outputs can be interpreted alongside qualitative reading, but that integration must be designed rather than assumed.
Choose it when a corpus is too large for exhaustive close reading and a clear textual construct or retrieval task can be operationalised, with access to appropriate validation material.
Strengths
- Scales consistent analysis across large document collections
- Can support transparent, reusable classification pipelines
Limitations
- Language, genre and domain shift can invalidate measures
- Automated labels may confuse textual cues with the intended construct
Know the boundary
A text classifier predicts labels from language; it does not establish the author’s true beliefs or causal effects of the text.