The SubData tool, implemented as a Python library, evaluates the alignment between large language models (LLMs) and human perspectives in subjective annotation tasks. It is particularly relevant during data preprocessing, where it helps to improve data quality by harmonizing heterogeneous datasets, providing standardized keyword mappings and taxonomies, and enabling theory-driven analyses of perspective alignment. In addition, SubData includes ten curated datasets that can be used to identify model biases, test generalizability, and contextualize hate speech datasets within the broader literature.
Further information: This tool was developed as part of the KODAQS project, a partnership between GESIS, the University of Mannheim, and LMU Munich.

