- In this project we artificially introduce representation bias to an automatic scoring algorithm by training three LLMs on single-demographic subsets of a larger dataset
- We then use the models to predict error on the more diverse dataset and examine whether there is differential error between demographic groups.
Holmes, L., Morris, W., Crossley, S., & Choi, J. S. (2026). Assessing fairness in finetuned scoring models with demographically restricted training data. Assessing Writing, 68, 101032. https://doi.org/10.1016/j.asw.2026.101032
Please see the following repository for guidance on obtaining the PERSUADE corpus: scrosseye/PERSUADE_corpus.