While language testers are becoming more and more enamoured with the application of AI scoring systems, this has come about despite an almost total lack of awareness of the potential of AI to add or significantly remove value from our tests. There is no doubt that AI has much to offer in addition to the scoring of tests, for example in auto text and item generation as well as auto difficulty estimation. However, in this presentation I will focus on some limitations (and solutions) to the ethical application of AI scoring models in our tests.
While the guidelines and standards we turn to when designing and operationalising new tests are of great value, they fail miserably when it comes to AI scoring systems, by hardly mentioning the subject. The same can be said of the language testing literature and theories of validation. Instead, I look to two initiatives from technology researchers in the USA: the concept of Datasheets for Datasets (Gebru et al., 2019) and Model Cards (Mitchell et al., 2019) which are designed to offer an in-depth understanding of an AI-scoring model and the validation of its use in a specific context for a specific purpose.
I conclude the presentation by reflecting on how these approaches can contribute to a more inclusive and open conversation about the appropriate use of AI, which in turn will enable us to work together with our EdTech colleagues to build fairer and more transparent AI scoring systems for our tests.