Abstract

There are numerous statistical procedures for detecting items that function differently across subgroups of examinees that take a test or survey. However, in endeavouring to detect items that may function differentially, selection of the statistical method is only one of many important decisions. In this article, we discuss the important decisions that affect investigations of differential item functioning (DIF) such as choice of method, sample size, effect size criteria, conditioning variable, purification, DIF amplification, DIF cancellation, and research designs for evaluating DIF. Our review highlights the necessity of matching the DIF procedure to the nature of the data analysed, the need to include effect size criteria, the need to consider the direction and balance of items flagged for DIF, and the need to use replication to reduce Type I errors whenever possible. Directions for future research and practice in using DIF to enhance the validity of test scores are provided.

Keywords

Differential item functioningPsychologyReplication (statistics)Matching (statistics)Sample size determinationTest (biology)Affect (linguistics)Sample (material)Differential (mechanical device)Function (biology)StatisticsApplied psychologySocial psychologyComputer sciencePsychometricsItem response theoryClinical psychologyMathematics

Related Publications

Publication Info

Year
2013
Type
article
Volume
19
Issue
2-3
Pages
170-187
Citations
64
Access
Closed

External Links

Social Impact

Social media, news, blog, policy document mentions

Citation Metrics

64
OpenAlex

Cite This

Stephen G. Sireci, Joseph A. Rios (2013). Decisions that make a difference in detecting differential item functioning. Educational Research and Evaluation , 19 (2-3) , 170-187. https://doi.org/10.1080/13803611.2013.767621

Identifiers

DOI
10.1080/13803611.2013.767621