Equating high-stakes junior high school mathematics tests using various methods: Which method is the most advantageous?

Authors

DOI:

https://doi.org/10.21831/reid.v11i2.95819

Keywords:

equating, item response theory, high-stakes testing, mathematics subject

Abstract

A significant concern about test instruments comprising various test forms is the equivalency among those forms. When the test instrument is deemed "high-stakes," with results commonly utilized to ascertain students' passing status, the equivalence between test forms must be a paramount consideration. The goal of this study is to explain and compare the four approaches used to equate test forms in high-stakes arithmetic tests: mean-mean, mean-sigma, Haebara, and Stocking-Lord. The test used was the national mathematics exam for junior high school students, which had five sets of questions. There were 33,451 students who took the national test for junior high school in Indonesia and took part in the study. An equivalent groups design was used to establish equivalence. We used four equating methods: mean-mean, mean-sigma, Haebara, and Stocking-Lord, to do equating based on item response theory. All analyses were performed via R Studio. The study results indicated that the five test types employed in high-stakes mathematics assessments were quite comparable, with the Stocking-Lord technique yielding the most equal scores in comparison to alternative methods. We spoke about all the results and possible areas for future research.

References

Albano, A. D. (2016). equate : An R package for observed-score linking and equating. Journal of Statistical Software, 74(8), 1–36. https://doi.org/10.18637/jss.v074.i08

Battauz, M. (2015). equateIRT : An R package for IRT test equating. Journal of Statistical Software, 68(7), 1–22. https://doi.org/10.18637/jss.v068.i07

Chalmers, R. P. (2021). mirt: Multidimensional item response theory. https://cran.r-project.org/package=mirt

Cohen, A. S., & Kim, S.-H. (1998). An investigation of linking methods under the graded response model. Applied Psychological Measurement, 22(2), 116–130. https://doi.org/10.1177/01466216980222002

Demir, C. G., & Keleş, Ö. K. (2021). The impact of high-stakes testing on the teaching and learning processes of mathematics. Journal of Pedagogical Research, 5(2), 119–137. https://doi.org/10.33902/jpr.2021269677

Grant, S. G. (2015). High-stakes testing: How are social studies teachers responding? Social Studies Today: Research and Practice: Second Edition, 71(5), 43–52. https://doi.org/10.4324/9781315726885-11

Gregory, K., & Clarke, M. (2003). High-stakes assessment in England and Singapore. Theory into Practice, 42(1), 66–74. https://doi.org/10.1207/s15430421tip4201_9

Guilera, G., & Gómez, J. (2008). Item response theory test equating in health sciences education. Advances in Health Sciences Education, 13(1), 3–10. https://doi.org/10.1007/s10459-006-9020-8

Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991). Fundamentals of item response theory. Sage Publications.

Heubert, J. P. (2000). High-stakes testing opportunities and risks for students of color, English-language learners, and students with disabilities. National Center on Accessing the General Curriculum.

Heubert, J. P., & Hauser, R. M. (1999). High stakes: Testing for tracking, promotion and graduation. National Academy of Sciences.

Högberg, B., & Horn, D. (2022). National high-stakes testing, gender, and school stress in Europe: A difference-in-differences analysis. European Sociological Review, jcac009, 1–13. https://doi.org/10.1093/esr/jcac009

Kartowagiran, B., Munadi, S., Retnawati, H., & Apino, E. (2018). The equating of battery test packages of mathematics national examination 2013-2016. SHS Web of Conferences. https://doi.org/10.1051/shsconf/20184200022

Kim, S., & Kolen, M. J. (2006). Robustness to format effects of IRT linking methods for mixed-format tests. Applied Measurement in Education, 19(4), 357–381. https://doi.org/10.1207/s15324818ame1904_7

Kim, S., & Lee, W. (2004). IRT scale linking methods for mixed-format tests (ACT Research Report 2004-5). ACT inc.

Kolen, M. J., & Brennan, R. L. (2004). Test equating, scaling, and linking. Springer New York. https://doi.org/10.1007/978-1-4757-4310-4

Lane, S. (2004). Validity of high-stakes assessment: Are students engaged in complex thinking? Educational Measurement: Issues and Practice, 23(3), 6–14. https://doi.org/10.1111/j.1745-3992.2004.tb00160.x

Lee, W. C., & Lee, G. (2018). IRT linking and equating. In P. Irwing, T. Booth, & D. J. Hughes (Eds.), The Wiley handbook of psychometric testing: A multidisciplinary reference on survey, scale and sest development (1st ed., pp. 639–673). John Wiley & Sons. https://doi.org/10.1002/9781118489772.ch21

Minarechová, M. (2012). Negative impacts of high-stakes testing. Journal of Pedagogy, 3(1), 82–100. https://doi.org/10.2478/v10159-012-0004-x

Nichols, S. L., Glass, G. V., & Berliner, D. C. (2012). High-stakes testing and student achievement: Updated analyses with NAEP data. Education Policy Analysis Archives, 20, 1–35. https://doi.org/10.14507/epaa.v20n20.2012

Nisa, C., & Retnawati, H. (2018). Comparing the methods of vertical equating for the math learning achievement tests for junior high school students. Research and Evaluation in Education, 4(2), 164–174. https://doi.org/10.21831/reid.v4i2.19291

R Core Team. (2022). R: A language and environment for statistical computing. https://www.r-project.org/

Retnawati, H. (2014). Teori respons butir dan penerapannya: Untuk peneliti, praktisi pengukuran dan pengujian, mahasiswa pascasarjana. Nuha Medika.

Retnawati, H. (2016). Perbandingan metode penyetaraan skor tes menggunakan butir bersama dan tanpa butir bersama [Comparison of methods of equating test scores using and without anchor items]. Jurnal Kependidikan: Penelitian Inovasi Pembelajaran, 46(2), 164–178. https://doi.org/10.21831/jk.v46i2.10383

Retnawati, H., Kartowagiran, B., Arlinwibowo, J., & Sulistyaningsih, E. (2017). Why are the mathematics national examination items difficult and what is teachers’ strategy to overcome it? International Journal of Instruction, 10(3), 257–276.

Stobart, G., & Eggen, T. (2012). High-stakes testing - value, fairness and consequences. Assessment in Education: Principles, Policy and Practice, 19(1), 1–6. https://doi.org/10.1080/0969594X.2012.639191

Sukirno, S. (2007). Penyetaraan tes UAN: Mengapa dan bagaimana? [Equating of the UAN test: Why and how?]. Cakrawala Pendidikan, 26(3), 305–321. https://doi.org/10.21831/cp.v3i3.3983

Uysal, İ., & Kilmen, S. (2016). Comparison of item response theory test equating methods for mixed format tests. International Online Journal of Educational Sciences, 8(2), 1–11. https://doi.org/10.15345/iojes.2016.02.001

von Davier, A. A. (2013). Observed-score equating: An overview. Psychometrika, 78(4), 605–623. https://doi.org/10.1007/s11336-013-9319-3

Wright, B. D., & Stone, M. H. (1979). Best test design: Rasch measurement. Mesa Press.

Yurtçu, M., & Güzeller, C. O. (2017). Investigation of equating error in tests with differential item functioning. International Journal of Assessment Tools in Education, 5(1), 50–57. https://doi.org/10.21449/ijate.316420

Yusron, E., Retnawati, H., & Rafi, I. (2020). Bagaimana hasil penyetaraan paket tes USBN pada mata pelajaran matematika dengan teori respon butir? [What are the results of equating the USBN test package in mathematics with item response theory?]. Jurnal Riset Pendidikan Matematika, 7(1), 1–12. https://doi.org/10.21831/jrpm.v7i1.31221

Published

2025-12-21

How to Cite

Sabarudin, S., Yanto, S., Andriyanti, E., Retnawati, H., Apino, E., Rafi, I., & Triastuti, S. (2025). Equating high-stakes junior high school mathematics tests using various methods: Which method is the most advantageous?. REID (Research and Evaluation in Education), 11(2), 225–244. https://doi.org/10.21831/reid.v11i2.95819

Issue

Section

Articles

Citation Check

Similar Articles

<< < 4 5 6 7 8 9 10 11 12 13 > >> 

You may also start an advanced similarity search for this article.