Optimal number of response categories for a Likert-type questionnaire: A simulation study within the frameworks of Classical Test and Item Response Theory

Authors

DOI:

https://doi.org/10.21831/reid.v12i1.95846

Keywords:

Likert scale, simulation study, number of categories response, classical test theory, item response theory

Abstract

Likert-type questionnaires have been widely used in educational contexts with many types of response categories. The use of many response categories raises the question of the optimal number of response categories for a Likert-type questionnaire. This issue continues to be investigated by researchers, and the current findings present an inconsistent answer. This study aims to find the optimal number of response categories in developing the Likert-type questionnaire based on classical test theory (CTT) and item response theory (IRT) framework. The simulation design used 8 (number of response categories: 3- to 10-point scale) × 3 (number of items: 20, 25, and 30 items) × 5 (sample size: 100, 200, 500, 1000, and 5000) = 120 conditions. We generated a dataset for each condition for 1000 replications using WinGen 3 and randomly took 10 data points from each dataset to estimate its psychometric properties using the CTT and IRT analysis frameworks. The optimal number of response categories was determined based on the highest reliability coefficient, total explained variance, and information function values. Using the CTT analysis framework, the simulation results indicated that the 8-point scale is optimal for Likert-type questionnaires with 20 and 30 items and for all sample size variations. However, for questionnaires with 25 items, the 8-point scale is only optimal for sample sizes of 100 and 5000, while for sample sizes of 200 to 1000, the 10-point scale is more optimal. Using the IRT analysis framework, the simulations reveal that it is necessary to consider the number of items and the sample size to determine the optimal number of response categories. In this context, 8- and 10-point scales are generally more recommended for Likert-type questionnaires. The findings provide researchers with insight into the optimal number of response categories and the analytical model for developing a Likert-type questionnaire.

References

Aiken, L. R. (1983). Number of response categories and statistics on a teacher rating scale. Educational and Psychological Measurement, 43(2), 397–401. https://doi.org/10.1177/001316448304300209

Allen, M. J., & Yen, W. M. (2002). Introduction to measurement theory. Belmont, CA: Waveland Press.

Anastasi, A., & Urbina, S. (1997). Psychological testing. In Psychology for Nurses (7th ed.). Upper Saddle Riper, NJ: Prentice Hall.

Assa’diyah, N. H., & Hadi, S. (2021). Developing student character assessment questionnaire on French subject in state high schools. REID (Research and Evaluation in Education), 7(2), 168–176. https://doi.org/10.21831/reid.v7i2.43196

Aybek, E. C., & Toraman, C. (2022). How many response categories are sufficient for Likert type scales? An empirical study based on the item response theory. International Journal of Assessment Tools in Education, 9(2), 534–547. https://doi.org/10.21449/ijate.1132931

Bahar, R., Setiawati, F. A., Sutarji, A., Hidayat, O., & Sudarna, N. (2021). Perbandingan metode Skala Thurstone dalam mengukur kompetensi kepribadian guru. Measurement In Educational Research, 1(2), 97–103. https://doi.org/10.33292/meter.v1i2.162

Bora, B. (2013). Pazarlama araştırmalarında kullanılan likert türü ölçeklerin uygulanabilirliğinin incelenmesi [A study on the applicability of the Likert type scales in marketing]. Doctoral Dissertation, Sakarya University, Turkey.

Chalmers, R. P. (2012). mirt: A multidimensional item response theory package for the R environment. Journal of Statistical Software, 48(6), 1–29. https://doi.org/10.18637/jss.v048.i06

Chang, L. (1994). A psychometric evaluation of 4-point and 6-point Likert-type scales in relation to reliability and validity. Applied Psychological Measurement, 18(3), 205–215. https://doi.org/10.1177/014662169401800302

Comrey, A. L., & Montag, I. (1982). Comparison of factor analytic results with two-choice and seven-choice personality item formats. Applied Psychological Measurement, 6(3), 285–289. https://doi.org/10.1177/014662168200600304

DeMars, C. (2010). Item response theory. New York, NY: Oxford University Press.

DeVellis, R. F. (2017). Scale development: Theory and applications (4th ed.). Los Angeles, CA: Sage Publication.

Dunn-Rankin, P., Knezek, G. A., Wallace, S. R., & Zhang, S. (2014). Scaling methods. In Scaling Methods. Lawrence Erlbaum Associates. https://doi.org/10.4324/9781410611048

Eutsler, J., & Lang, B. (2015). Rating scales in accounting research: The impact of scale points and labels. Behavioral Research in Accounting, 27(2), 35–51. https://doi.org/10.2308/bria-51219

Falani, I., Akbar, M., & Naga, D. S. (2020). The precision of students’ ability estimation on combinations of item response theory models. International Journal of Instruction, 13(4), 545–558. https://doi.org/10.29333/iji.2020.13434a

Feinberg, R. A., & Rubright, J. D. (2016). Conducting simulation studies in psychometrics. Educational Measurement: Issues and Practice, 35(2), 36–49. https://doi.org/10.1111/emip.12111

García-Pérez, M. A. (2024). Are the steps on Likert scales equidistant? Responses on visual analog scales allow estimating their distances. Educational and Psychological Measurement, 84(1), 91–122. https://doi.org/10.1177/00131644231164316

Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991). Fundamentals of item response theory. Newbury Park, CA: Sage Publications.

Han, K. T. (2007). WinGen: Windows software that generates item response theory parameters and item responses. Applied Psychological Measurement, 31(5), 457–459. https://doi.org/10.1177/0146621607299271

İlhan, M., & Güler, N. (2017). The number of response categories and the reverse scored item problem in Likert-type scales: A study with the Rasch model. Journal of Measurement and Evaluation in Education and Psychology, 8(3), 320–342. https://doi.org/10.21031/epod.321057

Jebb, A. T., Ng, V., & Tay, L. (2021). A review of key Likert scale development advances: 1995–2019. Frontiers in Psychology, 12(May), 1–14. https://doi.org/10.3389/fpsyg.2021.637547

Jones, W. P., & Loe, S. A. (2013). Optimal number of questionnaire response categories: More may not be better. SAGE Open, 3(2), 1–10. https://doi.org/10.1177/2158244013489691

Joshi, A., Kale, S., Chandel, S., & Pal, D. (2015). Likert scale: Explored and explained. British Journal of Applied Science & Technology, 7(4), 396–403. https://doi.org/10.9734/bjast/2015/14975

Kim, K. H. (1998). An analysis of optimum number of response categories for Korean consumers. Journal of Global Academy of Marketing Science, 1(1), 61–86. https://doi.org/10.1080/12297119.1998.9707386

King, L. A., King, D. W., & Klockars, A. J. (1983). Dichotomous and multipoint scales using bipolar adjectives. Applied Psychological Measurement, 7(2), 173–180. https://doi.org/10.1177/014662168300700205

Kusmaryono, I., Wijayanti, D., & Maharani, H. R. (2022). Number of response options, reliability, validity, and potential bias in the use of the Likert scale education and social science research: A literature review. International Journal of Educational Methodology, 8(4), 625–637. https://doi.org/10.12973/ijem.8.4.625

Lee, J., & Paek, I. (2014). In search of the optimal number of response categories in a rating scale. Journal of Psychoeducational Assessment, 32(7), 663–673. https://doi.org/10.1177/0734282914522200

Leung, S. O. (2011). A comparison of psychometric properties and normality in 4-, 5-, 6-, and 11-point likert scales. Journal of Social Service Research, 37(4), 412–421. https://doi.org/10.1080/01488376.2011.580697

Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 22, 2–55.

Lord, F. M. (1954). Chapter II: Scaling. Review of Educational Research, 24(5), 375–392. https://doi.org/10.3102/00346543024005375

Lozano, L. M., García-Cueto, E., & Muñiz, J. (2008). Effect of the number of response categories on the reliability and validity of rating scales. Methodology, 4(2), 73–79. https://doi.org/10.1027/1614-2241.4.2.73

Maydeu-Olivares, A., Kramp, U., García-Forero, C., Gallardo-Pujol, D., & Coffman, D. (2009). The effect of varying the number of response alternatives in rating scales: Experimental evidence from intra-individual effects. Behavior Research Methods, 41(2), 295–308. https://doi.org/10.3758/BRM.41.2.295

Mellenbergh, G. J. (1996). Measurement precision in test score and item response models. Psychological Methods, 1(3), 293–299. https://doi.org/10.1037/1082-989X.1.3.293

Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory. McGraw-Hill. https://doi.org/10.2307/1161962

Oppenheim, A. N., & Torgerson, W. (1961). Theory and methods of scaling. In The British Journal of Sociology (Vol. 12). John Willey & Sons. https://doi.org/10.2307/588039

Posit Team. (2023). RStudio: Integrated development environment for R. Boston, MA: Posit Software, PBC. Retrieved from http://www.posit.co/

Preston, C. C., & Colman, A. M. (2000). Optimal number of response categories in rating scales: Reliability, validity, discriminating power, and respondent preferences. Acta Psychologica, 104(1), 1–15. https://doi.org/10.1016/S0001-6918(99)00050-5

Price, L. R. (2017). Psychometric methods, theory into practice. Guilford Press.

Qasem, M. A. N., Almoshigah, T. S., & Gupta, S. (2014). The effect of number of alternatives on validity and reliability in Likert scale. International Journal of Innovative Research & Studies, 3(6), 324–333. https://doi.org/10.13140/2.1.2237.2803

Retnawati, H. (2014). Teori respons butir dan penerapannya: Untuk peneliti, praktisi pengukuran dan pengujian, mahasiswa pascasarjana [Item response theory and its applications: For researchers, measurement and testing practitioners, & graduate students]. Yogyakarta: Nuha Medika.

Retnawati, H. (2016). Analisis kuantitatif instrumen penelitian: Panduan peneliti, mahasiswa, dan psikometrian [Quantitative analysis of research instruments: Guide for researchers, students, and psychometricians]. Yogyakarta: Parama Publishing.

Robinson, S. (2004). Simulation: The practice of model development and use. West Sussex, UK: John Wiley & Sons.

Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores (psychometric monograph no. 17). Psychometrika, 34(1), 1–100. https://doi.org/10.1007/BF02290599

Santoso, P. H., Setiawati, F. A., Ismail, R., & Suhariyono, S. (2023). Comparing IRT properties among different category numbers: A case from attitudinal measurement on physics education research. Discover Psychology, 3(1), 39. https://doi.org/10.1007/s44202-023-00101-6

Sharma, H. (2022). How short or long should be a questionnaire for any research? Researchers dilemma in deciding the appropriate questionnaire length. Saudi Journal of Anaesthesia, 16(1), 65–68. https://doi.org/10.4103/sja.sja_163_21

Taherdoost, H. (2019). What is the best response scale for survey and questionnaire design: Review of different lengths of rating scale/attitude scale/Likert scale. International Journal of Academic Research in Management, 8(1), 2296–1747.

Tarka, P. (2015). Likert scale and change in range of response categories vs. The factors extraction in EFA model. Acta Universitatis Lodziensis: Folia Oeconomica, 1(311), 27–36. https://doi.org/10.18778/0208-6018.311.04

Wakita, T., Ueshima, N., & Noguchi, H. (2012). Psychological distance between categories in the likert scale. Educational and Psychological Measurement, 72(4), 533–546. https://doi.org/10.1177/0013164411431162

Weng, L.-J. (2004). Impact of the number of response categories and anchor labels on coefficient alpha and test-retest reliability. Educational and Psychological Measurement, 64(6), 956–972. https://doi.org/10.1177/0013164404268674

Willse, J. T. (2018). CTT: Classical test theory functions. Retrieved from https://cran.r-project.org/package=CTT

Published

2026-07-30

How to Cite

Soniawati, N. I., Apino, E., Retnawati, H., Setiawati, F. A., Rafi, I., Rosyada, M. N., & Rezkilaturahmi, R. (2026). Optimal number of response categories for a Likert-type questionnaire: A simulation study within the frameworks of Classical Test and Item Response Theory. REID (Research and Evaluation in Education), 12(1), 1–25. https://doi.org/10.21831/reid.v12i1.95846

Issue

Section

Articles

Citation Check

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.