Optimal number of response categories for a Likert-type questionnaire: A simulation study within the frameworks of Classical Test and Item Response Theory
DOI:
https://doi.org/10.21831/reid.v12i1.95846Keywords:
Likert scale, simulation study, number of categories response, classical test theory, item response theoryAbstract
Likert-type questionnaires have been widely used in educational contexts with many types of response categories. The use of many response categories raises the question of the optimal number of response categories for a Likert-type questionnaire. This issue continues to be investigated by researchers, and the current findings present an inconsistent answer. This study aims to find the optimal number of response categories in developing the Likert-type questionnaire based on classical test theory (CTT) and item response theory (IRT) framework. The simulation design used 8 (number of response categories: 3- to 10-point scale) × 3 (number of items: 20, 25, and 30 items) × 5 (sample size: 100, 200, 500, 1000, and 5000) = 120 conditions. We generated a dataset for each condition for 1000 replications using WinGen 3 and randomly took 10 data points from each dataset to estimate its psychometric properties using the CTT and IRT analysis frameworks. The optimal number of response categories was determined based on the highest reliability coefficient, total explained variance, and information function values. Using the CTT analysis framework, the simulation results indicated that the 8-point scale is optimal for Likert-type questionnaires with 20 and 30 items and for all sample size variations. However, for questionnaires with 25 items, the 8-point scale is only optimal for sample sizes of 100 and 5000, while for sample sizes of 200 to 1000, the 10-point scale is more optimal. Using the IRT analysis framework, the simulations reveal that it is necessary to consider the number of items and the sample size to determine the optimal number of response categories. In this context, 8- and 10-point scales are generally more recommended for Likert-type questionnaires. The findings provide researchers with insight into the optimal number of response categories and the analytical model for developing a Likert-type questionnaire.
References
Aiken, L. R. (1983). Number of response categories and statistics on a teacher rating scale. Educational and Psychological Measurement, 43(2), 397–401. https://doi.org/10.1177/001316448304300209
Allen, M. J., & Yen, W. M. (2002). Introduction to measurement theory. Belmont, CA: Waveland Press.
Anastasi, A., & Urbina, S. (1997). Psychological testing. In Psychology for Nurses (7th ed.). Upper Saddle Riper, NJ: Prentice Hall.
Assa’diyah, N. H., & Hadi, S. (2021). Developing student character assessment questionnaire on French subject in state high schools. REID (Research and Evaluation in Education), 7(2), 168–176. https://doi.org/10.21831/reid.v7i2.43196
Aybek, E. C., & Toraman, C. (2022). How many response categories are sufficient for Likert type scales? An empirical study based on the item response theory. International Journal of Assessment Tools in Education, 9(2), 534–547. https://doi.org/10.21449/ijate.1132931
Bahar, R., Setiawati, F. A., Sutarji, A., Hidayat, O., & Sudarna, N. (2021). Perbandingan metode Skala Thurstone dalam mengukur kompetensi kepribadian guru. Measurement In Educational Research, 1(2), 97–103. https://doi.org/10.33292/meter.v1i2.162
Bora, B. (2013). Pazarlama araştırmalarında kullanılan likert türü ölçeklerin uygulanabilirliğinin incelenmesi [A study on the applicability of the Likert type scales in marketing]. Doctoral Dissertation, Sakarya University, Turkey.
Chalmers, R. P. (2012). mirt: A multidimensional item response theory package for the R environment. Journal of Statistical Software, 48(6), 1–29. https://doi.org/10.18637/jss.v048.i06
Chang, L. (1994). A psychometric evaluation of 4-point and 6-point Likert-type scales in relation to reliability and validity. Applied Psychological Measurement, 18(3), 205–215. https://doi.org/10.1177/014662169401800302
Comrey, A. L., & Montag, I. (1982). Comparison of factor analytic results with two-choice and seven-choice personality item formats. Applied Psychological Measurement, 6(3), 285–289. https://doi.org/10.1177/014662168200600304
DeMars, C. (2010). Item response theory. New York, NY: Oxford University Press.
DeVellis, R. F. (2017). Scale development: Theory and applications (4th ed.). Los Angeles, CA: Sage Publication.
Dunn-Rankin, P., Knezek, G. A., Wallace, S. R., & Zhang, S. (2014). Scaling methods. In Scaling Methods. Lawrence Erlbaum Associates. https://doi.org/10.4324/9781410611048
Eutsler, J., & Lang, B. (2015). Rating scales in accounting research: The impact of scale points and labels. Behavioral Research in Accounting, 27(2), 35–51. https://doi.org/10.2308/bria-51219
Falani, I., Akbar, M., & Naga, D. S. (2020). The precision of students’ ability estimation on combinations of item response theory models. International Journal of Instruction, 13(4), 545–558. https://doi.org/10.29333/iji.2020.13434a
Feinberg, R. A., & Rubright, J. D. (2016). Conducting simulation studies in psychometrics. Educational Measurement: Issues and Practice, 35(2), 36–49. https://doi.org/10.1111/emip.12111
García-Pérez, M. A. (2024). Are the steps on Likert scales equidistant? Responses on visual analog scales allow estimating their distances. Educational and Psychological Measurement, 84(1), 91–122. https://doi.org/10.1177/00131644231164316
Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991). Fundamentals of item response theory. Newbury Park, CA: Sage Publications.
Han, K. T. (2007). WinGen: Windows software that generates item response theory parameters and item responses. Applied Psychological Measurement, 31(5), 457–459. https://doi.org/10.1177/0146621607299271
İlhan, M., & Güler, N. (2017). The number of response categories and the reverse scored item problem in Likert-type scales: A study with the Rasch model. Journal of Measurement and Evaluation in Education and Psychology, 8(3), 320–342. https://doi.org/10.21031/epod.321057
Jebb, A. T., Ng, V., & Tay, L. (2021). A review of key Likert scale development advances: 1995–2019. Frontiers in Psychology, 12(May), 1–14. https://doi.org/10.3389/fpsyg.2021.637547
Jones, W. P., & Loe, S. A. (2013). Optimal number of questionnaire response categories: More may not be better. SAGE Open, 3(2), 1–10. https://doi.org/10.1177/2158244013489691
Joshi, A., Kale, S., Chandel, S., & Pal, D. (2015). Likert scale: Explored and explained. British Journal of Applied Science & Technology, 7(4), 396–403. https://doi.org/10.9734/bjast/2015/14975
Kim, K. H. (1998). An analysis of optimum number of response categories for Korean consumers. Journal of Global Academy of Marketing Science, 1(1), 61–86. https://doi.org/10.1080/12297119.1998.9707386
King, L. A., King, D. W., & Klockars, A. J. (1983). Dichotomous and multipoint scales using bipolar adjectives. Applied Psychological Measurement, 7(2), 173–180. https://doi.org/10.1177/014662168300700205
Kusmaryono, I., Wijayanti, D., & Maharani, H. R. (2022). Number of response options, reliability, validity, and potential bias in the use of the Likert scale education and social science research: A literature review. International Journal of Educational Methodology, 8(4), 625–637. https://doi.org/10.12973/ijem.8.4.625
Lee, J., & Paek, I. (2014). In search of the optimal number of response categories in a rating scale. Journal of Psychoeducational Assessment, 32(7), 663–673. https://doi.org/10.1177/0734282914522200
Leung, S. O. (2011). A comparison of psychometric properties and normality in 4-, 5-, 6-, and 11-point likert scales. Journal of Social Service Research, 37(4), 412–421. https://doi.org/10.1080/01488376.2011.580697
Likert, R. (1932). A technique for the measurement of attitudes. Archives of Psychology, 22, 2–55.
Lord, F. M. (1954). Chapter II: Scaling. Review of Educational Research, 24(5), 375–392. https://doi.org/10.3102/00346543024005375
Lozano, L. M., García-Cueto, E., & Muñiz, J. (2008). Effect of the number of response categories on the reliability and validity of rating scales. Methodology, 4(2), 73–79. https://doi.org/10.1027/1614-2241.4.2.73
Maydeu-Olivares, A., Kramp, U., García-Forero, C., Gallardo-Pujol, D., & Coffman, D. (2009). The effect of varying the number of response alternatives in rating scales: Experimental evidence from intra-individual effects. Behavior Research Methods, 41(2), 295–308. https://doi.org/10.3758/BRM.41.2.295
Mellenbergh, G. J. (1996). Measurement precision in test score and item response models. Psychological Methods, 1(3), 293–299. https://doi.org/10.1037/1082-989X.1.3.293
Nunnally, J. C., & Bernstein, I. H. (1994). Psychometric theory. McGraw-Hill. https://doi.org/10.2307/1161962
Oppenheim, A. N., & Torgerson, W. (1961). Theory and methods of scaling. In The British Journal of Sociology (Vol. 12). John Willey & Sons. https://doi.org/10.2307/588039
Posit Team. (2023). RStudio: Integrated development environment for R. Boston, MA: Posit Software, PBC. Retrieved from http://www.posit.co/
Preston, C. C., & Colman, A. M. (2000). Optimal number of response categories in rating scales: Reliability, validity, discriminating power, and respondent preferences. Acta Psychologica, 104(1), 1–15. https://doi.org/10.1016/S0001-6918(99)00050-5
Price, L. R. (2017). Psychometric methods, theory into practice. Guilford Press.
Qasem, M. A. N., Almoshigah, T. S., & Gupta, S. (2014). The effect of number of alternatives on validity and reliability in Likert scale. International Journal of Innovative Research & Studies, 3(6), 324–333. https://doi.org/10.13140/2.1.2237.2803
Retnawati, H. (2014). Teori respons butir dan penerapannya: Untuk peneliti, praktisi pengukuran dan pengujian, mahasiswa pascasarjana [Item response theory and its applications: For researchers, measurement and testing practitioners, & graduate students]. Yogyakarta: Nuha Medika.
Retnawati, H. (2016). Analisis kuantitatif instrumen penelitian: Panduan peneliti, mahasiswa, dan psikometrian [Quantitative analysis of research instruments: Guide for researchers, students, and psychometricians]. Yogyakarta: Parama Publishing.
Robinson, S. (2004). Simulation: The practice of model development and use. West Sussex, UK: John Wiley & Sons.
Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores (psychometric monograph no. 17). Psychometrika, 34(1), 1–100. https://doi.org/10.1007/BF02290599
Santoso, P. H., Setiawati, F. A., Ismail, R., & Suhariyono, S. (2023). Comparing IRT properties among different category numbers: A case from attitudinal measurement on physics education research. Discover Psychology, 3(1), 39. https://doi.org/10.1007/s44202-023-00101-6
Sharma, H. (2022). How short or long should be a questionnaire for any research? Researchers dilemma in deciding the appropriate questionnaire length. Saudi Journal of Anaesthesia, 16(1), 65–68. https://doi.org/10.4103/sja.sja_163_21
Taherdoost, H. (2019). What is the best response scale for survey and questionnaire design: Review of different lengths of rating scale/attitude scale/Likert scale. International Journal of Academic Research in Management, 8(1), 2296–1747.
Tarka, P. (2015). Likert scale and change in range of response categories vs. The factors extraction in EFA model. Acta Universitatis Lodziensis: Folia Oeconomica, 1(311), 27–36. https://doi.org/10.18778/0208-6018.311.04
Wakita, T., Ueshima, N., & Noguchi, H. (2012). Psychological distance between categories in the likert scale. Educational and Psychological Measurement, 72(4), 533–546. https://doi.org/10.1177/0013164411431162
Weng, L.-J. (2004). Impact of the number of response categories and anchor labels on coefficient alpha and test-retest reliability. Educational and Psychological Measurement, 64(6), 956–972. https://doi.org/10.1177/0013164404268674
Willse, J. T. (2018). CTT: Classical test theory functions. Retrieved from https://cran.r-project.org/package=CTT
Published
How to Cite
Issue
Section
Citation Check
License
Copyright (c) 2026 REID (Research and Evaluation in Education)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
The authors submitting a manuscript to this journal agree that, if accepted for publication, copyright publishing of the submission shall be assigned to REID (Research and Evaluation in Education). However, even though the journal asks for a copyright transfer, the authors retain (or are granted back) significant scholarly rights.
The copyright transfer agreement form can be downloaded here: [REID Copyright Transfer Agreement Form]
The copyright form should be signed originally and sent to the Editorial Office through email to reid.ppsuny@uny.ac.id

REID (Research and Evaluation in Education) by http://journal.uny.ac.id/index.php/reid is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.



.png)




