Comparative analysis of language models in automated educational diagnostic tasks

Authors

DOI:

https://doi.org/10.65835/aj.2026.2.13

Keywords:

Automated assessment, large language models, educational diagnosis, formative feedback

Abstract

This study presents a comparative analysis of language models for automated educational diagnosis, evaluating two specialized models (a feature-based ML model and a fine-tuned BERT) against three state-of-the-art Large Language Models (GPT-4, Claude 3, Gemini Pro). Using a mixed-methods approach on a corpus of 450 annotated student texts, the research assessed diagnostic accuracy, feedback quality, robustness, and hallucination rates. Results indicate that advanced LLMs, particularly GPT-4 and Claude 3 Opus, achieve superior diagnostic performance (F1-Score ≥ 0.83) and generate pedagogically higher-quality feedback, rated as more useful and actionable by expert judges. However, these models exhibit a critical weakness: a significant propensity for factual hallucinations (5.2% - 12.3%). In contrast, specialized models, while less sophisticated, offer near-perfect reliability and greater cost-efficiency. The study concludes that LLMs represent a qualitative leap in automated assessment but are not yet reliable for autonomous application. The most viable implementation model is AI-augmented education, where LLMs act as diagnostic assistants to teaching professionals, who provide essential oversight, validation, and personalization. This human-in-the-loop paradigm balances technological potential with pedagogical integrity and ethical responsibility.

Downloads

Download data is not yet available.

References

Creswell, J. W., & Plano Clark, V. L. (2018). Designing and conducting mixed methods research (3rd ed.). SAGE Publications.

Lillo-Fuentes, F., Venegas, R., & Lobos, I. (2023). Evaluación automatizada y semiautomatizada de la calidad de textos escritos: Una revisión sistemática. Perspectiva Educacional, *62*(2), 5–36. https://www.redalyc.org/journal/3333/333379712002/html/

Salvagno, M., Taccone, F. S., & Gerli, A. G. (2023). Can artificial intelligence help for scientific writing? Critical Care, *27*(1), 75. https://doi.org/10.1186/s13054-023-04380-2

Shermis, M. D. (2020). The evolution of automated essay scoring. En Handbook of automated essay evaluation (pp. 1–15). Routledge.

Shyr, C., Grout, R., Kennedy, N., Akdas, Y., Tischbein, M., Milford, J., ... Harris, P. (2024). Evaluating large language models in generating and explaining diagnoses. Journal of the American Medical Informatics Association, ocae186. https://doi.org/10.1093/jamia/ocae186

Sociedad Suiza de Informática Médica. (2023). Los grandes modelos lingüísticos (LLM) en la práctica clínica diaria. Medizin Online. https://medizinonline.com/es/dr-chatgpt-grandes-modelos-lingueisticos-en-la-practica-clinica-diaria/

Downloads

Published

14.03.2026

How to Cite

Nagua Velepucha , M. T. . (2026). Comparative analysis of language models in automated educational diagnostic tasks. Applied Journal of Digital Innovation, 2, 13. https://doi.org/10.65835/aj.2026.2.13