Hello everyone. This is Tanaka, a first-year doctoral student.
I would like to introduce a paper I read at our recent English seminar.
Paper Title: Large language models fall short in classifying learners’ open-ended responses
Journal: Research Methods in Applied Linguistics
Year of Publication: 2025
Authors: Atsushi Mizumoto, Mark Feng Teng
1. Introduction
Large Language Models (LLMs) are being applied in various fields as tools capable of generating and analyzing text like humans, and in applied linguistics, their use in analyzing open-ended responses and interview data is progressing. In recent years, the high linguistic processing capabilities of LLMs have been demonstrated, and attempts are being made to introduce them into tasks such as classification and coding. On the other hand, there are limitations in understanding context and grasping relationships between concepts, and the necessity of collaboration with humans has been pointed out. This study aims to compare the classification accuracy of open-ended responses by LLMs with human judgment and to clarify the possibilities and challenges of using LLMs in qualitative research.
2. Theoretical Framework
When classifying open-ended data in qualitative research, it is required to ensure reliability and validity through clear category definitions, careful training, and repeated consensus-building among human coders. Researchers consider context and nuances of expression when assigning codes, and proceed with analysis based on coding manuals and theoretical models. In this study, the three processes of Self-Regulated Learning (SRL)—”planning,” “monitoring,” and “evaluating”—were classified into categories based on the questionnaire by Teng et al. (2022).
3. Research Question
The research question of this study is, “How accurately can LLMs classify open-ended data?” Specifically, it aims to verify the classification accuracy of LLMs by comparing it with human coders in a classification task based on the definitions of “planning,” “monitoring,” and “evaluating.”
4. Methodology
The subjects of this study were 143 first-year English majors (CEFR B1–B2 level) enrolled at a private university in Japan. Students were asked to write a single sentence describing how they approach writing an essay in English, and their responses were collected. The responses were translated into English and classified by two researchers with PhDs in applied linguistics, and back-translation was also performed. Classification was carried out according to the three categories of “planning,” “monitoring,” and “evaluating” based on Zimmerman’s (2000) SRL theory. If multiple elements were included, they were classified based on the most prominent process, and in cases of disagreement, the first author joined to reach a consensus.
Analysis by LLM
In this study, seven LLMs (GPT-4o, GPT-o1, GPT-o3mini, Llama3.3–70B, Gemini2.0-Flash, Claude3.5-Sonnet, and DeepSeek-V3) were used. Each model was accessed via API and performed the same classification task. The models were instructed to classify the 143 open-ended responses into one of “planning,” “monitoring,” or “evaluating.”
The accuracy of the LLM classification was evaluated using “simple agreement rate” and “Cohen’s kappa coefficient.”
Prompt Design
Initially, zero-shot classification presenting only definitions was attempted, but the accuracy was insufficient, so a structured prompt including category definitions and concrete examples was created. The prompt includes the following three points:
① Clear definition of each category
② Concrete response examples
③ Specification of output format (category name only)
This structure improved classification accuracy and output consistency.
5. Results
Comparison between models
・The model that showed the highest agreement rate with human coders was DeepSeek-V3, with an agreement rate of 83.2% and κ = 0.68.
・The runners-up were Llama3.3–70B (κ = 0.61) and GPT-o3mini (κ = 0.60), showing moderate agreement.
・GPT-4o, GPT-o1, Gemini2.0-Flash, and Claude3.5-Sonnet were judged to have weak agreement with κ = 0.37–0.49.
・Open-source models (DeepSeek-V3, Llama3.3–70B) showed higher accuracy than commercial models.
Trends in Misclassification
The main patterns of misclassification are as follows:
① Confusion between planning and monitoring
Example: “Write everything out and then check” → Originally corresponds to monitoring, but many models classified it as planning.
② Ambiguous interpretation of “revision”
“Revision during writing” and “review after writing” are confused, and there is a tendency to classify as evaluation based solely on the word “revision.”
③ Over-reliance on syntax
If there are expressions such as “first… then…”, it is easily judged as planning automatically.
6. Discussion
In this study, we verified how accurately LLMs can classify open-ended responses according to pre-defined categories. As a result, DeepSeek-V3 (κ = 0.68) and Llama3.3–70B (κ = 0.61) showed moderate agreement, but did not reach κ ≥ 0.8, which is considered a highly reliable classification standard. This reveals that it is difficult for current LLMs to make judgments equivalent to humans in context-dependent classification tasks like open-ended responses. On the other hand, LLMs are capable of consistent preliminary classification, and it is thought that combining this with subsequent human review can enable efficient and reliable qualitative analysis.
Differences between LLM and human judgment
A clear difference between LLM and human classification judgment is that while humans interpret the learner’s intention based on context and common sense, LLMs judge based on superficial linguistic patterns. For example, if there is an expression like “check after writing,” it is easily classified as “evaluation” regardless of the content. Also, syntax like “first… then…” tended to be automatically interpreted as “planning.”
Direction for prompt improvement
The following points are considered for future improvement:
・Few-shot prompting: In addition to typical examples, present examples that are easily misclassified.
・Chain-of-thought prompting: Prompt the thought process of classification step-by-step.
・Contrastive examples: Show examples where similar sentences result in different classifications.
・Utilization of confidence scores: Prompt human re-confirmation for ambiguous judgments.
Factors for performance differences between models
The reason why open-source models had higher accuracy than commercial models is thought to be related to the latest architecture, high-quality training data, and designs like Mixture-of-experts. On the other hand, it also became clear that even models excellent at explanation generation and dialogue, such as Claude3.5-Sonnet, can be unsuitable for structured classification tasks.
7. Implications and Future Prospects
This study provides the following practical implications when using LLMs for qualitative analysis:
・Usefulness as an auxiliary tool: Especially in large datasets, it is possible to reduce the workload.
・Pre-verification is necessary: LLM classification results must be confirmed by comparing them with human coding.
・Ensuring transparency: It is necessary to clearly state that an LLM was used and ensure reproducibility.
・Reliance on superficial judgment: LLMs are more easily influenced by the form of expression than the meaning of sentences, and human judgment is indispensable in ambiguous cases.
Future research could take the following directions:
・Advanced prompt design (introduction of few-shot, chain-of-thought, contrastive examples, etc.)
・Identification of ambiguous judgments using confidence scores
・Fine-tuning models for specific fields
・Performance comparison across diverse datasets and classification frameworks
Through these, the construction of a hybrid classification support system where LLMs and humans collaborate is expected.
8. Impressions
The reason I chose this paper was to gain empirical knowledge for designing more accurate prompts when using LLMs in language education, and to deepen my methodological understanding when qualitatively coding open-ended responses. In this study, seven LLM models were compared, which was also useful for grasping the latest trends in generative AI. In addition, the data and prompts used in the research are made public, contributing to ensuring the transparency and reproducibility of the research. I believe this study provides important insights for understanding the current status of using LLMs in educational settings and research practices, and for clarifying the issues that should be addressed in the future.
A challenge is that the open-ended responses to be coded were all limited to “one sentence.” Due to this constraint, LLM output is likely to rely on superficial linguistic cues such as word order patterns like “if” and “then,” and it is thought that the accuracy of meaning understanding and judgment based on context has not been sufficiently verified. Since it is common for learners’ open-ended responses to be composed of compound sentences or paragraphs in actual educational settings, I believe it is necessary to verify the classification accuracy of LLMs for descriptions including implications and complex structures in the future.




