Hello everyone.
I would like to introduce a paper we read at the recent English seminar.
Paper Title: Exploring ChatGPT as a writing assessment tool
Journal: Innovations in Education and Teaching International
Publication Year: 2024
Authors: Junifer Leal Bucol, Napattanissa Sangkawong
Abstract
This study examined the potential of using ChatGPT as an Automated Writing Evaluation (AWE) tool in an English course at a university in Thailand. Using ChatGPT, students’ writing was graded based on pre-prepared prompts and rubrics, and the results were compared with those of human evaluators. Furthermore, the study qualitatively analyzed teachers’ reflections on the evaluation process to clarify the strengths and weaknesses of ChatGPT as an assessment tool.
1. Introduction
While writing is a critical process in language acquisition, assessing student writing and providing feedback has long been a challenge for educators. With increasing class sizes and time constraints, there is a growing need for objective and efficient writing assessment systems. Traditional teacher-led assessment is prone to subjectivity and potential inconsistency, leading to increased interest in standardized rubrics and the introduction of AWE as a means to enhance the consistency, fairness, and efficiency of assessment.
Evaluating AWE: Benefits and Challenges
The benefits of AWE include immediate feedback and grading for students, the promotion of iterative learning, convenience and efficiency, increased student engagement in writing and revision, and improved writing accuracy. However, some AWE software cannot fully replicate the nuanced grading capabilities, critical thinking, or creativity of human evaluators. There are also concerns regarding assessment accuracy and challenges related to implementation, such as costs. Some argue that AWE is most effective when used in conjunction with teacher feedback and should not be used as a replacement.
The research gaps are identified as the following two points:
• While there is a significant body of research on AWE, there are few empirical studies validating ChatGPT as an AWE tool.
• There is almost no research comparing human evaluation using customized rubrics.
2. Research Questions
This study established the following research questions to compare the student writing assessment capabilities of ChatGPT and human evaluators, and to identify the strengths and challenges of using ChatGPT as an assessment tool:
RQ1: To what extent can ChatGPT accurately assess short essays using a pre-designed analytical rubric compared to human evaluators?
RQ2: What are the strengths and weaknesses of using ChatGPT as a writing assessment tool?
3. Methodology
This study employed an exploratory methodology combining quantitative and qualitative approaches. Quantitative data were obtained from the scores of essays (Topic: “My Favorite Place,” 90-250 words) written by 10 students (CEFR A2-B1), which were evaluated by ChatGPT and human evaluators (10 EFL instructors working at a Thai university). The instructors were divided into the following two groups:
① Group 1 (5 instructors using ChatGPT for evaluation)
② Group 2 (5 human evaluators)
Both groups conducted evaluations based on a rubric containing five customized assessment criteria (Task Achievement, Grammar Usage, Vocabulary Selection, Coherence, and Mechanics). Reliability tests and correlation analyses were performed on the scores using SPSS. As qualitative data, the participating instructors recorded their observations during the evaluation process using ChatGPT.
4. Results and Discussion
Although there was some variation in the scores obtained from ChatGPT and human evaluators, the scores generated by ChatGPT were relatively high and showed a consistent trend.
In the analysis using Cronbach’s alpha coefficient to assess the internal consistency of the entire dataset, ChatGPT scores showed α = .980, human evaluator scores showed α = .926, and the overall score was α = .954, all of which were high values, suggesting consistency and reliability in the overall scoring between evaluators.
Intraclass Correlation Coefficient (ICC) analysis yielded a p-value of less than .001, confirming a statistically significant and consistent relationship between evaluators. Pearson correlation analysis also showed a strong positive correlation, particularly among scores generated by ChatGPT. Notably, the consistency rate was highest between “ChatGPT and ChatGPT,” followed by “ChatGPT and human,” and then “human and human.”
Qualitative data from instructor observations also revealed challenges with ChatGPT in the evaluation process.
For example,
• While consistent evaluation based on rubrics is possible, there are instances where fine details are overlooked.
• In terms of text comprehension, it excels at grasping themes and main points, but has difficulty understanding complex content.
• Regarding corrective feedback, it provides errors and improvement suggestions, but rarely mentions the root causes of the errors.
• While rapid evaluation and continuous feedback are possible, periodic monitoring and adjustment are required.
The results of this study indicate that although there was some variation in scores, ChatGPT showed high internal consistency and significant correlation between evaluators, suggesting that ChatGPT has the potential to be a reliable assessment tool. The benefit of scalability, which allows for the efficient grading of a large number of essays, was also confirmed. However, teachers must use the technology with caution as the AI may tend to be lenient with scores. Other challenges include a tendency to ignore specific requirements of some rubric criteria and a lack of information regarding the causes of identified errors.
5. Conclusion
ChatGPT demonstrated the ability to evaluate writing as an AWE tool and is considered useful in terms of efficiency, consistency, speed, and scalability. On the other hand, the challenges of ChatGPT must also be recognized. These include errors in interpreting specific information within essays, an insufficient ability to understand complex and creative writing styles, and the fact that the comments provided do not cover all aspects of writing. The limitations of this study include the use of a small number of evaluators and writing samples, and the restriction to the evaluation of short essays. Future research needs to examine AI-human collaborative assessment targeting more complex tasks.
An approach that combines automated assessment with human review can improve the accuracy of assessment, especially in writing that requires detailed evaluation. By integrating ChatGPT with human expertise, it is possible to effectively reinforce the evaluation process and provide beneficial support for improving students’ writing skills.
6. Reflections
This study is a practical attempt to explore the potential of using ChatGPT as an AWE tool, and I believe it is possible to implement it at the individual teacher level without the need for special technical equipment. The analysis results confirm the consistency of evaluation between humans and AI, suggesting its utility in terms of streamlining grading and improving reliability. On the other hand, I believe this study has several limitations.
First, the focus is solely on the efficiency and consistency of the grader, and there is a lack of perspective on how to return the evaluation to the learner and utilize it for subsequent learning. Also, although a customized rubric is used, some of the published evaluation criteria include quantitative elements based on word count (e.g., in “Task Response,” an essay under 100 words = 0.5 points, 200 words or more = 2 points). With such criteria, there is a possibility that the AI judges the score based on word count without considering the qualitative elements of the essay. I believe it is necessary to fully verify the validity and applicability of using rubrics in AWE.
Furthermore, the consistency rate between AI and human evaluators is analyzed based only on the “total score” of the rubric, and it is not specified in which aspects (e.g., grammar, logic) the discrepancies occurred. I believe that further analysis by aspect is necessary to aim for the qualitative improvement of AWE.
Contributor: Sayo Tanaka




