Hello everyone! I am Naohiro Higuchi, a second-year master’s student at Yamada Laboratory.
With the spread of personalized learning and Learning Analytics (LA), we have entered an era where “data literacy”—the ability to interpret data—is indispensable. But how can we objectively evaluate and measure this skill?
In this post, I will introduce a review paper that systematically analyzes evaluation methods used across various fields worldwide. This paper provided significant hints for the evaluation design of my own research, “Analysis of Teacher Evaluation and Behavioral Logs.”
Paper Information
・Title: Data literacy assessments: a systematic literature review
・Authors: Ying Cui, Fu Chen, Alina Lutsyk, Jacqueline P. Leighton & Maria Cutumisu
・Journal: Assessment in Education: Principles, Policy & Practice, 30(1), 76-96.
・Year of Publication: 2023
・DOI: https://doi.org/10.1080/0969594X.2023.2182737
1. Introduction and Background
With the advancement of 21st-century technology, the amount of data available to us in our daily lives and work has exploded. Consequently, the importance of “data literacy skills”—the ability to access data and correctly understand its meaning—is rapidly increasing across society.
This trend is the same in educational settings, where movements to utilize data for educational improvement are flourishing. However, it has been pointed out as a major issue that data literacy assessments conducted in different educational environments and for different purposes have not been systematically reviewed.
To address this issue, the authors set the following five research questions (RQs) with the aim of clarifying the commonalities, unique features, and respective strengths and weaknesses of existing data literacy assessments, and providing guidelines for future evaluation design.
RQ1: Who are the target audiences for each data literacy assessment?
RQ2: What definitions of data literacy were adopted in each study?
RQ3: Which data literacy competencies were focused on in each evaluation?
RQ4: What are the evaluation formats and item types of each assessment tool?
RQ5: Do the evaluations provide evidence of reliability and validity through verification using actual data?
2. Methodology
In this study, to cover the multifaceted nature of data literacy (which spans education, social sciences, engineering, etc.), search keywords and databases were selected by referring to five previous review papers.
Searches were conducted using six global international databases: PsycINFO, ERIC, Scopus, IEEE Xplore, SpringerLink, and Web of Science. From literature published between 2000 and May 2021, 616 items were extracted as initial data.
After applying strict inclusion and exclusion criteria, such as “whether the study actually conducts data literacy evaluation using an empirical approach,” 25 papers were finally selected for review. Information was extracted from these papers based on pre-set categories, then integrated and summarized.
3. Results and Discussion
The analysis of the review highlighted the current state of data literacy assessment and future challenges. The results and discussion are organized according to the five RQs.
RQ1: Who are the target audiences for the assessment?
Results:
The target audiences for the assessments were mainly classified into the following five groups:
・K-12 (Kindergarten to High School) students: 10 cases (middle and high school students were the primary focus)
・Teachers and educational professionals (including pre-service teachers): 7 cases
・Researchers and data librarians: 4 cases
・Higher education (university students): 2 cases
・Other professionals: 2 cases
Discussion:
The most common target group was K-12 students (especially at the secondary education level), followed by assessments targeting teachers. However, it became clear that there is extremely little research targeting “higher education levels such as university students,” which should be crucial for bridging the skills gap in the labor market.
RQ2: What definitions of “data literacy” are adopted?
Results: There were clear differences in definition trends depending on the target reader community.
・Teachers and educational professionals: The definition by Gummer and Mandinach (2015) was mainly adopted. This refers to the ability to collect, analyze, and interpret “all kinds of data”—not just test scores, but also school climate, behavioral data, longitudinal data, and ad-hoc data—and convert them into actionable instructional knowledge and practices to determine teaching steps. The cognitive elements of this definition are highly multifaceted, consisting of three interactive domains (data use for teaching, content knowledge, and pedagogical content knowledge) and six inquiry cycles (identifying problems, framing questions, using data, transforming data into information, transforming information into decisions, and evaluating results), incorporating 59 subdivided elements related to knowledge and skills.
・Students (K-12 and Higher Education): There was a strong tendency for authors of each study to create their own definitions, often defined as the ability to collect, analyze, and interpret data for field-specific or real-world problem solving.
・Researchers and data librarians: Data literacy was positioned as part of information literacy or digital literacy, and tended to be defined as more practical and granular “research data management” skills (file naming conventions, data storage, sharing, etc.).
Discussion:
Overall, it is evident that definitions emphasizing the “inquiry process”—identifying problems, collecting and analyzing data, and applying it to decision-making and problem-solving—are widely used, rather than just isolated skills like “reading a graph.”
RQ3: Which data literacy competencies were focused on?
Results:
・Teachers and educational professionals: The focus was on how to utilize data in the instructional process and classroom decision-making.
・Students: “Data visualization,” which is directly linked to problem-solving, was frequently emphasized, and in higher education, “critical thinking” was also prioritized.
・Researchers and librarians: Practical data management skills, such as standard file naming and secure data storage and collaborative sharing, were emphasized.
・Other professionals: Competencies specific to each profession were required, such as journalists writing articles based on data.
Discussion:
Because data literacy contains a vast number of complex elements (mathematical knowledge, statistics, critical thinking, etc.), it is suggested that trying to cover all of these with a single evaluation tool leads to test tasks that are superficial and unnatural.
RQ4: What are the “formats” and “item types” of the evaluations?
Results:
Evaluation methods were broadly classified into “self-reflective approaches (8 cases)” and “objective measurements (17 cases).”
Self-reflective approaches: Methods where individuals subjectively evaluate their own data skills, such as questionnaires, surveys, semi-structured interviews, and think-aloud protocols.
Objective measurements: Methods where others objectively measure ability, such as traditional paper tests (multiple-choice/open-ended), digital game-based assessments, and participant observation tracking classroom behavior.
Discussion:
“Objective measurement” was most commonly used in evaluations for teachers and K-12. However, most of these were traditional approaches like multiple-choice or open-ended tests, and advanced, practical digital evaluation methods that realistically handle real-world data and problems remain limited to a few attempts.
RQ5: Is the verification of reliability and validity sufficiently conducted?
Results:
Only 3 studies (Larasati & Yunanta, 2020; Pratama et al., 2020; Wu et al., 2021) empirically verified the reliability and validity of the evaluation tools themselves.
Discussion:
The majority of studies had the primary purpose of measuring the effectiveness of some training program or intervention, and did not pay attention to verifying whether the tool used for that measurement actually measured the intended ability. The three studies that did conduct verification used advanced test theories such as Item Response Theory (IRT) and cognitive diagnostic models to scientifically confirm validity, but overall, the current situation is that quality verification of evaluation tools is still lacking.
4. Conclusion and Limitations
The biggest lesson learned from this study is that “in data literacy education, the development of high-quality, verified assessment tools is still in its infancy.”
Previous research has been mainly biased toward the secondary education level, and evaluation tools lack versatility because they are tightly coupled with specific subjects. Furthermore, they are biased toward multiple-choice tests and subjective questionnaires, failing to fully capture the multifaceted skills that data literacy possesses.
Limitations of the Study
Limitations of this review include the small number of target papers (25), many of which are small-scale initial studies, and the fact that it covers literature only up to May 2021. However, its value as a roadmap for organizing challenges in future evaluation tool development is extremely high.
5. Reasons for Selection and Impressions
Currently, I am developing a “class support and reflection system for teachers in the field.” One of the biggest challenges in this project was “how to evaluate the transformation of the subjects (teachers) who use the system.” While considering evaluation methods tailored to the context of teachers’ data literacy, I reviewed this paper to comprehensively grasp recent assessment trends.
By reading this paper, I gained great confidence and new perspectives for the evaluation design I am currently planning.
The other day, I formulated “what data to collect and how to analyze it” for ethical review. In my research, I plan to combine not only the implementation of questionnaires (subjective surveys) for subjects, but also actual classroom observation and analysis of teachers’ behavioral logs within the app (which data they viewed and how they moved).
The insight pointed out in this paper—that “self-reflective approaches alone have limitations in verifying validity, and objective behavioral observation and measurement should be combined”—coincides with the evaluation policy I am aiming for. I am confident that if this evaluation method, which combines multifaceted behavioral logs during class and classroom observation, functions well, it could be a very valuable step in the field of teacher data literacy research and, by extension, Educational Technology.
While taking into account the educational level of the evaluation target and the challenges of tool development, I will strive toward future system development and implementation evaluation so that it becomes a system that is truly easy for teachers in the field to use and encourages reflection!
Written by: Naohiro Higuchi




