The role of AI in the future of assessment makes the most sense when viewed against the full history of educational testing, because every major shift in assessment has followed a familiar pattern: new tools promise better measurement, institutions adapt slowly, and educators eventually discover that the value of a test depends less on novelty than on validity, fairness, and usefulness. Educational assessment refers to the systematic process of gathering evidence about what learners know, can do, and are ready to learn next. Educational testing is one part of that process, usually involving standardized tasks scored under defined conditions. Artificial intelligence adds another layer by using machine learning, natural language processing, computer vision, and predictive analytics to score work, generate questions, detect patterns, and support decision-making.
This matters now because schools, universities, certification bodies, and employers are all under pressure to assess more complex skills at greater scale. In my own work reviewing assessment programs, I have seen the same challenge repeat across contexts: leaders want richer evidence than multiple-choice tests alone can provide, but they also need consistency, speed, and defensible results. AI appears to solve that tension. It can score essays in seconds, adapt item difficulty in real time, flag disengaged test behavior, and help teachers identify misconceptions earlier. Yet AI also inherits longstanding problems from the history of testing, including bias, construct underrepresentation, surveillance concerns, and the temptation to confuse efficient scoring with meaningful learning evidence.
As a hub within Foundations of Educational Assessment, this article covers the history of educational testing comprehensively while showing where AI fits in the next phase. The key terms are essential. Reliability means score consistency. Validity concerns whether evidence supports the intended interpretation and use of scores. Standardization means uniform administration and scoring conditions. Norm-referenced interpretation compares learners with peers, while criterion-referenced interpretation compares performance against defined standards. These concepts shaped paper examinations, intelligence tests, large-scale accountability systems, computer-based testing, and now AI-enhanced assessment. Understanding that progression helps educators make better choices about what should be automated, what should remain human-led, and how future assessment can serve learning rather than simply sort students more efficiently.
From oral examinations to mass schooling: the early history of educational testing
The history of educational testing begins long before modern psychometrics. In ancient China, imperial civil service examinations selected officials through structured written assessments, establishing one of the earliest large-scale examples of standardized testing. In medieval and early modern Europe, universities relied heavily on oral disputations, recitations, and essays to judge mastery. These systems were labor intensive, locally defined, and closely tied to curriculum. They assessed what institutions valued, but they were often inconsistent because scoring depended on examiner judgment. Even so, they reveal an enduring truth: tests have always reflected administrative needs as much as educational philosophy.
The rise of mass schooling in the nineteenth century changed assessment fundamentally. As student populations expanded, education systems needed methods that were more scalable than oral examination. Written exams became common because they allowed multiple students to be assessed at once and created a record for later review. In Britain, civil service reforms and university entrance examinations encouraged more formalized written testing. In the United States, reformers such as Horace Mann promoted common schooling and used comparative test results to argue for accountability. By the late 1800s, testing was becoming an instrument for system management, not just classroom judgment.
This period also introduced a central tension that still shapes AI in assessment today. Once tests are used to compare schools, rank students, or allocate opportunities, efficiency pressures grow quickly. Administrators want scoring procedures that are faster, more uniform, and easier to defend. Teachers often want assessments that capture deeper understanding and support instruction. Those goals overlap, but they are not identical. The modern history of educational testing can be read as a series of attempts to reconcile them, first through standardization, then through statistical measurement, later through digital delivery, and now through AI-supported inference.
Psychometrics, intelligence testing, and the standardization era
The late nineteenth and early twentieth centuries established the scientific language that still governs assessment practice. Psychometrics emerged as the field devoted to measuring mental attributes, using statistical models to study test quality. Alfred Binet and Théodore Simon developed early intelligence scales in France to identify students needing additional support. Their work was later adapted and often overextended, especially in the United States, where intelligence testing became entangled with ranking, selection, and deeply flawed assumptions about fixed ability. The lesson is not that measurement is inherently suspect; it is that measurement becomes dangerous when technical tools are treated as neutral while social assumptions go unexamined.
During this era, standardized achievement tests also expanded. The College Board introduced early forms of admissions testing, and large-scale testing accelerated during World War I with the Army Alpha and Beta tests. These assessments demonstrated how standardized instruments could be administered to very large populations and analyzed statistically. At the same time, they exposed recurring fairness problems related to language, culture, and access. Modern validity theory grew partly in response to those concerns. Test developers learned that a score is not meaningful on its own; it must be interpreted within a documented argument about the construct being measured, the population tested, and the decisions the score will support.
By the mid-twentieth century, core concepts such as reliability coefficients, standard error of measurement, item difficulty, discrimination, and equating had become standard practice. Multiple-choice formats spread because they offered broad content sampling and machine scoring. This was not simply a convenience. Optical mark recognition made statewide and national testing programs feasible. However, widespread use of selected-response items also narrowed what many systems measured well. Knowledge recall and procedural fluency were easier to test consistently than writing quality, collaboration, or scientific inquiry. That narrowing is one reason AI has gained traction: institutions have spent decades wanting broader evidence without sacrificing comparability.
Accountability, performance assessment, and the move to digital testing
From the 1980s onward, educational testing became more tightly linked to policy accountability. In the United States, standards-based reform, No Child Left Behind, and later the Every Student Succeeds Act made annual testing central to school evaluation. International programs such as PISA, TIMSS, and PIRLS expanded comparative measurement across countries. These programs generated useful trend data, but they also intensified teaching-to-the-test concerns. When stakes rise, assessment shapes curriculum, pacing, and classroom time. That effect is not always harmful; clear standards can improve coherence. Yet high stakes can also distort instruction if test formats reward narrow performance proxies over richer learning goals.
In response, many systems renewed interest in performance assessment. Portfolios, extended writing, labs, exhibitions, and capstone projects promised more authentic evidence of student competence. I have seen districts adopt these approaches with strong instructional benefits, especially when common rubrics and moderation protocols are in place. The challenge has always been scale. Human scoring is expensive, slower, and vulnerable to drift unless raters are trained and recalibrated. Digital platforms helped by streamlining administration, capturing student process data, and supporting distributed scoring, but they did not eliminate tradeoffs between authenticity and efficiency.
Computer-based testing changed the infrastructure again. Online delivery enabled multimedia items, simulation tasks, faster reporting, and adaptive testing. Item response theory became more operationally important because it supports computerized adaptive testing, where item selection adjusts to student performance in real time. The Graduate Record Examination and many licensure exams demonstrated that adaptive designs could shorten tests while preserving measurement precision. Once assessment moved online, AI was the next logical layer. Systems now had item banks, response data, keystroke logs, timing patterns, and scalable computing. The question shifted from whether technology could assist assessment to how far automation should extend into judgment.
How AI is changing assessment design, scoring, and feedback
AI is reshaping assessment in four major areas: task design, scoring, personalization, and quality control. In task design, generative models can draft items aligned to standards, produce reading passages at controlled complexity levels, and create parallel forms more quickly than traditional workflows. Good programs still require human review for content accuracy, bias, accessibility, and blueprint fit, but development cycles are shorter. In scoring, automated essay scoring systems such as e-rater and IntelliMetric have shown that machine scoring can reach useful agreement with trained human raters on constrained writing tasks. Similar methods support short-answer scoring through natural language processing and semantic matching.
Personalization is another major shift. Adaptive platforms already adjust difficulty; AI extends this by diagnosing likely misconceptions, recommending next tasks, and generating formative feedback. In mathematics, systems can identify error patterns such as place-value confusion or sign errors. In language learning, speech recognition can analyze pronunciation, fluency, and vocabulary usage. In simulation-based assessment, AI can infer strategy use from clickstreams, sequence choices, and time allocation. This opens the possibility of measuring process, not just final answers. For teachers, that can be genuinely useful because intervention becomes more targeted.
Quality control may become AI’s most underappreciated contribution. AI can detect anomalous item behavior, possible item leakage, unusual response patterns, and scoring drift across raters or sites. It can help flag differential item functioning for further review, though final fairness decisions must remain with trained psychometricians. The table below shows how major assessment eras compare.
| Era | Primary method | Main strength | Main limitation | AI implication |
|---|---|---|---|---|
| Pre-modern | Oral exams, recitations | Rich judgment of understanding | Low consistency and scale | AI may support richer evidence at scale |
| Standardization era | Written and multiple-choice tests | Efficiency and comparability | Narrow skill coverage | AI can expand assessable constructs |
| Accountability era | Large-scale standardized testing | Trend data and policy visibility | High-stakes distortion risks | AI can improve reporting but may amplify misuse |
| Digital era | Computer-based and adaptive tests | Faster scoring and precision | Infrastructure and access gaps | AI builds on existing data-rich platforms |
| Emerging future | AI-assisted multimodal assessment | Personalized feedback and broader evidence | Bias, privacy, and transparency concerns | Human oversight remains essential |
The risks, limits, and governance rules that will define responsible use
AI in assessment is valuable only when governance is stronger than enthusiasm. The first risk is validity drift. A model may score features correlated with quality rather than the intended construct itself. For example, essay systems can overweight length, formulaic structure, or surface fluency unless carefully trained and audited. The second risk is bias. If training data reflect historical inequities, model outputs may disadvantage multilingual learners, students with disabilities, or groups underrepresented in the source data. The third risk is opacity. Black-box scoring undermines trust when educators cannot explain why a response received a given score.
There are also operational risks. Remote proctoring tools using computer vision have produced false flags related to lighting, skin tone, eye movement, and assistive technology use. Predictive models that estimate success, dropout risk, or proficiency can harden expectations if treated as destiny rather than probabilistic signals. Data privacy is another major concern because AI systems often depend on large volumes of student data, including writing samples, audio, video, and behavioral traces. In education, consent, retention rules, cybersecurity, and vendor contracts are not side issues. They are core assessment design requirements.
Responsible implementation follows established principles. Start with a clear construct definition and use argument. Require human review for high-stakes decisions. Audit subgroup performance regularly. Conduct bias and accessibility testing before launch, not after complaints. Document training data sources, model versioning, and score interpretation limits. Align practices with recognized standards such as the Standards for Educational and Psychological Testing, published by AERA, APA, and NCME, and with privacy obligations under laws such as FERPA where applicable. The future of assessment will not be decided by AI capability alone. It will be decided by whether institutions can pair technical innovation with disciplined evidence, transparent policy, and educator judgment.
The history of educational testing shows that assessment evolves when institutions need broader reach, better comparability, or richer evidence of learning. Oral examinations emphasized expert judgment. Standardized tests delivered scale. Psychometrics brought statistical discipline. Accountability systems expanded policy influence. Digital platforms increased speed and adaptivity. AI now enters this lineage as the next operational and methodological shift, but it does not erase earlier lessons. Every assessment system still stands or falls on validity, reliability, fairness, and fitness for purpose.
For educators and leaders, the practical takeaway is simple. Use AI where it improves assessment quality, not merely where it reduces labor. It is well suited to item generation support, low-stakes formative feedback, scoring assistance with strong auditing, and pattern detection across large datasets. It is not a substitute for clear learning goals, expert rubric design, accessibility planning, or human accountability in high-stakes decisions. The strongest assessment programs will combine machine efficiency with professional judgment and explicit governance.
As this hub for the history of educational testing within Foundations of Educational Assessment, the article provides the context needed to evaluate every related topic in the subpillar, from early examinations and intelligence testing to standardization, accountability, digital delivery, and emerging AI methods. If you are building, selecting, or revising assessments, start by tracing the historical problem your test is trying to solve, then ask whether AI truly improves the evidence you gather. That question leads to better assessment decisions and better learning outcomes.
Frequently Asked Questions
1. What role is AI likely to play in the future of educational assessment?
AI is likely to play a significant but carefully bounded role in the future of educational assessment. At its best, it can help educators gather richer evidence of student learning, score certain kinds of responses more efficiently, identify patterns that may be difficult to see manually, and support faster feedback cycles. For example, AI can assist with scoring short written responses, analyzing common misconceptions across large groups of learners, generating practice items aligned to specific skills, and adapting question difficulty based on student performance. These capabilities can make assessment more responsive and potentially more useful for instruction.
However, the deeper lesson from the history of testing is that new tools do not automatically produce better assessment. Every major shift in educational measurement has promised greater precision or efficiency, but the real value of any assessment still depends on validity, fairness, and usefulness. In other words, the central question is not whether AI is advanced, but whether it helps educators make sounder judgments about what learners know and can do. If an AI-supported assessment cannot validly measure the intended skill, treats groups of students inequitably, or produces results that are difficult to interpret and act on, then its technological sophistication does not matter.
In practice, AI is most likely to augment rather than replace human judgment. Teachers, assessment designers, psychometricians, and school leaders will still be responsible for deciding what should be measured, how evidence should be interpreted, and what consequences should follow from assessment results. AI may become a powerful tool within that process, but the future of assessment will still hinge on human decisions about quality, ethics, and educational purpose.
2. Can AI make assessments more fair and accurate, or does it create new risks?
AI can do both. It has the potential to improve fairness and accuracy in some contexts, but it also introduces serious risks that educational institutions must address directly. On the positive side, AI systems can help standardize scoring, reduce some forms of human inconsistency, and provide multiple opportunities for students to demonstrate learning. AI can also support more accessible assessment experiences by offering features such as text-to-speech, translation assistance, alternative item formats, and adaptive pathways that better match student readiness levels. When thoughtfully designed, these tools may help reduce barriers that have historically limited some learners’ ability to show what they know.
At the same time, AI systems can reflect or amplify bias if they are trained on incomplete, unrepresentative, or historically skewed data. An automated scoring model, for instance, may privilege certain language patterns, cultural references, or writing styles that do not necessarily reflect stronger understanding. Similarly, predictive systems may overstate risk for some student groups if they rely on variables that correlate with disadvantage rather than actual learning. These issues are especially concerning in high-stakes settings, where flawed outputs can affect placement, advancement, admissions, or support decisions.
That is why fairness in AI assessment cannot be assumed; it must be demonstrated. Institutions need ongoing validation studies, bias audits, transparency about how systems work, and clear policies for human review and appeal. Accuracy also has to be defined carefully. A system may appear accurate statistically while still missing important dimensions of learning, such as creativity, reasoning quality, or disciplinary judgment. The strongest approach is to treat AI as one component in a broader assessment strategy, not as an unquestioned authority. Used responsibly, AI can contribute to fairer and more accurate assessment. Used carelessly, it can make old problems harder to detect and more difficult to correct.
3. How could AI change the way students are tested and receive feedback?
AI could shift assessment away from isolated, one-time testing events and toward more continuous, interactive, and feedback-rich models of evaluation. Instead of relying only on end-of-unit exams or standardized formats, educators may increasingly use AI-supported tools to collect evidence of learning throughout the instructional process. This could include adaptive quizzes, simulations, writing platforms, problem-solving environments, and digital tasks that adjust in real time to student responses. In theory, that allows assessment to become more closely connected to learning rather than functioning only as a retrospective judgment.
One of AI’s most promising contributions is faster and more tailored feedback. Students often learn best when they receive timely guidance they can immediately apply. AI systems can provide comments on writing structure, flag errors in mathematical reasoning, suggest next steps in a practice task, or identify areas where a student may need additional instruction. For teachers, this can reduce the time spent on routine feedback and create more space for deeper instructional support, conferencing, and intervention.
Still, the quality of feedback matters more than the speed of delivery. Feedback that is generic, inaccurate, overly corrective, or disconnected from learning goals can do little to improve performance. There is also a risk that students may become overly reliant on AI hints or revisions without fully developing their own judgment and independence. The most effective future model is likely one in which AI handles parts of the feedback workflow, while teachers remain central in interpreting performance, setting expectations, and helping students understand how to improve. In that sense, AI may change the mechanics of testing and feedback, but not the educational principle that assessment should support learning in meaningful and understandable ways.
4. Will AI eventually replace teachers or human evaluators in assessment?
It is unlikely that AI will fully replace teachers or human evaluators in any responsible vision of educational assessment. Assessment is not just a technical process of scoring responses; it is also a professional and ethical process of deciding what counts as evidence, what standards matter, and how results should be used. Those decisions involve context, judgment, and sensitivity to learners that automated systems do not truly possess. Even when AI can score efficiently or identify patterns at scale, it does not understand student growth, motivation, classroom dynamics, or local instructional goals in the way educators do.
There are many parts of the assessment process where human expertise remains essential. Teachers design tasks that reflect actual learning objectives, recognize when a student’s unusual response demonstrates insight rather than error, and interpret results in light of background knowledge about the learner. Human evaluators are also needed to monitor whether AI tools are functioning as intended, to review contested results, and to guard against inappropriate uses of automated outputs. In high-stakes settings especially, relying solely on AI would raise major concerns about accountability, transparency, and due process.
A more realistic future is one of partnership. AI may take over selected routine functions, such as first-pass scoring, item generation, data summarization, or preliminary feedback. Human professionals, meanwhile, will continue to provide oversight, make final judgments, and ensure that assessment remains educationally meaningful. In fact, as AI systems become more common, the need for human assessment literacy may increase rather than decrease. Educators will need to know not only how to assess learning, but also how to evaluate the tools doing part of that work.
5. What should schools and universities consider before adopting AI-based assessment tools?
Schools and universities should begin with purpose, not technology. Before adopting any AI-based assessment tool, institutions need to ask what problem they are trying to solve and whether AI is genuinely the best solution. Are they seeking faster feedback, more scalable scoring, broader evidence of student performance, better accessibility, or improved instructional decision-making? Clear goals matter because AI tools can look impressive while offering little real educational value. Adoption should be guided by learning outcomes and assessment quality, not by novelty or marketing claims.
Once the purpose is clear, institutions should evaluate the tool through several critical lenses: validity, fairness, reliability, transparency, privacy, accessibility, and usability. They need evidence that the tool measures the intended construct, performs consistently, and does not disadvantage particular student groups. They should understand how the system was trained, what data it uses, how outputs are generated, and what limitations are known. Student data protection is especially important, since assessment information can be sensitive and consequential. Institutions should also ensure that AI-supported assessments remain accessible to learners with diverse needs and do not create unnecessary technical barriers.
Implementation planning is just as important as product selection. Faculty and staff need training, governance structures need to be defined, and policies must clarify when human review is required. Institutions should also establish procedures for monitoring performance over time, auditing for bias, and revising or withdrawing tools that do not meet educational standards. Perhaps most importantly, schools and universities should remember the historical pattern seen across educational testing: tools change, but the core questions remain the same. Does this assessment produce meaningful evidence? Is it fair? Does it help learners and educators make better decisions? If the answer to those questions is yes, AI may be a valuable addition. If not, adopting it simply adds complexity without improving assessment.
