Anchor items in test equating are the common questions used to connect scores from different test forms so results remain comparable across administrations. In psychometrics, an anchor item set serves as a statistical bridge: examinees may take different versions of a test, yet scaled scores can still be interpreted on the same reporting scale. That function makes anchor items central to scaling and equating, the branch of measurement theory concerned with maintaining fairness when assessments change over time. I have worked on operational testing programs where a form refresh looked simple on paper but would have broken trend lines without a carefully designed anchor block, reviewed under both classical and item response theory assumptions.
To understand why anchor items matter, it helps to define the surrounding terms. Scaling is the process of placing scores or item parameters onto a defined metric, such as a reporting scale from 200 to 800 or an ability scale centered at zero. Equating is the statistical adjustment that makes scores from alternate forms interchangeable, accounting for small differences in difficulty. Linking is a broader term for establishing relationships between scores, while calibration refers to estimating item characteristics such as difficulty, discrimination, and sometimes guessing. In many programs, common-item nonequivalent groups designs rely on anchor items because test takers in one administration are not randomly equivalent to those in another. The anchor provides the evidence needed to separate true form difficulty from shifts in examinee ability.
This topic matters because high-stakes decisions depend on comparable scores. College admissions tests, licensure exams, interim assessments, and statewide accountability systems all rotate forms to protect security and refresh content. Without equating, a harder spring form could unfairly suppress pass rates, while an easier fall form could inflate them. Anchor items reduce that risk by preserving the reporting meaning of a score. They also support trend reporting, standard setting maintenance, and vertical scaling across grade levels. As the hub for scaling and equating within psychometrics and measurement theory, this article explains how anchor items are selected, how they function in common equating designs, what can go wrong, and how testing organizations monitor quality in practice.
What anchor items are and how they support score comparability
An anchor item is administered in more than one test form and scored consistently so its statistical behavior can be compared across forms. The simplest case is two parallel forms sharing a subset of identical operational questions. If the common questions perform similarly, analysts can estimate the relative difficulty shift between the forms and adjust raw-to-scale conversions accordingly. In a classical framework, this may involve equating observed scores through linear or equipercentile methods. In an item response theory framework, anchor items help place item parameters and examinee abilities on a common scale using concurrent calibration or separate calibration with linking constants such as the Stocking-Lord or Haebara methods.
Anchor items are not random leftovers from previous tests. Good anchors represent the content blueprint, cognitive demand, statistical difficulty range, and item format mix of the full form. If a mathematics assessment contains algebra, geometry, and data analysis, the anchor set should reflect those domains rather than overrepresenting one strand. If the total test mixes easy, medium, and difficult items, anchors should cover that spread. In practice, I look for anchors that are stable, free of known exposure problems, and supported by solid point-biserial correlations, acceptable fit statistics, and no unresolved fairness concerns from differential item functioning reviews.
Score comparability depends on that representativeness. Suppose a reading test refreshes many inference items but the anchor set contains mostly literal comprehension questions. The common block may suggest one difficulty relationship, while the full form behaves differently because the anchor does not mirror the operational construct. The result is weak equating and noisy scale conversions. This is why technical manuals often specify anchor blueprints and minimum information targets. The common items must capture enough of the measurement signal to stabilize the equating relationship, not merely satisfy a numeric requirement.
Major equating designs in scaling and equating
Scaling and equating use several designs, but anchor items are especially important in common-item designs. In a random groups design, different forms are given to randomly equivalent groups, so group differences average out by design. In a single-group design, the same examinees take both forms, often with counterbalancing to control order effects. Those designs can produce strong equating evidence but are often impractical operationally. Large testing programs usually administer one form at a time to naturally occurring groups, which are rarely identical in ability. That is where the common-item nonequivalent groups design becomes the workhorse.
In the common-item nonequivalent groups design, Form X and Form Y are given to different populations, but both include an anchor set. Analysts use the common items to infer how much of the score difference comes from form difficulty versus population ability differences. This design is standard in statewide assessment, admissions testing, and credentialing. It is operationally feasible, supports security, and fits recurring administrations. Its quality, however, rests heavily on anchor integrity. If anchor items drift because of curriculum changes, coaching, compromised exposure, or wording differences in reused contexts, the design weakens quickly.
Vertical scaling adds another layer. When programs want scores across grades to support growth interpretations, anchor items or anchor testlets may connect adjacent grade forms. Because content legitimately changes by grade, vertical links must balance overlap with developmental appropriateness. In practice, this often means using items near the boundary of adjacent grade standards, then checking whether the resulting scale behaves sensibly across the score range. Horizontal equating across the same grade is typically simpler than vertical scaling because the construct is more constant.
Selecting strong anchor items
Strong anchor items meet content, statistical, and operational requirements simultaneously. Content specialists first verify alignment to standards and blueprint cells. Psychometricians then screen for stable p-values, strong discrimination, acceptable model fit, and no evidence of local dependence that would overweight one skill. Item writers and security teams review whether previous exposure has been too high. Fairness specialists examine differential item functioning across relevant subgroups. The best anchors are often straightforward operational items that measure the intended construct cleanly; flashy items with multimedia complexity may be valid for testing but less stable for equating.
A practical rule used in many programs is to allocate enough anchor material to capture 15 percent to 25 percent of the scored test, though the right proportion depends on test length, score precision targets, and model choice. Very short tests usually need a relatively stronger anchor because each item carries more information. Longer tests can sometimes support a smaller percentage if the anchor is highly representative and statistically informative. Item response theory also encourages attention to information functions. An anchor concentrated only around average ability may link the center well but leave the tails unstable, which matters for pass-fail decisions near cut scores or growth claims at the extremes.
| Criterion | Why it matters | Example of good practice |
|---|---|---|
| Content representation | Prevents the anchor from reflecting only one domain or process | Match anchor distribution to blueprint strands and cognitive levels |
| Difficulty spread | Supports linking across the score scale rather than only at the center | Include easy, medium, and hard items with known stable performance |
| Statistical quality | Reduces noise in equating constants and score conversions | Retain items with solid discrimination and acceptable fit indices |
| Security stability | Limits exposure effects that can make anchor items artificially easier | Rotate anchor pools and retire compromised items promptly |
| Fairness evidence | Protects comparability across demographic groups | Exclude items with unresolved differential item functioning concerns |
Context matters as well. If the anchor appears in a different section order, adjacent item context can change performance through fatigue, priming, or passage effects. Passage-based tests often anchor whole sets rather than isolated items for that reason, though testlets introduce local dependence concerns that must be modeled or managed. Equating is not just about reusing text; it is about preserving the measurement meaning of that text under comparable administration conditions.
How anchor items work in classical and item response theory methods
In classical test theory, anchor items may support linear, mean, sigma, or equipercentile equating, depending on the design and score scale. Analysts summarize how groups performed on the common items and use that information to adjust observed-score relationships between forms. These methods can work well, especially for large samples and stable forms, but they are tied more closely to the specific score distributions observed in each administration. They are also less flexible when forms differ in information across the ability range.
Item response theory uses anchor items to identify the transformation between item parameter estimates from separate calibrations or to stabilize concurrent calibration. For dichotomous items, common models include the one-parameter, two-parameter, and three-parameter logistic models; for polytomous items, the generalized partial credit model and graded response model are common. If Form X and Form Y are calibrated separately, the anchor items provide paired parameter estimates. Linking methods such as Stocking-Lord then find constants that minimize differences between test characteristic curves, placing both forms on one metric. Once items are aligned, raw scores can be converted to scaled scores through the common ability scale.
Operational programs often compare multiple equating methods rather than relying on one output blindly. If concurrent calibration, separate calibration with Stocking-Lord, and observed-score equipercentile equating all point to similar scale conversions, confidence rises. If methods diverge sharply, the anchor set, sample composition, model fit, or dimensionality assumptions deserve scrutiny. Good psychometric practice includes sensitivity analyses, standard errors of equating, and reviews of subgroup consistency, not merely a single reported conversion table.
Common threats to anchor validity
The biggest threat to anchor items is parameter drift, sometimes called item drift, where an item changes in difficulty or discrimination across administrations for reasons unrelated to form difficulty. Drift can follow curriculum emphasis changes, public sharing of items, altered testing mode, revised accessibility features, or subtle wording edits. A classic example is a science item that becomes easier after a widely adopted curriculum unit starts teaching the exact graph type used in the question. If that item remains in the anchor, it can bias the link and make the new form appear harder than it really is.
Mode effects are another common problem. When a program moves from paper to computer delivery, even identical items may not behave the same way. Reading passages on screen, drag-and-drop interactions, calculators embedded in software, or timing displays can all alter performance. In those transitions, programs often conduct bridge studies with dedicated anchors and conservative reporting rules before declaring full comparability. The same caution applies to translated forms, accommodated administrations, and changes in scoring rubrics for constructed-response tasks.
Security breaches can be especially damaging because anchors are reused by design. If tutoring networks obtain anchor content, those items may show inflated performance while the rest of the form remains unaffected. The equating then overcorrects. That is why secure pool management, exposure monitoring, forensic data review, and timely retirement policies are not side issues; they are part of equating validity. I have seen otherwise well-built links rejected because an anchor subset showed suspiciously low response times and abnormally high accuracy concentrated in a handful of sites.
Best practices for operational testing programs
Strong scaling and equating systems treat anchor items as governed assets, not convenient statistics. Start with an anchor plan in the test design phase: blueprint targets, content constraints, information targets, exposure limits, and retirement rules should all be specified before form assembly. After administration, evaluate anchor performance using preestablished criteria, including item fit, differential item functioning, residual analyses, and comparisons of equating outcomes with and without flagged items. Document every decision in a technical manual so users understand the basis for score comparability claims.
Programs should also maintain link chains carefully. If Form A links to B, and B links to C, error can accumulate over time. Periodic refresh studies, external audits, or designs that reconnect back to a stable reference form help control scale drift. For large programs, common-item pools are often replenished continuously so anchors remain representative without becoming overexposed. Small-volume credentialing exams may need different strategies, including pretesting pipelines, matrix sampling, or stronger use of IRT calibration banks when sample sizes are limited.
The practical takeaway is simple: anchor items make score comparability possible, but only when they are representative, secure, stable, and reviewed with discipline. They sit at the center of scaling and equating because they connect test forms, support reporting continuity, and protect fairness for examinees. If you are building or evaluating an assessment program, audit your anchor design, monitoring rules, and equating evidence before trusting score trends. Strong anchors do not guarantee perfect equating, but weak anchors almost always undermine it.
Frequently Asked Questions
What are anchor items in test equating, and why are they so important?
Anchor items are questions that appear in more than one version of a test so psychometricians can link those forms statistically. Their main purpose is to act as a stable reference point when different groups of examinees take different test forms across administrations. If every test form were scored in isolation, even small differences in difficulty could make score comparisons misleading. A student who took a slightly harder form might appear to perform worse than a student with the same underlying ability who took an easier form. Anchor items help correct for that problem.
In practice, the anchor set serves as a bridge between forms. Because the same items are administered again, analysts can examine how those common questions function across administrations and estimate the relative difficulty of each form. That information is then used in scaling and equating so reported scores stay comparable over time. This is what allows testing programs to say that a score earned on one administration has the same meaning as the same score earned on another administration, even when the exact questions are not identical.
Anchor items are especially important in large-scale educational and credentialing assessments, where fairness, consistency, and score interpretability are essential. Without a strong anchor design, score trends could reflect shifts in form difficulty rather than real changes in performance. For that reason, anchor items are not just a technical detail in psychometrics; they are one of the core tools used to support valid score comparisons.
How do anchor items help make scores from different test forms comparable?
Anchor items make comparability possible by providing direct evidence about how two or more test forms relate to one another in difficulty and scale location. Suppose Form A and Form B are both intended to measure the same construct, but they are assembled separately. Even when test developers work carefully to match specifications, one form may still turn out a bit easier or harder than the other. By including a set of common items on both forms, psychometricians gain a shared metric for evaluating those differences.
The logic is straightforward: if the anchor items are truly common and function similarly across groups, performance on those items can reveal whether one form was relatively more difficult. For example, if examinees perform as expected on the anchor set but overall scores differ across forms, part of that difference may reflect form difficulty rather than true ability differences. Statistical equating methods then use the anchor item data to adjust raw-to-scale score conversions so equivalent performance receives equivalent reported scores.
In item response theory, anchor items are often used to place item parameters and examinee ability estimates from separate calibrations onto a common scale. In classical equating frameworks, common-item designs support methods that estimate score relationships between forms. Although the exact procedures vary, the principle is the same: common items provide the evidence needed to align forms. When done well, this process protects the meaning of the score scale and helps ensure that decisions based on scores remain fair across administrations.
What makes a good anchor item set in psychometric practice?
A good anchor item set is representative, stable, secure, and statistically informative. Representativeness matters because the anchor should reflect the content and cognitive demands of the total test. If the anchor items cover only a narrow slice of the blueprint or emphasize one skill type too heavily, they may not provide an accurate basis for linking forms. Ideally, the anchor mirrors the broader assessment in content balance, item format, and difficulty distribution.
Stability is equally important. Anchor items should function consistently across administrations and groups, meaning they should measure the intended construct in the same way over time. Items that show unusual shifts in difficulty, wording sensitivity, or subgroup performance can weaken the equating. Psychometricians therefore review anchor items for differential item functioning, parameter drift, and other signs that the item may not be behaving as a reliable link.
Item quality also matters. Strong anchor items tend to have good discrimination, clear wording, and well-established performance histories. They should provide useful measurement information and contribute meaningfully to the statistical relationship between forms. At the same time, security is critical. Because anchor items may be reused, overexposure can threaten their usefulness. If content becomes widely known, item performance may change for reasons unrelated to ability, undermining the equating design.
Finally, the anchor set must be large enough to support accurate linking. There is no single ideal number in every testing program, but too few anchor items can produce unstable equating results, while a thoughtfully constructed set gives analysts a stronger basis for estimating form differences. In short, a good anchor set is not merely reused content; it is a carefully selected measurement tool designed to preserve the validity of score comparisons.
What can go wrong if anchor items are poorly chosen or do not perform well?
When anchor items are poorly chosen, the entire equating process can become less accurate. One major risk is that the anchor set may not represent the test as a whole. If the common items are much easier, harder, or narrower in content than the operational form, the resulting equating may misstate the relationship between forms. That can lead to score adjustments that are technically precise but substantively misleading.
Another problem is item drift, which occurs when an anchor item’s statistical characteristics change over time. Drift can happen for many reasons, including curriculum changes, increased familiarity with the content, shifts in instruction, item exposure, or subtle wording effects across populations. If an item becomes easier or harder for reasons unrelated to the construct being measured, it stops being a neutral bridge. In that case, the equating may absorb noise instead of capturing true form differences.
Differential item functioning is another concern. If anchor items behave differently for subgroups after controlling for ability, they may distort linking relationships and raise fairness issues. Security breaches are also serious. Once anchor items are overexposed, examinees may answer them correctly because they have seen them before, not because they possess the intended skill or knowledge. That weakens the validity of the anchor and, by extension, the comparability of reported scores.
For these reasons, testing programs routinely monitor anchor performance using statistical diagnostics and content review. Items that show instability, drift, or unusual subgroup patterns may be removed from the anchor set or replaced in future administrations. Strong equating depends not just on having common items, but on having common items that continue to function as trustworthy indicators of form relationships.
Are anchor items the same as pretest items or reused items on an exam?
Not exactly. Anchor items, pretest items, and reused items can overlap, but they are not identical concepts. Anchor items are included specifically to support linking and equating across forms. Their defining feature is their role in maintaining a common scale. Pretest items, by contrast, are typically included to collect data on new questions before they are used operationally. They help test developers evaluate item difficulty, discrimination, and fit, but they are not necessarily part of the equating design.
Reused items is the broadest category. Any item that appears again on a later test form is, in a general sense, reused. However, not every reused item is functioning as an anchor. To serve as an anchor, the item must be intentionally selected and incorporated into a formal equating plan. It must also be analyzed as part of the statistical linking process. Simply repeating a question does not automatically make it an anchor item.
In some testing programs, an item can serve multiple roles over time. For example, a question may first appear as a pretest item, later become operational, and then eventually be selected as part of an anchor set if it demonstrates strong psychometric performance. Even in those cases, the role of the item depends on how it is being used in the measurement system. Understanding that distinction is important because equating is a highly structured process, and anchor items are chosen not merely for convenience, but for their ability to support valid, fair, and stable score interpretations across different test forms.
