Skip to content

  • Home
  • Assessment Design & Development
    • Assessment Formats
    • Pilot Testing & Field Testing
    • Rubric Development
    • Pilot Testing & Field Testing
    • Test Construction Fundamentals
  • Assessment in Practice (K–12 & Higher Ed)
    • Assessment for Learning (AfL)
    • Classroom Assessment Strategies
    • Grading & Reporting Systems
    • Higher Education Assessment
  • Careers, Certifications & Professional Development
    • Academic Publishing & Peer Review
    • Careers in Educational Assessment
    • Continuing Education Resources
    • Degrees & Certifications
  • Data Analysis & Interpretation
    • Data Visualization
    • Descriptive Statistics
    • Inferential Statistics
    • Interpreting Assessment Results
  • Toggle search form

Criterion-Related Validity: Predictive vs. Concurrent

Posted on September 9, 2026 By

Criterion-related validity explains how well a test, scale, or assessment corresponds with an external outcome that matters in the real world. In psychometrics, it is one branch of validity evidence, and it becomes especially useful when researchers, employers, clinicians, and educators need to know whether a score predicts future performance or aligns with a trusted measure collected at the same time. The two classic forms are predictive validity and concurrent validity. Both ask whether test scores relate to a criterion, but they differ in timing, design, and practical purpose.

Within the broader domain of validity and reliability, criterion-related validity sits alongside content validity and construct validity, while reliability addresses score consistency across items, raters, or time. I have seen teams confuse these ideas repeatedly: a highly reliable test can still be invalid, and a test with respectable construct evidence can still fail to predict an outcome that stakeholders care about. That distinction matters because decisions about admission, hiring, diagnosis, promotion, and intervention often depend on whether a measure has demonstrated a meaningful link to real performance.

For a hub article in psychometrics and measurement theory, criterion-related validity is the practical bridge between abstract measurement and applied decision-making. It helps answer direct questions: Will this admissions test forecast first-year GPA? Does this depression screener match clinical diagnosis today? Should an aptitude battery replace a structured interview, or complement it? These are not academic details. They affect fairness, legal defensibility, resource allocation, and the quality of conclusions drawn from data.

Understanding predictive versus concurrent validity also clarifies how to evaluate evidence across the wider validity and reliability landscape. Researchers need appropriate criteria, sufficient sample sizes, sound reliability estimates, and transparent interpretation of correlation coefficients, regression weights, classification accuracy, and incremental validity. Without those elements, claims about “validated” instruments are weak. With them, criterion-related validity becomes one of the clearest ways to judge whether a measure earns its place in practice.

What criterion-related validity means in measurement practice

Criterion-related validity is the extent to which scores on a measure relate to a criterion variable, usually an important behavioral, clinical, educational, or organizational outcome. The criterion should be relevant, independently measured, and itself reasonably accurate. In employment testing, common criteria include supervisor ratings, training completion, sales numbers, absenteeism, and safety incidents. In education, criteria often include course grades, standardized test scores, retention, or licensure performance. In clinical settings, criteria may involve diagnosis, symptom severity ratings, treatment response, or functional impairment.

The core statistic is often a correlation coefficient, though modern studies may also report regression results, area under the ROC curve, sensitivity and specificity, odds ratios, or cross-validated prediction error. In plain terms, stronger criterion-related validity means the test score moves in a useful way with the outcome of interest. A cognitive ability test correlated at .35 with job performance has more practical value than one correlated at .08, assuming the study design is sound and range restriction is considered. The Standards for Educational and Psychological Testing treat relations with other variables as a major source of validity evidence, and criterion evidence is one of the most decision-relevant forms within that category.

Criterion-related validity is never a permanent property of a test in the abstract. It is evidence tied to a specific use, population, criterion, and context. A nursing entrance exam may predict licensing outcomes well in one region but less well in another if preparation pathways, language demands, or score distributions differ. That is why validation should be local whenever stakes are high.

Predictive validity: forecasting future outcomes

Predictive validity evaluates whether current test scores forecast a criterion measured later. The timing is the defining feature. You administer the predictor first, wait, then collect the outcome. This design is common when organizations want to know whether a pre-hire assessment predicts later job performance, whether SAT or ACT scores predict college achievement, or whether an early literacy screener predicts reading proficiency by year’s end.

In practice, predictive validity studies are harder to run than many people expect. You need a longitudinal design, participant follow-up, and clear rules about attrition and missing data. Suppose a company uses a customer service simulation for applicants and later compares applicant scores with six-month call quality ratings, average handle time, and customer satisfaction. That can produce valuable evidence, but only if the criterion data are standardized and the sample is not distorted by selecting only top scorers. Restriction of range is a serious issue here because once low scorers are screened out, observed validity coefficients often shrink. Psychometricians may apply corrections, but corrected values must be interpreted cautiously and reported transparently.

Predictive validity is especially persuasive for high-stakes selection because it mirrors the actual decision sequence. You test first, decide, then observe outcomes. It supports claims that a measure helps forecast future performance beyond guesswork. However, it does not prove causation. A predictor may correlate with success because it captures background preparation, opportunity, language proficiency, or other related factors. That is why subgroup analyses, adverse impact review, and incremental validity testing are essential parts of responsible validation.

Concurrent validity: comparing measures in the present

Concurrent validity examines whether test scores relate to a criterion measured at roughly the same time. Instead of forecasting the future, it asks whether a measure agrees with a trusted standard now. This approach is common when building shorter forms, screening tools, translations, or replacement instruments. For example, a new five-minute anxiety screener may be validated against an established clinician-administered scale collected during the same visit. A new typing test may be compared with current productivity records for existing staff.

Concurrent studies are often faster and less expensive than predictive designs because the criterion is immediately available. That makes them attractive during instrument development. In my own measurement work, concurrent designs were often the first checkpoint before investing in a long follow-up study. If a new scale cannot align with a credible present-day criterion, there is little reason to assume it will predict meaningful outcomes later.

Still, concurrent validity has limitations. When predictor and criterion are gathered together, shared method variance can inflate associations. If both rely on self-report, correlations may partly reflect response style rather than the underlying trait. In personnel settings, validating a test on current employees can also create contamination because incumbents have already adapted to the job, and poor performers may have left. That means concurrent evidence can be useful, but it is not automatically interchangeable with predictive evidence.

Predictive vs. concurrent validity: key differences

The simplest distinction is timing: predictive validity uses a future criterion, while concurrent validity uses a present criterion. The deeper difference is purpose. Predictive validity supports decisions about what will happen next. Concurrent validity supports decisions about what a score means right now relative to an accepted benchmark. Both are valuable, but they answer different operational questions.

Aspect Predictive validity Concurrent validity
Criterion timing Measured after the predictor Measured at the same time as the predictor
Main use Selection, admission, forecasting Screening, instrument comparison, short-form validation
Design demands Longitudinal follow-up, attrition management Cross-sectional or same-window data collection
Common risk Range restriction after selection Shared method variance and criterion contamination
Interpretive strength Best for future outcome claims Best for present agreement claims

A practical example makes the contrast clear. If a medical school wants to know whether an admissions interview predicts residency performance three years later, that is predictive validity. If the school wants to know whether a brief situational judgment test aligns with an established professionalism rating administered during the application cycle, that is concurrent validity. The statistical tools may overlap, but the inferential claim is different.

How criterion-related validity connects to validity and reliability

As a hub topic, validity and reliability must be considered together. Reliability concerns whether scores are consistent. Internal consistency, test-retest reliability, interrater reliability, and parallel-forms reliability all affect how much criterion validity you can observe. Classical test theory makes this point directly: unreliability attenuates correlations. If a predictor has alpha or omega around .55, its correlation with a criterion will be artificially limited even if the underlying construct matters. The same applies when the criterion itself is noisy, such as supervisor ratings with low interrater agreement.

Validity is broader than criterion evidence alone. Content validity asks whether items adequately sample the domain. Construct validity asks whether scores behave as theory predicts across convergent, discriminant, and structural evidence. Face validity concerns superficial credibility, which can matter for acceptance but is not scientific proof. Consequential considerations address the social impact of testing. Good validation integrates these strands rather than treating any single coefficient as decisive.

This is why a selection test can show respectable predictive validity and still require revision. If item content underrepresents the job, subgroup differences are unexplained, or score interpretation drifts across languages, the evidence base remains incomplete. Conversely, a scale can be internally consistent and theoretically elegant yet practically weak if it fails to relate to outcomes anyone needs it to predict.

How to evaluate evidence quality in criterion studies

Strong criterion-related validity studies start with a defensible criterion. The criterion should be relevant, uncontaminated, and as objective as feasible. In job analysis, that means aligning outcomes with essential duties. In clinical research, it means using diagnoses established through recognized procedures such as structured interviews rather than casual impressions. In education, it means checking whether grades are comparable across instructors and whether standardized outcomes have enough score variance.

Next, inspect reliability for both predictor and criterion. Then examine sample size, representativeness, restriction of range, missing data, and model choice. Correlations around .20 can matter in large-scale selection contexts, but interpretation depends on base rates, selection ratios, and alternative predictors. Schmidt and Hunter’s personnel selection research showed that modest validity coefficients can still create substantial economic gains when applied across many hires, especially when combined with structured interviews or work samples.

Cross-validation is another nonnegotiable step. A validity coefficient estimated on one sample often drops in a new sample because of capitalization on chance. Split-sample validation, k-fold cross-validation, or external replication helps protect against overly optimistic claims. For classification tools, calibration matters too. A risk score can show acceptable discrimination yet systematically overpredict outcomes for some groups.

Finally, look for incremental validity. Does the measure add information beyond existing tools? A new admissions essay score that correlates with GPA is not automatically useful if it contributes nothing beyond prior grades and standardized tests. The best evidence shows unique predictive value, not just redundancy.

Common mistakes and better decisions

The most common mistake is treating any significant correlation as proof that a test is valid everywhere. Statistical significance is easy to achieve in large samples, but practical utility may still be trivial. Another mistake is using convenience criteria because they are available rather than because they represent the construct of interest. I have also seen organizations validate against supervisor ratings without checking rater training, halo effects, or criterion contamination from knowledge of test scores.

Better practice begins with a clear use case. Define the decision, identify the criterion, estimate reliability, and choose predictive or concurrent evidence based on the claim you need to support. Use multiple sources when possible. For example, a nursing program might combine concurrent evidence against faculty clinical ratings with predictive evidence for board exam pass rates. That gives a fuller picture than either design alone.

For readers building a stronger foundation in psychometrics and measurement theory, criterion-related validity is the working tool that connects validity and reliability to actual outcomes. Predictive validity tells you whether scores forecast future performance. Concurrent validity tells you whether scores align with an accepted current benchmark. The best assessments often need both, supported by reliable measurement, credible criteria, cross-validation, and careful interpretation.

If you are reviewing a test, survey, screener, or selection process, start by asking one direct question: what criterion must this measure relate to for the decision to be justified? From there, evaluate the timing, design, reliability, and utility of the evidence. That disciplined approach leads to better instruments, better decisions, and more defensible measurement practice across education, work, and health.

Frequently Asked Questions

What is criterion-related validity, and why does it matter?

Criterion-related validity refers to how well a test, scale, or assessment relates to an external criterion that has practical importance in the real world. In simple terms, it asks whether a score actually connects to something meaningful outside the test itself, such as job performance, academic success, clinical diagnosis, customer satisfaction, or another trusted measure. This matters because a test is only useful if its results help researchers and decision-makers understand, predict, or classify outcomes that truly matter. In psychometrics, criterion-related validity is one important source of validity evidence because it moves beyond internal test features and focuses on whether the measure works in practice.

For example, an employer might want to know whether a pre-employment assessment predicts later sales performance. A school may want evidence that an entrance exam relates to first-year grades. A clinician may ask whether a screening tool aligns with a well-established diagnostic interview. In each case, criterion-related validity helps answer the practical question: “Do the scores tell us something accurate and useful about an outcome we care about?” Without that evidence, even a well-designed assessment may have limited value in real-world settings.

What is the difference between predictive validity and concurrent validity?

The main difference lies in timing. Predictive validity examines whether test scores collected now can forecast a future outcome. Concurrent validity examines whether test scores relate to an outcome or established measure collected at about the same time. Both are forms of criterion-related validity, and both rely on comparing test scores with an external criterion, but they serve different purposes.

Predictive validity is especially important when the goal is forecasting. For instance, SAT scores might be evaluated based on how well they predict college GPA one year later, or a hiring assessment might be judged by whether it predicts job performance after six months. The emphasis is on future performance, future behavior, or a later criterion. If the relationship is strong, the test is considered more useful for selection, admissions, or other forward-looking decisions.

Concurrent validity, by contrast, is used when the goal is to show that a new measure agrees with an accepted benchmark measured in the present. For example, a shorter depression screening questionnaire might be compared with a gold-standard clinical interview administered during the same week. If the two measures show a strong relationship, that provides evidence of concurrent validity. This is often useful when validating a new instrument, replacing a more expensive test, or checking whether a simpler measure can stand in for a trusted one.

How is criterion-related validity usually measured?

Criterion-related validity is typically evaluated by examining the statistical relationship between test scores and the criterion. In many studies, that relationship is summarized with a correlation coefficient, which shows how strongly the assessment is associated with the external outcome. A higher positive correlation generally indicates that higher test scores go along with better performance on the criterion, while a lower correlation suggests weaker practical usefulness. Researchers may also use regression analysis, classification accuracy, sensitivity and specificity, or other methods depending on the type of measure and the nature of the criterion.

The strength of the evidence depends on more than just a single number. Researchers also consider the quality of the criterion itself, sample size, range restriction, reliability of both measures, and whether the relationship is meaningful in the intended context. For example, a modest correlation may still be valuable in complex settings like employee selection, where many factors influence performance. On the other hand, a strong correlation with a weak or poorly defined criterion may not provide convincing evidence. Good validation work therefore requires both sound statistics and careful judgment about what outcome is being measured and why it matters.

When should someone use predictive validity instead of concurrent validity?

Predictive validity should be used when the assessment is intended to support decisions about future outcomes. If an organization wants to know whether a test can help identify applicants who will succeed later, or whether a student assessment can forecast later academic achievement, predictive validity is the right framework. It directly addresses whether current scores can serve as a useful basis for future-oriented choices such as admissions, hiring, placement, certification, or intervention planning. Because it requires waiting for the criterion to occur, predictive validation can take more time and resources, but it often provides the strongest evidence for decision-making purposes.

Concurrent validity is more appropriate when the need is immediate or when a future outcome is not necessary to answer the research question. It is often used when validating a new measure against an established one, especially if the established measure is expensive, lengthy, invasive, or difficult to administer. For example, a researcher might compare a new brief anxiety scale with a widely accepted clinical assessment collected at the same time. Concurrent validity is practical and efficient, but it does not prove that the test predicts future success or future behavior. That is why it is best matched to situations where present-time agreement is the main concern.

What are common mistakes or limitations when evaluating criterion-related validity?

One common mistake is assuming that any correlation automatically demonstrates strong validity. In reality, the interpretation depends on context, purpose, and the quality of the criterion. If the criterion is vague, biased, or unreliable, then the validity evidence will be weakened no matter how promising the numbers appear. Another frequent issue is range restriction, which happens when the sample is too narrow, such as studying only top-performing applicants after selection has already occurred. This can artificially reduce observed relationships and lead people to underestimate a test’s usefulness.

Another limitation is overgeneralizing results from one setting to another. A test that predicts success in one company, one school, or one clinical population may not perform the same way elsewhere. Criterion-related validity is not a permanent label attached to a test; it is evidence that depends on the population, the criterion, and the use case. Timing can also create problems. For predictive validity, too short or too long a delay between test and outcome may distort the relationship. For concurrent validity, collecting measures “at the same time” does not guarantee they capture exactly the same construct. Strong evaluation therefore requires clear criteria, appropriate samples, careful methodology, and realistic expectations about what validity evidence can and cannot show.

Psychometrics & Measurement Theory, Validity & Reliability

Post navigation

Previous Post: How to Establish Content Validity in Assessments
Next Post: Threats to Validity in Educational Research

Related Posts

What Is Classical Test Theory (CTT)? A Complete Guide Classical Test Theory (CTT)
Key Concepts of Classical Test Theory Explained Classical Test Theory (CTT)
Understanding True Score Theory in CTT Classical Test Theory (CTT)
Observed Score vs. True Score: What’s the Difference? Classical Test Theory (CTT)
What Is Item Difficulty in Classical Test Theory? Classical Test Theory (CTT)
How to Calculate Item Difficulty Step-by-Step Classical Test Theory (CTT)
  • Educational Assessment & Evaluation Resource Hub
  • Privacy Policy

Copyright © 2026 .

Powered by PressBook Grid Blogs theme