Validity & Research

Psychometric Validity Evidence

Rigorous psychometric evidence is the foundation of any trustworthy assessment platform. This page presents b4skills' ongoing concurrent validity research — how our adaptive CEFR scores correlate with gold-standard external tests.

Concurrent Validity

Pearson r and Spearman ρ between b4skills θ estimates and criterion scores. Industry benchmark for high-stakes certification: r ≥ 0.85 (IELTS/TOEFL).

Loading validity data…

Annual Item Bank Health Report

Aggregate item statistics for the 2026 assessment year. Only ACTIVE items (≥ 200 calibration responses, IQS ≥ 65) are included.

Loading item bank statistics…

Psychometric Framework

Item Response Theory (3PL)

Every item is calibrated under the three-parameter logistic model. Discrimination (a), difficulty (b), and pseudo-guessing (c) parameters are estimated via marginal maximum likelihood using ≥ 200 responses per item.

CAT — EAP θ Estimation

The adaptive engine uses Expected A Posteriori (EAP) θ estimation with a normal N(0,1) prior. Item selection maximises Fisher information at the current θ estimate. Sessions terminate when SEM ≤ 0.30.

CEFR Alignment

Cut scores are anchored to the CEFR via a standard-setting study using the Bookmark method. Provisional thresholds are reviewed against concurrent validity evidence and updated when n ≥ 200 per level boundary.

DIF & Fairness Analysis

Differential Item Functioning (DIF) is monitored using the Mantel-Haenszel statistic across gender, age group, and L1 language. Items flagged for DIF are reviewed and retired if bias is confirmed.

AI Scoring Governance

Open-ended responses (Writing & Speaking) are scored by a multi-rater ensemble (Gemini, GPT-4, Claude). Low-confidence outputs are escalated to a human examiner queue. Inter-rater reliability (QWK) is monitored monthly.

Concurrent Validity Design

Concurrent validity pairs require the external test to be taken within ±30 days of the b4skills assessment. Both self-reported and verified (official document) data sources are tracked separately in all analyses.

Key References

  • Baker, F. B., & Kim, S. H. (2004). Item Response Theory: Parameter Estimation Techniques (2nd ed.). Marcel Dekker.
  • Bland, J. M., & Altman, D. G. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 327, 307–310.
  • Council of Europe. (2001). Common European Framework of Reference for Languages: Learning, Teaching, Assessment. Cambridge University Press.
  • Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.
  • Lord, F. M. (1980). Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum.
  • Reckase, M. D. (2009). Multidimensional Item Response Theory. Springer.
  • Wainer, H. (Ed.). (2000). Computerized Adaptive Testing: A Primer (2nd ed.). Lawrence Erlbaum.

Collaborate with Us

We are actively seeking university partners for concurrent validity studies. If your institution can provide IELTS/TOEFL/Cambridge scores alongside b4skills assessments, we would like to hear from you.

research@b4skills.com