Psychometric Validity
Evidence
Rigorous psychometric evidence is the foundation of any trustworthy assessment platform. This page presents b4skills' ongoing concurrent validity research — how our adaptive CEFR scores correlate with gold-standard external tests.
Concurrent Validity
Pearson r and Spearman ρ between b4skills θ estimates and criterion scores. Industry benchmark for high-stakes certification: r ≥ 0.85 (IELTS/TOEFL).
Annual Item Bank Health Report
Aggregate item statistics for the 2026 assessment year. Only ACTIVE items (≥ 200 calibration responses, IQS ≥ 65) are included.
Psychometric Framework
Item Response Theory (3PL)
Every item is calibrated under the three-parameter logistic model. Discrimination (a), difficulty (b), and pseudo-guessing (c) parameters are estimated via marginal maximum likelihood using ≥ 200 responses per item.
CAT — EAP θ Estimation
The adaptive engine uses Expected A Posteriori (EAP) θ estimation with a normal N(0,1) prior. Item selection maximises Fisher information at the current θ estimate. Sessions terminate when SEM ≤ 0.30.
CEFR Alignment
Cut scores are anchored to the CEFR via a standard-setting study using the Bookmark method. Provisional thresholds are reviewed against concurrent validity evidence and updated when n ≥ 200 per level boundary.
DIF & Fairness Analysis
Differential Item Functioning (DIF) is monitored using the Mantel-Haenszel statistic across gender, age group, and L1 language. Items flagged for DIF are reviewed and retired if bias is confirmed.
AI Scoring Governance
Open-ended responses (Writing & Speaking) are scored by a multi-rater ensemble (Gemini, GPT-4, Claude). Low-confidence outputs are escalated to a human examiner queue. Inter-rater reliability (QWK) is monitored monthly.
Concurrent Validity Design
Concurrent validity pairs require the external test to be taken within ±30 days of the b4skills assessment. Both self-reported and verified (official document) data sources are tracked separately in all analyses.
Key References
- Baker, F. B., & Kim, S. H. (2004). Item Response Theory: Parameter Estimation Techniques (2nd ed.). Marcel Dekker.
- Bland, J. M., & Altman, D. G. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 327, 307–310.
- Council of Europe. (2001). Common European Framework of Reference for Languages: Learning, Teaching, Assessment. Cambridge University Press.
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.
- Lord, F. M. (1980). Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum.
- Reckase, M. D. (2009). Multidimensional Item Response Theory. Springer.
- Wainer, H. (Ed.). (2000). Computerized Adaptive Testing: A Primer (2nd ed.). Lawrence Erlbaum.
Collaborate with Us
We are actively seeking university partners for concurrent validity studies. If your institution can provide IELTS/TOEFL/Cambridge scores alongside b4skills assessments, we would like to hear from you.
research@b4skills.com