IQ Testing Limitations in Intellectual Disability Assessment
This paper critically evaluates the use of IQ scores as the primary tool for assessing intellectual disability (ID), formerly termed mental retardation. Drawing on the 2002 AAMR definitional revisions and a review of empirical studies, the paper highlights significant limitations of IQ-based assessment, including score instability over time, floor effects in standardized subtests, and substantial discrepancies between widely used instruments such as the WISC, WAIS, and Stanford-Binet. Evidence from multiple studies demonstrates that the same individual may receive markedly different IQ scores depending on which test is administered, with consequences for special education placement, resource allocation, and legal outcomes such as capital punishment eligibility. The paper concludes by calling for a standardized IQ measure applicable across all assessment contexts.
- Introduction: AAMR Revisions and State Guidelines: AAMR definitional changes and inconsistent state adoption
- IQ Score Instability and Floor Effects: IQ scores vary over time and overestimate ability
- Discrepancies Between the WISC and WAIS: Same-manufacturer tests yield significantly different scores
- WAIS Versus Stanford-Binet in Adults with Intellectual Disabilities: WAIS overestimates ability relative to Stanford-Binet
- Ruling Out Confounding Factors: Alternative explanations eliminated from findings
- Conclusion: The Need for a Standardized IQ Measure: Calls for a single standardized IQ assessment tool
✍️ How to write this paper — guide, tools & examples ▾
What makes this paper effective
- The paper builds a cumulative argument across four empirical studies, with each source adding a new dimension of measurement error — temporal instability, cross-instrument disparity within the same manufacturer, and cross-instrument disparity across different publishers.
- It anchors abstract psychometric concerns to concrete real-world stakes (special education placement, capital punishment eligibility), giving the technical discussion practical urgency.
- The author acknowledges and preemptively dismisses alternative explanations (practice effects, Flynn effect, differential minimum scores), strengthening the central claim.
Key academic technique demonstrated
The paper exemplifies evidence synthesis for a critical-evaluation argument. Rather than simply summarizing individual studies, the author sequences them to show an escalating problem: instability within one test, then disagreement between two tests from the same publisher, then disagreement between tests from different publishers. This layered structure transforms a collection of separate findings into a single, coherent indictment of IQ-based assessment.
Structure breakdown
The paper opens with policy context (AAMR revisions and Polloway et al.'s state-guidelines survey), identifies the gap those guidelines leave unexamined (IQ test reliability), and then moves through three empirical critiques in ascending severity. A brief methodological-controls paragraph preempts counterarguments before the conclusion issues a policy recommendation. The structure follows a classic problem-evidence-recommendation pattern well suited to applied assessment topics.
Introduction: AAMR Revisions and State Guidelines
In 2002, the American Association on Mental Retardation (AAMR) made changes to their manuals regarding the assessment of mental retardation (MR). The revisions were designed to affect changes in professional practice regarding the assessment of MR, public policy, and the scientific understanding of MR. Key among these changes was the attempted shift from the term "mental retardation" to the more current term "intellectual disability." Assessment was to consider both IQ scores and adaptive behavior (AB) — referred to as "adaptive skills" — as well as the individual's cultural background and associated strengths. Rather than following a deficit model of explanation, the goal was to follow a needs model. The definition of intellectual disability then includes three core criteria: significant impairment of intellectual functioning (defined by decreased IQ scores), significant impairment of adaptive and social functioning, and onset before adulthood.
Polloway et al. (2009) examined the impact of these changes on state guidelines for the assessment and treatment of MR. Interestingly, 27 states still formally used the MR term in their guidelines and definitions, with only four using the recommended term. Most states still adopted older formal definitions of MR. With respect to IQ scores — the primary focus of this paper — only 12 states reported that no specific IQ cutoff score was needed for diagnosis. States that did report a specific IQ score typically used a cutoff of 70 or 70–75, either by formal definition or by maintaining that a score of two standard deviations below the mean was the threshold. Forty-nine states required deficits in adaptive behavior. Age guidelines and a classification system (e.g., designating mild MR, moderate MR, and so on) were variably employed.
Polloway et al. (2009) largely leave the IQ assessment issue alone, suggesting that it is a cornerstone of the recognition and assessment of MR. However, IQ scores as a method of assessing MR have severe limitations and certainly warrant more attention regarding accurate assessment than whether a state terms the condition "mental retardation," "cognitive impairment," or "intellectual disability." This paper focuses primarily on Full Scale IQ scores and their equivalents.
IQ Score Instability and Floor Effects
One significant issue with IQ scores is their consistency over time, especially at the lower end of the IQ distribution. Changes in IQ scores across time are often attributed to random or systematic error, which is why it is often preferable to report confidence intervals rather than point estimates of IQ. However, Whitaker (2008) performed a meta-analysis of individuals who obtained low IQ scores (Full Scale IQ below 80) and were retested at a mean interval of 2.8 years. Despite most IQ test manuals reporting 95% confidence intervals of approximately five IQ points on either side of the obtained score, Whitaker found that 14% of the scores in the meta-analysis changed by 12.5 points or more.
Moreover, stability at the lower ends of the subtest score distribution is questionable, and floor effects exist. For instance, in the Wechsler IQ protocols, a raw score of zero on a subtest often still yields a scaled score of one or higher in older individuals administered the WAIS rather than the WISC. This means that scaled scores will frequently overestimate a person's ability in that particular domain. When these overestimates occur, there are serious consequences for the classification of individuals with intellectual disabilities who require special education. If IQ scores cannot demonstrate constancy over time, their utility as the primary assessment tool for classifying students with MR is questionable — and the stakes extend beyond education. In cases where the death penalty is involved, a cognitive assessment can be crucial to a person's life, as it is unlawful to execute an individual with intellectual disability who has been convicted of a crime.
Discrepancies Between the WISC and WAIS
There are no standardized or formalized recommendations specifying which IQ test provides the best estimate of intellectual disability, and different tests yield different results even when administered to the same person. This effect has been observed even among tests produced by the same manufacturer. Gordon, Duff, Davison, and Whitaker (2010) noted that the WISC and WAIS have historically provided disparate Full Scale IQ scores when administered to the same subjects. Because the WISC has a ceiling age of 16 and the WAIS can be administered to subjects as young as 16 years old, Gordon et al. administered both — the WISC-IV (UK) and the WAIS-III (UK) — to 17 students with a mean age of 16.2 years who had been identified as having intellectual disabilities. A counterbalanced repeated-measures design was used to avoid order effects.
Despite finding significant correlations between the IQ scores for the two tests, the mean WAIS-III Full Scale IQ was 64, whereas the mean WISC-IV score was 53 — a significant difference that would result in two different classifications of MR severity. All participants scored lower on the WISC-IV than on the WAIS-III, with the smallest difference between any participant's scores being five points. The Index Scores, with the exception of the Working Memory Index, were also significantly higher on the WAIS-III, with differences ranging between 9.50 and 12.58 points. The primary difference in Full Scale IQ scores was reflected in the verbal domain: the mean Verbal Comprehension Index score was 67.59 on the WAIS-III and 55.76 on the WISC-IV.
Conclusion: The Need for a Standardized IQ Measure
IQ tests have been used to measure low intellectual ability ever since Binet and Simon produced their original IQ test in 1905. IQ is viewed as a measurable construct, and the concept of intellectual disability has been viewed as a measurable entity since the development of these tests. State guidelines regarding the assessment of MR — in order to provide special services to clients who need them — typically specify IQ as a determining factor. Those states that do not publish formal IQ cutoff scores most often rely on their respective educational systems to determine the appropriate cutoff for assessing MR (Polloway et al., 2009).
However, as the evidence reviewed here demonstrates, the assessment of IQ is fraught with inaccuracies. These measurement errors affect who receives what resources, who is placed in the appropriate environment, and in some states, who lives and who dies. It is time for a standardized IQ measure to be adopted and used consistently across all situations in which intellectual disability must be assessed.
References
Gordon, S., Duff, S., Davison, T., & Whitaker, S. (2010). Comparison of the WAIS-III and WISC-IV in 16-year-old special education students. Journal of Applied Research and Intellectual Disabilities, 23, 197–200.
Polloway, E. A., Patton, J. R., Smith, J. D., Antoine, K., & Lubin, J. (2009). State guidelines for mental retardation and intellectual disabilities: A revisitation of previous analyses in light of changes in the field. Education and Training in Developmental Disabilities, 44, 14–24.
Silverman, W., Miezejeski, C., Ryan, R., Zigman, R., Krinsky-McHale, S., & Urv, T. (2010). Stanford-Binet & WAIS IQ differences and their implications for adults with intellectual disability. Intelligence, 38(2), 242–248.
Whitaker, S. (2008). The stability of IQ in people with low intellectual ability: An analysis of the literature. Intellectual and Developmental Disabilities, 46, 120–128.
Always verify citation format against your institution’s current style guide requirements.