Skip to main content
Essay Undergraduate 1,743 words

Limitations of Norms in Psychological Testing

~9 min read 6 sections Psychology
Abstract

This paper examines the limitations of norm-referenced psychological tests, beginning with the foundational assumptions required for valid norm construction—including ordinality, transitivity, and population homogeneity. It provides an overview of how standardization samples and norm groups are developed and how raw scores are converted into percentiles for comparison. The paper then reviews real-world examples illustrating problems that arise when norms are inadequately constructed or applied, including sociodemographic bias, small normative samples, poor construct validity, and misaligned severity cutoff criteria. Case studies drawn from the Denver Developmental Screening Test, the Rey Auditory Verbal Learning Test, and assessments used in South African and U.S. educational contexts highlight why periodic revision and culturally appropriate norms are essential for accurate psychological interpretation.

Key Takeaways
  • Introduction: Benefits and Imperfections of Norm-Referenced Tests: Overview of norm-referenced test benefits and shortcomings
  • Assumptions Underlying Norm Construction: Key mathematical and conceptual assumptions for valid norms
  • Overview of Norms in Psychological Testing: How standardization samples and norm groups are developed
  • Limitations of Norms on Psychological Test Interpretation: Real-world failures in normative test revision and standards
  • Reconciling Limitations and Appropriate Use of Norms: Research examples of norm misuse and psychometric deficiencies
  • Conclusion: Score interpretation depends on appropriate norm group membership
✍️ How to write this paper — guide, tools & examples

What makes this paper effective

  • Grounds abstract concepts in concrete case studies—the Denver II revision, the Rey AVLT re-norming, and the South African cross-cultural study—making the argument about normative limitations tangible and evidence-based.
  • Organizes the discussion logically: from theoretical assumptions, to how norms are built, to documented real-world failures, giving the reader a clear conceptual journey.
  • Balances technical vocabulary (ordinality, transitivity, ANCOVAs, construct validity) with accessible explanations, appropriate for an undergraduate psychology audience.

Key academic technique demonstrated

The paper uses a literature synthesis approach, drawing on multiple primary and secondary sources to build a cumulative argument. Rather than relying on a single study, it triangulates evidence across developmental, neuropsychological, and educational contexts to demonstrate that norm limitations are a systemic, cross-domain problem—not an isolated issue.

Structure breakdown

The paper opens with a brief statement of the topic and its significance, then dedicates two sections to foundational concepts (assumptions and norm development). The analytical core follows in two sections that present and reconcile documented limitations through specific research examples. A short conclusion reinforces the central claim that valid score interpretation depends entirely on the appropriateness of the norm group. The structure moves clearly from theory to practice to implications.

Essay 1,743 words

Introduction: Benefits and Imperfections of Norm-Referenced Tests

Tests that are norm-referenced provide a number of benefits over non-norm-referenced tests. Psychological tests enable the gathering of valuable information about individual functioning across many different areas. Most norm-referenced tests are relatively quick to administer, such that a psychologist can obtain a sampling of behavior with a small investment of time and resources. A primary advantage of psychological testing is that rich and detailed information is revealed through the testing process that would otherwise be unavailable to the psychologist. However, norm-referenced tests are far from perfect, and the quality, reliability, and validity of norm-referenced tests varies substantially in some very important ways.

Assumptions Underlying Norm Construction

A number of assumptions are important to the construction of norms. The characteristic being measured must accommodate the ordering of individuals from low to high along an asymmetrical continuum that should at least be ordinal (Angoff, 1984). In addition, the relation of scores must be transitive. The mathematical definition of transitive states that if a condition "applies between two successive members of a sequence, it must also apply between any two members taken in order," such that, for example, "if A is larger than B, and B is larger than C, then A is larger than C" ("Transitive," n.d.) (Angoff, 1984).

The operational definition of the characteristic being measured must be reasonably clear and valid to a degree that it yields similar orderings of the characteristic among individuals (Angoff, 1984). The range of scores for a characteristic must all evaluate that same characteristic (Angoff, 1984). There must be a good match between the group(s), the target characteristics, and the test design and purpose (Angoff, 1984).

Norms are meaningful and useful only to the extent that they have been carefully defined. The norming population must be appropriate to the subject being tested and to the test itself; the challenge is to define the concept of appropriateness without conflating it with the concept of difficulty (Angoff, 1984). This means that a test or a subject can be difficult for many of the test takers, yet the test or the subject can still be considered appropriate for that population of test takers (Angoff, 1984).

Normative data should be developed for each distinct norming population for which it is meaningful to make comparisons with individuals or groups (Angoff, 1984). The test items themselves must be subject to pilot testing in which data about the items is drawn from samples of the population for which the test is being developed—that is, for the groups for which the norms will be provided (Angoff, 1984). Populations that serve as the basis for a set of norms should evidence homogeneity (Angoff, 1984). This means that all individuals are clearly members of the group and are logical and/or actual competitors in the same arena (Angoff, 1984).

Overview of Norms in Psychological Testing

A variety of norms exist, including the following: national norms, local norms, age and grade equivalents, item norms, school mean norms, user-selected norms, special study norms, and norms that yield direct meaning. This discussion centers on norms that are used for psychological tests, for which the following section provides an overview of how norms are developed.

Standardization samples are generated for psychological tests so that tests can be referenced to a normal distribution used to compare scores on specific future tests. Standardization relies on the creation of a large sample of test takers who are representative of the larger population for which the test is being developed. This standardization sample is referred to as the norm group or norming group. The raw scores of a sample group are converted into percentiles, which can be associated with a constructed normal distribution that will be used to rank the relative standing of individuals who take the test in the future.

Norms function as frames of reference for the interpretation of test scores, but they are not performance standards or clinical ideals. The size of norm groups varies widely, ranging from just a few hundred up to a hundred thousand people. As with other types of samples, the more individuals that are included in the norm group, the closer the sample approximates a normal population distribution. Moreover, normative data illustrates how the dimensions of major population subgroups differ and the extent to which test variables are associated with population classifications.

2 Sections Hidden · 540 words
Limitations of Norms on Psychological Test Interpretation230 words
Norms are developed according to the assumptions that underlie psychological testing and in accordance with criteria from a variety of sources. The involvement of several different authoritative sources in determining normative criteria…
Reconciling Limitations and Appropriate Use of Norms310 words
Several studies illustrate the problems encountered in the field when psychologists encounter lax standards for ensuring that the norm-referencing process is of consistently high quality. Sociodemographic factors can profoundly influence the accuracy of neuropsychological tests, as…

Conclusion

It is important to have information about the sample used to norm an instrument. An individual's score on a psychological test is only meaningful in the context of the standardization sample. Scores on psychological tests do not generally carry any concrete meaning on their own, which means that the interpretation of a score must always be relative to the scores received by other individuals on that same test. Moreover, it is essential to understand that the interpretation of an individual score is not possible if that person does not belong to the population that was used to norm the psychological test.

References

Angoff, W. H. (1984). Scales, norms, and equivalent scores. Princeton, NJ: Educational Testing Service. Retrieved from https://www.ets.org/Media/Research/pdf/Angoff.Scales.Norms.Equiv.Scores.pdf

Ferrett, H. L., Thomas, K. G., Tapert, S. F., Carey, P. D., Conradie, S., Cuzen, N. L., Stein, D. J., and Fein, G. (2014, June). The cross-cultural utility of foreign- and locally-derived normative data for three WHO-endorsed neuropsychological tests for South African adolescents. Metabolic Brain Disease, 29(2), 395–408. DOI: 10.1007/s11011.014.9495-6.

Frankenburg, W. K., Dodds, J. A., Shapiro, H., and Bresnick, B. (1992, January). The Denver II: A major revision and re-standardization of the Denver Developmental Screening Test. Pediatrics, 89(1), 91–97.

Kirk, C., and Vigeland, K. C. (2014, October). A psychometric review of norm-referenced tests used to assess phonological error patterns. Language, Speech, and Hearing Services in Schools, 45(4), 365–377. DOI: 10.1044/2014_LSHSS-13-0053.

Spaulding, T. J., Swartwout Szulga, M., and Figueroa, C. (2012, April). Using norm-referenced tests to determine the severity of language impairment between U.S. policy makers and test developers. Language, Speech, and Hearing Services in Schools, 43(2), 365–377. DOI: 10.1044/-1461(2011/10-0103).

Transitive. Google. Retrieved from https://www.google.com/webhp

Vakil, E., Greenstein, Y., and Blachstein, H. (2010). Normative data for composite scores for children and adults derived from the Rey Auditory Verbal Learning Test. Clinical Neuropsychologist, 24(4), 662–677.

Key Concepts in This Paper
Norm-Referenced Tests Standardization Sample Normative Data Construct Validity Transitivity Assumption Cultural Appropriateness Developmental Screening Verbal Memory Assessment Language Impairment Psychometric Properties
Cite This Paper
PaperDue. (2026). Limitations of Norms in Psychological Testing. PaperDue. https://www.paperdue.com/study-guide/limitations-norms-psychological-testing-2150113

Always verify citation format against your institution’s current style guide requirements.