Skip to main content
Research Paper Master's 641 words

Test-retest and inter-rater reliability methods in psychological testing

~4 min read
✍️ How to write this paper — guide & tools
Essay 641 words

¶ … improve the reliability of any test item. What are the factors that you should keep in mind while formulating these measures? Examine the parameters against which you should gauge the extent of improvement in the test item.

One of the most effective methods to improve test reliability is simply performing the test a second or multiple times: if a test can be performed again with similar results, this supports claims that the test's results were due to the manipulation of the variables under scrutiny, not to random occurrences. "Test-retest reliability is a measure of reliability obtained by administering the same test twice over a period of time to a group of individuals. The scores from Time 1 and Time 2 can then be correlated in order to evaluate the test for stability over time" (Phelan & Wren 2006). The different test trials should occur under relatively similar conditions.

The great value of this method is its simplicity and also its relative cheapness. However, "this definition relies upon there being no confounding factor during the intervening time interval" (Shuttleworth 2009). For some types of tests, such as educational aptitude tests, this method is not suitable, because students will have gained in knowledge during the interim. In theory, however, IQ or aptitude tests should be relatively similar between different test administrations, since IQ is supposed to be a stable variable and less affected by the acquisition of pure knowledge. In the absence of other confounding factors, the relative coherence between results should be the method by which this standard of reliability is judged. While theoretically "a perfect correlation between the test and the retest" is ideal, "perfection is impossible and most researchers accept a lower level, either 0.7, 0.8 or 0.9" of correlation (Shuttleworth 2009).

Another method is that of "inter-rater reliability" which is designed to eliminate personal bias affecting the results by having two assessors grade the same data separately: "Inter-rater reliability is useful because human observers will not necessarily interpret answers the same way; raters may disagree as to how well certain responses or material demonstrate knowledge of the construct or skill being assessed" (Phelan & Wren 2006). Inter-rater reliability may draw upon rates not originally involved in the study. "You probably should establish inter-rater reliability outside of the context of the measurement in your study" (Trochim 2006).

Using inter-rater reliability demonstrates a commitment to unbiased research. This is a critical issue, given that establishing that a researcher pursued an ethical course of action will be an area of scrutiny when the findings are presented. There are two major ways to establish inter-rater reliability. The first is to calculate the percentage of agreement between raters. The second is to "calculate the correlation between the ratings of the two observers. If your measurement consists of categories -- the raters are checking off which category each observation falls in -- you can calculate the percent of agreement between the raters" (Trochim 2006). The main disadvantage of inter-rater reliability as an estimate is that it requires two sets of raters equally informed about the research being conducted. This can be expensive or unfeasible to obtain.

78 Words Hidden
A study's budget, design, and type will all determine the most appropriate methods for the researcher to select. For example, "the test-retest estimator is especially…
Cite This Paper
PaperDue. (2015). Test-retest and inter-rater reliability methods in psychological testing. PaperDue. https://www.paperdue.com/essay/improve-the-reliability-of-any-2151690

Always verify citation format against your institution’s current style guide requirements.