Reliability, validity, and standardization in psychological test development
✍️ How to write this paper — guide & tools ▾
WHAT MAKES A GOOD TEST 4 What Makes a Good Test Developers should ensure that the tests they develop are reliable and valid. Validity means that the test is able to measure what it is supposed to measure. Validity determines the usefulness of a test. The scores on the test must be related to some other behavior that is reflective of ability, personality, or interest. On the other hand, reliability refers to the test producing the same results when taken at different times by the same individuals. This means the test is consistent all the time (Noble & Smith, 2015). Developers have to ensure that their tests will produce similar results every time it is used to measure the same individual. Standardization is another way that can be used to make a good test. Standardization for psychology tests requires that a test is administered to a large sample derived from the intended population. After the test is administered computation for certain group statistics namely mean and standard deviation. Tying out the test will allow the developer to see the scores that they will typically obtain after conducting the test. This will allow the developer to make sense of their scores by comparing them to typical scores. In order to ensure validity and reliability, the developer should ensure that they have a large random sample of participants whenever possible or when conducting the study. To ensure validity the developer should also ensure that the results of the test agree with those of another test. In order for a test to be considered valid, it should give out the same values as those of an established test. This would guarantee the validity of the new test. Predictive validity should also be considered where predictions can be made that would agree with what one expects as a measure of characteristic. When the test is examined, it is vital that one sees that is measures what the developer claims is measured by the test. For example, if the test is for mathematical aptitude, the test should contain logical and mathematical problems to solve. Reliability is guaranteed by ensuring that the test would produce similar results when the individuals are tested repeatedly under the same conditions. Therefore, the developer should ensure that the test does not produce different results if an individual is tested again. Test-retest reliability should be ensured and the developer should aim at having the test-retest reliability correlation being 0.95 or better (Farrelly, 2013). The developer can also use the split-half reliability to ensure the reliability of their test. Any test that will be carried out on human subjects should ensure that their proper consent was given by the participant before they can be allowed to take the test. Also, the developer should ensure that the test is not biased in any way to result in the discrimination of certain individuals based on their culture or gender. The test should be aimed at only focusing on what the developer is interested in testing and should not be impacted upon if the participants are from different cultures. The nature of the test should provide the participants with ample assurance that their information will not be disclosed in any way apart from the results of the test. Personally identifiable information that is captured during the test should not be included in the final results or presentation of the findings. This would ensure that the developer is ethical and has not gone against the requirements of research. In order to create a norm for a test, the test developer will administer the test to a large group of subjects across different age groups. This is aimed at providing a scaled score that can be used to determine how other test takers performed as compared to the initial subjects. The main goal of norming a test is to get an estimate that can be applied to other subjects or test takers. The norming of tests is done to reflect how a test taker performed as compared to other subjects of the same age. Norming of tests is not done to establish the subject’s mastery of a particular subject or their cognitive abilities. Norming a test for a population requires that the developer collects data from thousands of subjects within the population in order to identify and establish norms for the population. A self-reported test is one where the respondent reads and selects an answer themselves without the interference of the researcher (Dyrstad, Hansen, Holme, & Anderssen, 2014). The researcher does not get involved and they solely rely on the answers provided by the respondent. There is no chance that the researcher can clarify the answers provided. Most of the self-reported tests are closed-ended. Administered tests, on the other hand, require the researcher to ask the questions and the respondent provides answers. The researcher then notes down or records what the respondent has said in regards to each question asked. With administered tests, the researcher can seek clarification from the respondent and allow them to expound or explain their response. In self-reported tests, the researcher can use the split testing to determine the reliability of the test. This way the same participant would be answering two halves of the test and if the results are identical, the test would be easily deemed reliable. This would counter the disadvantage of self-reports where participants are more likely to exaggerate or embarrassed to reveal certain details. Administered tests have the advantage of allowing the researcher to clarify the provided response and seek further information from the participant. One of the disadvantages of administered tests is that they are time-consuming because the researcher can only administer the test to one participant at a time.
Create your account
Always verify citation format against your institution’s current style guide requirements.