When researchers in population and family health studies build a questionnaire to measure something like contraceptive knowledge, maternal health awareness, or attitudes toward immunisation, the first question they must answer is simple but unforgiving: can we trust this tool to give consistent results? A measurement instrument that yields different answers every time it’s used is worthless, no matter how clever its design. This is where reliability enters the picture, and three classical methods help us determine it: the test-retest method, the alternative form method, and the split-half method. Each has its own logic, its own strengths, and its own limitations.
Table of Contents
- What reliability means in research
- Test-retest method
- Advantages of the test-retest method
- Limitations of the test-retest method
- Alternative form method
- Advantages of the alternative form method
- Limitations of the alternative form method
- Split-half method
- Why the Spearman-Brown correction matters
- Advantages of the split-half method
- Limitations of the split-half method
- Choosing the right method in population and family health research
- Beyond the three classical methods
What reliability means in research
Reliability refers to the consistency, stability, and dependability of a measurement instrument. According to classical test theory, any observed score on a test is made up of two components: the respondent’s true score and measurement error. The smaller the error, the more reliable the instrument. A reliable scale measuring postpartum depression, for instance, should produce nearly identical scores if administered to the same woman twice within a short interval, assuming her condition hasn’t changed.
In population and family health research, where large-scale surveys like the National Family Health Survey (NFHS) shape national policy, reliability is non-negotiable. The numbers on contraceptive prevalence, child immunisation, and anaemia must be reproducible – otherwise the policies built on them stand on quicksand. Let’s now look at the three foundational methods researchers use to estimate reliability.
Test-retest method
The test-retest method is the most intuitive way to check reliability. The same test is administered to the same group of people on two separate occasions, and the scores from both administrations are correlated. The resulting correlation coefficient – usually Pearson’s product-moment coefficient – is the reliability estimate, often called the coefficient of stability.
If a researcher develops a 20-item questionnaire measuring adolescent knowledge of reproductive health and administers it to the same 100 students in March and again in April, a high correlation (say 0.85 or above) between the two sets of scores would suggest the instrument is stable over time.
Advantages of the test-retest method
The method is conceptually simple and easy to implement. It directly measures the stability of scores across time, which is particularly important for traits assumed to be relatively enduring – like personality factors, health beliefs, or cultural attitudes toward family planning. Only one version of the test is needed, which saves the researcher the work of developing parallel instruments.
Limitations of the test-retest method
The biggest problem is the memory effect. Respondents often remember their earlier answers, especially on cognitive or knowledge-based tests, and reproduce them on the second attempt. This inflates the correlation artificially and gives a misleadingly rosy picture of reliability.
The time gap is equally tricky. If the interval is too short, memory bias dominates. If it’s too long, genuine changes in the respondent – new learning, changes in opinion, life events – reduce the correlation, even though the instrument itself may be perfectly consistent. Practical events like fatigue, illness, or changed circumstances on the second testing day can also depress the score. Beyond this, asking the same people to take the same test twice is demanding on participants and resources, which often leads to attrition in field-based health research.
Alternative form method
To sidestep the memory problem, researchers can use the alternative form method, also called the parallel form or equivalent form method. Here, two different but equivalent versions of the test are constructed. Both versions measure the same underlying construct, contain the same number of items at comparable difficulty levels, and have similar formats – but the specific questions differ. Both forms are administered to the same group, usually with a short gap between them, and the two sets of scores are correlated. The resulting reliability coefficient is called the coefficient of equivalence.
For example, a research team studying nutritional knowledge among pregnant women might construct Form A with 30 items and Form B with 30 different but psychometrically matched items. A woman who scores 22 on Form A should score close to 22 on Form B if the instruments are truly equivalent.
Advantages of the alternative form method
This method effectively eliminates the memory bias that plagues test-retest, because respondents cannot simply recall and reproduce their earlier answers – the questions are different. For this reason it is often considered the most satisfactory method for well-standardised tests. The use of two parallel forms also broadens the sample of items covering the construct, which strengthens content validity at the same time.
Limitations of the alternative form method
The catch is severe: constructing two genuinely equivalent forms is extraordinarily difficult. The two forms must match in content coverage, item difficulty, item discrimination, and even item format. In specialised projective tests like the Rorschach, building a parallel form is virtually impossible. If the two forms differ even slightly in difficulty, the resulting correlation can be misleading.
The method also doubles the workload – for the researcher (who must develop and validate two versions) and for the respondents (who must sit through two tests). In field-based population health surveys conducted across rural and urban settings, this kind of dual administration is often impractical. The core challenge is generating enough items reflecting the same construct while ensuring the two sets are statistically interchangeable.
Split-half method
The split-half method offers an elegant workaround. Instead of administering the test twice or building two parallel forms, the researcher administers the test only once, then splits the items into two halves and correlates the scores on the two halves. This makes it a method of internal consistency rather than stability – it tells you whether the items within the test hang together as measures of the same construct.
The most common way to split is the odd-even split: all odd-numbered items form one half, all even-numbered items the other. Other approaches include splitting the first half versus the second half, or random splitting. The scores on the two halves are correlated, and the resulting coefficient is then corrected using the Spearman-Brown prophecy formula.
Why the Spearman-Brown correction matters
When you split a 40-item test into two halves of 20 items each, you’re effectively measuring the reliability of a shorter test. Since reliability generally increases with test length, the raw split-half correlation underestimates the reliability of the full instrument. The Spearman-Brown formula corrects for this by projecting what the reliability would be if the test were at its full length. In its simplest form for halving: rfull = 2rhalf / (1 + rhalf), where rhalf is the correlation between the two halves.
Advantages of the split-half method
The single biggest advantage is efficiency. Only one test administration is required, which eliminates memory effects, removes the burden of constructing parallel forms, and conserves time and resources – invaluable in low-budget public health studies. It also provides insight into the internal coherence of the instrument: are all the items pulling in the same direction, or are some measuring something quite different?
Limitations of the split-half method
The reliability estimate depends heavily on how the test is split. Different splits – odd-even, first-half versus second-half, or random – can produce different reliability coefficients on the same data. The method also assumes the two halves are equivalent, which isn’t always true, especially when items vary in difficulty or content.
The method is unsuitable for tests where item order matters, such as progressive difficulty tests or speed tests, because splitting them produces halves that aren’t comparable. It also doesn’t capture stability over time – a scale could have high internal consistency yet still produce different scores on different days. For this reason, modern researchers often supplement split-half estimates with Cronbach’s alpha, which is essentially the average of all possible split-half correlations.
Choosing the right method in population and family health research
No single method is universally best. The choice depends on what kind of error you most want to control, how much time and resources you have, and the nature of the construct being measured.
If you’re studying a stable trait like health-seeking behaviour or attitudes toward immunisation, test-retest is informative. If you’re worried about memory bias and have the resources to develop parallel forms – say, when assessing knowledge before and after a health-education intervention – the alternative form method is stronger. If you’re developing a new scale and need a quick, internal estimate, split-half (or its extension, Cronbach’s alpha) is the practical choice.
Large-scale Indian studies often combine methods. The Indian version of the HLS-EU-Q16 health literacy questionnaire, for instance, used Pearson’s correlation for test-retest reliability and Cronbach’s alpha for internal consistency, reporting alpha values of 0.97 and 0.98 for the Kannada and Hindi versions – excellent reliability by any standard. Similarly, the oral health KAB questionnaire developed for Indian adults tested internal consistency across knowledge, attitude, and behaviour domains using Cronbach’s alpha, demonstrating how researchers blend classical and modern reliability checks.
Beyond the three classical methods
While these three methods form the bedrock of reliability assessment, contemporary research increasingly relies on more sophisticated approaches like Cronbach’s alpha, intraclass correlation coefficients (ICC), and item response theory. These extensions address some of the limitations of the classical trio – for instance, Cronbach’s alpha generalises the split-half logic across all possible splits, giving a more stable estimate of internal consistency. Still, understanding the three classical methods remains essential, because they reveal the basic conceptual building blocks: stability, equivalence, and internal consistency.
A well-designed family health study often reports more than one type of reliability – for instance, both test-retest and Cronbach’s alpha – to give a fuller picture of how dependable the instrument is.
What do you think? If you were designing a questionnaire to measure rural women’s awareness of postnatal care services, which reliability method would you choose first, and why? And how would you handle the trade-off between memory bias in the test-retest method and the practical difficulty of building parallel forms?
References
- https://www.scribbr.com/methodology/types-of-reliability/
- https://www.dataforindia.com/nfhs-explainer/
- https://science.jrank.org/pages/5566/Psychometry-Reliability.html
- https://navclasses.in/reliability-and-its-types-test-retest-method-split-half-method-equivalent-form-parallel-form/
- https://www.yourarticlelibrary.com/statistics-2/determining-reliability-of-a-test-4-methods/92574
- https://conjointly.com/kb/types-of-reliability/
- https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7827499/
- https://assess.com/spearman-brown-prediction-formula/
- https://www.sciencedirect.com/science/article/abs/pii/S0895435617302494
- https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8780344/

Leave a Reply