When researchers in population and family health studies build a questionnaire to measure something like contraceptive knowledge, maternal health awareness, or attitudes toward immunisation, the first question they must answer is simple but unforgiving: can we trust this tool to give consistent results? A measurement instrument that yields different answers every time it’s used is worthless, no matter how clever its design. This is where reliability enters the picture, and three classical methods help us determine it: the test-retest method, the alternative form method, and the split-half method. Each has its own logic, its own strengths, and its own limitations.

Table of Contents

What reliability means in research

Reliability refers to the consistency, stability, and dependability of a measurement instrument. According to classical test theory, any observed score on a test is made up of two components: the respondent’s true score and measurement error. The smaller the error, the more reliable the instrument. A reliable scale measuring postpartum depression, for instance, should produce nearly identical scores if administered to the same woman twice within a short interval, assuming her condition hasn’t changed.

In population and family health research, where large-scale surveys like the National Family Health Survey (NFHS) shape national policy, reliability is non-negotiable. The numbers on contraceptive prevalence, child immunisation, and anaemia must be reproducible – otherwise the policies built on them stand on quicksand. Let’s now look at the three foundational methods researchers use to estimate reliability.

Test-retest method

The test-retest method is the most intuitive way to check reliability. The same test is administered to the same group of people on two separate occasions, and the scores from both administrations are correlated. The resulting correlation coefficient – usually Pearson’s product-moment coefficient – is the reliability estimate, often called the coefficient of stability.

If a researcher develops a 20-item questionnaire measuring adolescent knowledge of reproductive health and administers it to the same 100 students in March and again in April, a high correlation (say 0.85 or above) between the two sets of scores would suggest the instrument is stable over time.

Advantages of the test-retest method

The method is conceptually simple and easy to implement. It directly measures the stability of scores across time, which is particularly important for traits assumed to be relatively enduring – like personality factors, health beliefs, or cultural attitudes toward family planning. Only one version of the test is needed, which saves the researcher the work of developing parallel instruments.

Limitations of the test-retest method

The biggest problem is the memory effect. Respondents often remember their earlier answers, especially on cognitive or knowledge-based tests, and reproduce them on the second attempt. This inflates the correlation artificially and gives a misleadingly rosy picture of reliability.

The time gap is equally tricky. If the interval is too short, memory bias dominates. If it’s too long, genuine changes in the respondent – new learning, changes in opinion, life events – reduce the correlation, even though the instrument itself may be perfectly consistent. Practical events like fatigue, illness, or changed circumstances on the second testing day can also depress the score. Beyond this, asking the same people to take the same test twice is demanding on participants and resources, which often leads to attrition in field-based health research.

Alternative form method

To sidestep the memory problem, researchers can use the alternative form method, also called the parallel form or equivalent form method. Here, two different but equivalent versions of the test are constructed. Both versions measure the same underlying construct, contain the same number of items at comparable difficulty levels, and have similar formats – but the specific questions differ. Both forms are administered to the same group, usually with a short gap between them, and the two sets of scores are correlated. The resulting reliability coefficient is called the coefficient of equivalence.

For example, a research team studying nutritional knowledge among pregnant women might construct Form A with 30 items and Form B with 30 different but psychometrically matched items. A woman who scores 22 on Form A should score close to 22 on Form B if the instruments are truly equivalent.

Advantages of the alternative form method

This method effectively eliminates the memory bias that plagues test-retest, because respondents cannot simply recall and reproduce their earlier answers – the questions are different. For this reason it is often considered the most satisfactory method for well-standardised tests. The use of two parallel forms also broadens the sample of items covering the construct, which strengthens content validity at the same time.

Limitations of the alternative form method

The catch is severe: constructing two genuinely equivalent forms is extraordinarily difficult. The two forms must match in content coverage, item difficulty, item discrimination, and even item format. In specialised projective tests like the Rorschach, building a parallel form is virtually impossible. If the two forms differ even slightly in difficulty, the resulting correlation can be misleading.

The method also doubles the workload – for the researcher (who must develop and validate two versions) and for the respondents (who must sit through two tests). In field-based population health surveys conducted across rural and urban settings, this kind of dual administration is often impractical. The core challenge is generating enough items reflecting the same construct while ensuring the two sets are statistically interchangeable.

Split-half method

The split-half method offers an elegant workaround. Instead of administering the test twice or building two parallel forms, the researcher administers the test only once, then splits the items into two halves and correlates the scores on the two halves. This makes it a method of internal consistency rather than stability – it tells you whether the items within the test hang together as measures of the same construct.

The most common way to split is the odd-even split: all odd-numbered items form one half, all even-numbered items the other. Other approaches include splitting the first half versus the second half, or random splitting. The scores on the two halves are correlated, and the resulting coefficient is then corrected using the Spearman-Brown prophecy formula.

Why the Spearman-Brown correction matters

When you split a 40-item test into two halves of 20 items each, you’re effectively measuring the reliability of a shorter test. Since reliability generally increases with test length, the raw split-half correlation underestimates the reliability of the full instrument. The Spearman-Brown formula corrects for this by projecting what the reliability would be if the test were at its full length. In its simplest form for halving: rfull = 2rhalf / (1 + rhalf), where rhalf is the correlation between the two halves.

Advantages of the split-half method

The single biggest advantage is efficiency. Only one test administration is required, which eliminates memory effects, removes the burden of constructing parallel forms, and conserves time and resources – invaluable in low-budget public health studies. It also provides insight into the internal coherence of the instrument: are all the items pulling in the same direction, or are some measuring something quite different?

Limitations of the split-half method

The reliability estimate depends heavily on how the test is split. Different splits – odd-even, first-half versus second-half, or random – can produce different reliability coefficients on the same data. The method also assumes the two halves are equivalent, which isn’t always true, especially when items vary in difficulty or content.

The method is unsuitable for tests where item order matters, such as progressive difficulty tests or speed tests, because splitting them produces halves that aren’t comparable. It also doesn’t capture stability over time – a scale could have high internal consistency yet still produce different scores on different days. For this reason, modern researchers often supplement split-half estimates with Cronbach’s alpha, which is essentially the average of all possible split-half correlations.

Choosing the right method in population and family health research

No single method is universally best. The choice depends on what kind of error you most want to control, how much time and resources you have, and the nature of the construct being measured.

If you’re studying a stable trait like health-seeking behaviour or attitudes toward immunisation, test-retest is informative. If you’re worried about memory bias and have the resources to develop parallel forms – say, when assessing knowledge before and after a health-education intervention – the alternative form method is stronger. If you’re developing a new scale and need a quick, internal estimate, split-half (or its extension, Cronbach’s alpha) is the practical choice.

Large-scale Indian studies often combine methods. The Indian version of the HLS-EU-Q16 health literacy questionnaire, for instance, used Pearson’s correlation for test-retest reliability and Cronbach’s alpha for internal consistency, reporting alpha values of 0.97 and 0.98 for the Kannada and Hindi versions – excellent reliability by any standard. Similarly, the oral health KAB questionnaire developed for Indian adults tested internal consistency across knowledge, attitude, and behaviour domains using Cronbach’s alpha, demonstrating how researchers blend classical and modern reliability checks.

Beyond the three classical methods

While these three methods form the bedrock of reliability assessment, contemporary research increasingly relies on more sophisticated approaches like Cronbach’s alpha, intraclass correlation coefficients (ICC), and item response theory. These extensions address some of the limitations of the classical trio – for instance, Cronbach’s alpha generalises the split-half logic across all possible splits, giving a more stable estimate of internal consistency. Still, understanding the three classical methods remains essential, because they reveal the basic conceptual building blocks: stability, equivalence, and internal consistency.

A well-designed family health study often reports more than one type of reliability – for instance, both test-retest and Cronbach’s alpha – to give a fuller picture of how dependable the instrument is.

What do you think? If you were designing a questionnaire to measure rural women’s awareness of postnatal care services, which reliability method would you choose first, and why? And how would you handle the trade-off between memory bias in the test-retest method and the practical difficulty of building parallel forms?

How useful was this post?

Click on a star to rate it!

Average rating 4 / 5. Vote count: 1

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://www.scribbr.com/methodology/types-of-reliability/
  2. https://www.dataforindia.com/nfhs-explainer/
  3. https://science.jrank.org/pages/5566/Psychometry-Reliability.html
  4. https://navclasses.in/reliability-and-its-types-test-retest-method-split-half-method-equivalent-form-parallel-form/
  5. https://www.yourarticlelibrary.com/statistics-2/determining-reliability-of-a-test-4-methods/92574
  6. https://conjointly.com/kb/types-of-reliability/
  7. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7827499/
  8. https://assess.com/spearman-brown-prediction-formula/
  9. https://www.sciencedirect.com/science/article/abs/pii/S0895435617302494
  10. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8780344/

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Research Methodology in Population and Family Health Studies

1 Social Science Research- An Overview

  1. The Meaning and Concept of Social Science Research
  2. The Differences between Natural and Social Science Research
  3. Approaches to Social Science Research
  4. Types of Social Science Research

2 Components of Social Science Research

  1. Concept
  2. Objectives
  3. Definition
  4. Hypothesis
  5. Variables

3 Research Designs

  1. Research Design – Meaning and Concept
  2. Functions of Research Design
  3. The Need for Research Design
  4. Features of Research Design
  5. Types of Research Design

4 Research Project Formulation

  1. Steps in the Formulation of a Research Project Proposal
  2. The Title of a Research Project
  3. Problem Statement
  4. Review of Literature
  5. Objectives of Research
  6. Methodology
  7. Work Schedule/Time Frame
  8. Budget
  9. Dissemination Strategy

5 Measurement

  1. Measurement โ€” Meaning and Concept
  2. Importance of Measurement
  3. Measurement Postulates
  4. Kinds of Measurement
  5. Admissible Statistical Tests for Measurement
  6. Criteria for Judging the Measuring Instruments
  7. Sources of Errors in Measurement

6 Scales and Tests

  1. Scales: Meaning and Techniques
  2. Types of Rating Scales
  3. Uses and Guidelines for Construction of Rating Scales
  4. Rating Errors
  5. Tests
  6. Types of Objective Test Questions
  7. Test Construction

7 Reliability and Validity

  1. Reliability
  2. Methods of Determining the Reliability
  3. Validity
  4. Types of Validity
  5. Reliability or Validity – Which is More Important?

8 Sampling

  1. Sampling: Meaning and Concept
  2. Types of Sampling
  3. Sample Design Process
  4. Errors in Sampling
  5. Determination of Sample Size

9 Quantitative Data Collection Methods and Devices

  1. Primary Data Collection: Meaning and Methods
  2. Questionnaire Method of Data Collection
  3. Interview Schedule
  4. Secondary Methods of Data Collection

10 Qualitative Data Collection Methods and Devices

  1. Qualitative Data – Meaning and Concept
  2. Methods and Techniques of Qualitative Data Collection
  3. Features of Qualitative and Quantitative Research

11 Data Sources- Primary and Secondary

  1. Sources of Data
  2. Process of Sourcing Data
  3. Qualities of Data Source
  4. Data Sources for Agriculture
  5. Data Sources for Infrastructure
  6. Data Sources for Service Sector
  7. Global Data Sources

12 Use of ICT in Data Collection and Processing

  1. ICT: Meaning and Attributes
  2. ICT and Development Interface
  3. ICT and Sectoral Development
  4. E-Development and its Strategies

13 Overview of Statistical Tools and Techniques

  1. The Data: Meaning and Types
  2. Frequency Distributions
  3. Measures of Central Tendency
  4. Measures of Dispersion
  5. Hypothesis Testing and Inferential Statistics
  6. Statistical Tests
  7. Correlation
  8. Regression

14 Data Processing and Analysis

  1. Data Measurement and Its Type
  2. Tabulation and Interpretation of Data
  3. Data Coding, Editing and Feeding
  4. Data Tabulation
  5. Graphical Presentation of Data

15 Report Writing

  1. Types of Report
  2. Writing the Research Report
  3. Preliminary Pages of Research Report
  4. Main Components or Chapterizing of Research Report
  5. Style and Layout of the Report

16 Dissemination of Findings

  1. Concept and Definition of Dissemination of Findings
  2. Importance of Dissemination
  3. Various Strategies of Dissemination of Findings
  4. Challenges in Dissemination of Findings
  5. Approaches for Dissemination

17 Project Cycle Management

  1. Projects: Meaning and Concept
  2. Difference between a Project and a Programme
  3. Project Preparation
  4. Project Cycle Management
  5. Project Appraisal Techniques

18 Monitoring

  1. Meaning and Scope of Monitoring
  2. Monitoring: What, Why, When and by Whom
  3. Basic Concepts and Elements in Monitoring
  4. Types of Monitoring
  5. The Techniques of Monitoring

19 Evaluation

  1. What is Evaluation?
  2. Appraisal vs. Monitoring vs. Evaluation vs. Impact Assessment
  3. Evaluation – Types and Designs
  4. Evaluation – Data Collection Methods
  5. Evaluation Approaches

20 Impact Assessment of Projects and Programmes

  1. Impact Assessment: Meaning and Importance
  2. Types of Impact Assessment
  3. Tools and Techniques used in Impact Assessment
  4. Steps in Implementing an Impact Assessment
  5. Associated Terms Related to Impact Assessment

21 Introduction to GIS and RS in Population Studies

  1. Basic Concepts of Geoinformatics
  2. Geospatial Data
  3. Overview of Applications of RS and GIS
  4. Application in Population Studies
  5. RS and GIS in Population Studies: Indian Examples