When researchers design a questionnaire to study maternal stress, fertility preferences, or adolescent health knowledge, one question matters more than any other: is the tool actually measuring what it claims to measure? This is the heart of validity. A reliable instrument can produce consistent results every time, but if those results do not reflect the real concept under study, the entire research falls apart. Validity is not a single property either; it has several types, each addressing a different way a measurement can go wrong. Understanding these types is essential for anyone designing surveys, scales, or assessments in population and family health studies.

Table of Contents

What validity really means in research

Validity refers to the accuracy of a measurement, that is, whether a tool truly captures the concept it was built to assess. A scale that claims to measure depression but only asks about sleep patterns would have weak validity, since depression includes much more than disturbed sleep. According to a paper in the Indian Journal of Psychiatry, validity refers to whether the tool measures “what it purports to measure,” and is typically broken down into content validity, criterion validity, and construct validity.

These three forms are not competing alternatives. They are complementary lenses. The most rigorous research instruments demonstrate evidence across all three. Let us examine each in detail.

Content validity: Does the test cover the full topic?

Content validity asks a deceptively simple question: does the measurement adequately represent the entire domain of the concept being studied? If a researcher wants to measure mathematical aptitude, the test must cover algebra, geometry, arithmetic, and other key branches, not just one of them. A test that focuses only on arithmetic would have poor content validity because it leaves out major dimensions of mathematics.

The same logic applies to language assessments, which is one of the most cited examples in textbooks. A vocabulary test has strong content validity only if it samples words across different frequency bands rather than just common words, and a grammar test is valid only when it proportionally covers all the grammar points in the curriculum rather than oversampling rules that are easy to test. A reading test must use text types that represent real reading contexts, not just literary passages.

How researchers establish content validity

Unlike other forms of validity that depend on statistical analysis, content validity is usually established through expert judgement. A panel of subject matter experts reviews every item on the instrument and decides whether it is essential, useful, or unnecessary for measuring the construct. The higher the agreement among experts that each item is essential, the stronger the content validity.

Researchers also use structured tools called test blueprints or tables of specifications, which map test items to the dimensions of the construct and ensure proportional coverage. As Statistics By Jim explains, this process compares the test against its goals and the theoretical properties of the construct, ensuring that no important aspect is overlooked.

Why content validity matters in population studies

In family health research, content validity becomes critical when measuring concepts like contraceptive knowledge or prenatal care practices. If a questionnaire on contraceptive awareness only asks about oral pills and condoms, it ignores intrauterine devices, sterilisation, and emergency contraception, all of which are part of the actual domain. Such a tool would underestimate awareness levels and mislead policy decisions.

Criterion validity: Does the measure match a known standard?

Criterion validity examines how well a measurement tool correlates with another established measure or predicts a specific outcome. The “criterion” is essentially a gold-standard benchmark, an existing measure that is widely accepted as the best representation of the construct. Criterion validity has two important subtypes, and the difference between them lies in timing.

Predictive validity: Looking into the future

Predictive validity assesses how well a tool forecasts future outcomes. A classic example is the use of entrance examinations to predict student achievement. If scores on a college entrance test correlate strongly with grades two or three years later, the test has high predictive validity. This is exactly why standardised tests like the SAT, GRE, or India’s NEET are evaluated for their ability to predict later academic performance.

In public health, predictive validity helps evaluate screening tools. For instance, a domestic violence screening questionnaire used by community health workers has predictive validity if women identified as at-risk by the screening tool actually report or experience violence in follow-up studies.

Concurrent validity: Comparing in the present

Concurrent validity, on the other hand, looks at correlations happening at the same point in time. As the LEADERS Project at Columbia University explains, concurrent validity is derived from one test’s results agreeing with another test’s results that measure the same ability or quality. When researchers develop a new language test, they often compare its results with an already-published version of the same test administered at the same time.

The same principle is used when a new, shorter scale for measuring depression is validated against a longer, clinically established scale administered to the same people on the same day. Strong agreement signals high concurrent validity.

Statistical approach to criterion validity

Criterion validity is typically expressed through correlation coefficients. Pearson’s correlation coefficient is used for continuous variables, while the phi coefficient is used for dichotomous (yes/no) variables. The higher the correlation between the new measure and the gold standard, the stronger the criterion validity.

Construct validity: Measuring the unmeasurable

Construct validity is perhaps the most complex but also the most important type of validity. It addresses whether a measurement tool truly represents the theoretical concept it claims to measure. This becomes essential when researchers study abstract ideas that cannot be directly observed, such as family cohesion, maternal stress, intimate partner violence acceptance, or gender role attitudes.

Construct validity is especially crucial when there is no clear content universe to sample from and no obvious external criterion to compare against. You cannot point to a “true” measurement of empathy lying somewhere in the world. You have to build the concept theoretically and then test whether your tool behaves the way the theory predicts.

The two faces of construct validity

Construct validity is typically broken down into convergent validity and discriminant validity, a framework first introduced by Campbell and Fiske in 1959 through their Multitrait-Multimethod Matrix.

Convergent validity is the extent to which a measure correlates with other measures of the same or similar constructs. If a new scale for measuring self-esteem correlates strongly with an established self-esteem scale, the new tool shows good convergent validity. Discriminant validity is the opposite. It checks that the measure does not correlate strongly with measures of unrelated concepts. A self-esteem scale should not correlate strongly with, say, a measure of mathematical ability. To claim construct validity, researchers must demonstrate both.

Construct validity in family health research

Consider a study trying to measure “family cohesion” in joint households. A researcher might check whether the cohesion scale correlates positively with measures of family satisfaction (convergent validity) and negatively with measures of family conflict (discriminant validity). If both patterns appear, the scale is likely capturing the actual construct of family cohesion rather than a vague mix of attitudes.

Construct validity often involves additional statistical tools like factor analysis, which examines whether the items in a scale group together in ways that match the theoretical structure of the construct.

Choosing the right type of validity

No single type of validity is enough on its own. The choice depends on what the research is trying to accomplish. Content validity is the priority when developing comprehensive assessments of well-defined domains, such as knowledge of immunisation schedules or maternal nutrition practices. Criterion validity matters most when creating screening tools or predictive instruments that will guide interventions, such as identifying women at risk of postpartum depression. Construct validity is essential when studying abstract concepts at the heart of theories about behaviour change, family dynamics, or health disparities.

Cultural sensitivity and validity

A measure developed in one cultural setting may lose its validity when used in another. Questions about family decision-making that work well in individualistic societies may miss important dynamics in collectivist cultures where extended family members shape choices. A scale measuring autonomy in reproductive decisions may need substantial reworking when moving from a Western context to a setting where joint family structures influence decisions. Researchers must revalidate instruments through translation, expert review, and pilot testing whenever they cross cultural boundaries.

Common threats to validity

Several issues can undermine validity even in carefully designed studies. Construct underrepresentation occurs when a measure fails to capture important dimensions of the concept, like a poverty scale that ignores non-monetary deprivation. Construct-irrelevant variance happens when a measure inadvertently captures things outside the construct, such as a literacy test that ends up measuring test-taking speed rather than reading ability. Social desirability bias is another common threat, especially in family health research where respondents may give answers they think are socially acceptable rather than truthful.

The strongest research instruments do not chase one type of validity in isolation. They build evidence across content, criterion, and construct validity in an integrated way. As the foundational work in measurement theory has long argued, this combined approach gives the most comprehensive picture of whether a tool can truly be trusted.

What do you think? If you were designing a survey to measure women’s empowerment in rural India, which type of validity would you focus on first, and why might construct validity be particularly tricky in such a study? And how would you adapt a Western-developed scale on family cohesion to fit Indian joint-family contexts without losing its original meaning?

How useful was this post?

Click on a star to rate it!

Average rating 0 / 5. Vote count: 0

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://journals.lww.com/indianjpsychiatry/fulltext/2025/09000/criterion_validity,_construct_validity,_and_factor.13.aspx
  2. https://mikeydoes.com/glossary/content-validity/
  3. https://statisticsbyjim.com/basics/content-validity/
  4. https://www.leadersproject.org/2013/03/01/types-of-validity-in-testing/
  5. https://en.wikipedia.org/wiki/Convergent_validity

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Research Methodology in Population and Family Health Studies

1 Social Science Research- An Overview

  1. The Meaning and Concept of Social Science Research
  2. The Differences between Natural and Social Science Research
  3. Approaches to Social Science Research
  4. Types of Social Science Research

2 Components of Social Science Research

  1. Concept
  2. Objectives
  3. Definition
  4. Hypothesis
  5. Variables

3 Research Designs

  1. Research Design – Meaning and Concept
  2. Functions of Research Design
  3. The Need for Research Design
  4. Features of Research Design
  5. Types of Research Design

4 Research Project Formulation

  1. Steps in the Formulation of a Research Project Proposal
  2. The Title of a Research Project
  3. Problem Statement
  4. Review of Literature
  5. Objectives of Research
  6. Methodology
  7. Work Schedule/Time Frame
  8. Budget
  9. Dissemination Strategy

5 Measurement

  1. Measurement โ€” Meaning and Concept
  2. Importance of Measurement
  3. Measurement Postulates
  4. Kinds of Measurement
  5. Admissible Statistical Tests for Measurement
  6. Criteria for Judging the Measuring Instruments
  7. Sources of Errors in Measurement

6 Scales and Tests

  1. Scales: Meaning and Techniques
  2. Types of Rating Scales
  3. Uses and Guidelines for Construction of Rating Scales
  4. Rating Errors
  5. Tests
  6. Types of Objective Test Questions
  7. Test Construction

7 Reliability and Validity

  1. Reliability
  2. Methods of Determining the Reliability
  3. Validity
  4. Types of Validity
  5. Reliability or Validity – Which is More Important?

8 Sampling

  1. Sampling: Meaning and Concept
  2. Types of Sampling
  3. Sample Design Process
  4. Errors in Sampling
  5. Determination of Sample Size

9 Quantitative Data Collection Methods and Devices

  1. Primary Data Collection: Meaning and Methods
  2. Questionnaire Method of Data Collection
  3. Interview Schedule
  4. Secondary Methods of Data Collection

10 Qualitative Data Collection Methods and Devices

  1. Qualitative Data – Meaning and Concept
  2. Methods and Techniques of Qualitative Data Collection
  3. Features of Qualitative and Quantitative Research

11 Data Sources- Primary and Secondary

  1. Sources of Data
  2. Process of Sourcing Data
  3. Qualities of Data Source
  4. Data Sources for Agriculture
  5. Data Sources for Infrastructure
  6. Data Sources for Service Sector
  7. Global Data Sources

12 Use of ICT in Data Collection and Processing

  1. ICT: Meaning and Attributes
  2. ICT and Development Interface
  3. ICT and Sectoral Development
  4. E-Development and its Strategies

13 Overview of Statistical Tools and Techniques

  1. The Data: Meaning and Types
  2. Frequency Distributions
  3. Measures of Central Tendency
  4. Measures of Dispersion
  5. Hypothesis Testing and Inferential Statistics
  6. Statistical Tests
  7. Correlation
  8. Regression

14 Data Processing and Analysis

  1. Data Measurement and Its Type
  2. Tabulation and Interpretation of Data
  3. Data Coding, Editing and Feeding
  4. Data Tabulation
  5. Graphical Presentation of Data

15 Report Writing

  1. Types of Report
  2. Writing the Research Report
  3. Preliminary Pages of Research Report
  4. Main Components or Chapterizing of Research Report
  5. Style and Layout of the Report

16 Dissemination of Findings

  1. Concept and Definition of Dissemination of Findings
  2. Importance of Dissemination
  3. Various Strategies of Dissemination of Findings
  4. Challenges in Dissemination of Findings
  5. Approaches for Dissemination

17 Project Cycle Management

  1. Projects: Meaning and Concept
  2. Difference between a Project and a Programme
  3. Project Preparation
  4. Project Cycle Management
  5. Project Appraisal Techniques

18 Monitoring

  1. Meaning and Scope of Monitoring
  2. Monitoring: What, Why, When and by Whom
  3. Basic Concepts and Elements in Monitoring
  4. Types of Monitoring
  5. The Techniques of Monitoring

19 Evaluation

  1. What is Evaluation?
  2. Appraisal vs. Monitoring vs. Evaluation vs. Impact Assessment
  3. Evaluation – Types and Designs
  4. Evaluation – Data Collection Methods
  5. Evaluation Approaches

20 Impact Assessment of Projects and Programmes

  1. Impact Assessment: Meaning and Importance
  2. Types of Impact Assessment
  3. Tools and Techniques used in Impact Assessment
  4. Steps in Implementing an Impact Assessment
  5. Associated Terms Related to Impact Assessment

21 Introduction to GIS and RS in Population Studies

  1. Basic Concepts of Geoinformatics
  2. Geospatial Data
  3. Overview of Applications of RS and GIS
  4. Application in Population Studies
  5. RS and GIS in Population Studies: Indian Examples