Every good research study begins with a deceptively simple question: how many participants do I actually need? Pick too few, and your findings may not represent the population. Pick too many, and you waste time, money, and the goodwill of respondents. Determining the right sample size is the bridge between a curious idea and a credible conclusion, and getting it right is one of the most important skills a researcher in population and family health can develop.

Table of Contents

Why sample size matters so much

Sample size is not just a number you pull out of a hat or copy from a previous study. It directly affects the statistical power of your research, the precision of your estimates, and the ethical soundness of your design. A study that is too small may fail to detect a real effect, while one that is too large unnecessarily exposes participants to interview burden, blood draws, or other procedures.

According to a practical guide published in the Journal of General and Family Medicine, the appropriate sample size should be determined based on the study’s statistical analysis plan, which in turn depends on the study design, the research question, and the primary outcome. In simpler terms, the math you use depends on what you are trying to find out.

Calculating sample size for large populations (above 10,000)

When your target population is very large-say, all adolescent girls in a state or all married women of reproductive age in a metropolitan area-the population size has very little effect on the calculation. In such cases, researchers commonly use a formula attributed to W.G. Cochran:

n = (Zยฒ ร— p ร— q) / dยฒ

Here is what each variable means:

n is the required sample size.
Z is the standard normal deviate corresponding to your chosen confidence level (1.96 for 95% confidence, 2.576 for 99% confidence).
p is the estimated proportion of the attribute present in the population (for example, the expected prevalence of anemia among pregnant women).
q is 1 โˆ’ p, the proportion of the population that does not have the attribute.
d is the absolute precision or margin of error you are willing to accept, expressed as a proportion (commonly 0.05 or 5%).

A worked example

Suppose you want to estimate the proportion of women using modern contraceptives in a city of more than 500,000 people. You set the confidence level at 95% (Z = 1.96), assume p = 0.5 (the most conservative estimate when the true proportion is unknown), and choose a 5% margin of error (d = 0.05).

Plugging the values in: n = (1.96)ยฒ ร— 0.5 ร— 0.5 รท (0.05)ยฒ = 384.16, which rounds up to 385. As confirmed by a survey methodology guide on Cochran’s formula, this is why the figure of 385 (or roughly 384) appears so often in sample-size discussions-it is the default minimum at 95% confidence with a 5% margin and unknown proportion.

If you instead wanted 99% confidence with the same margin and proportion, the sample size would jump to roughly 663, illustrating that greater certainty always costs more participants.

Calculating sample size for smaller populations (below 10,000)

When the target population is finite and relatively small-say, all ASHA workers in a district, or all patients registered at a particular primary health centre-the formula above tends to overestimate the required sample. Researchers then apply the finite population correction (FPC).

The two-step process works like this:

Step 1: Calculate nโ‚€ using the Cochran formula above.
Step 2: Adjust it using:

n = nโ‚€ / [1 + (nโ‚€ โˆ’ 1)/N]

Where N is the known population size and nโ‚€ is the sample size from step 1.

Why the correction matters

A methodological note on Cochran’s calculator explains that when the initial sample exceeds roughly 5% of the population, applying the FPC produces a meaningfully smaller and more realistic sample. For instance, with nโ‚€ = 385 and a population of 5,000, the adjusted sample drops to 358. For a population of 2,000 employees in a workplace health study, the corrected sample shrinks from 384 to about 322 respondents.

This adjustment is logical: sampling a large proportion of a small population provides more information per participant, so fewer observations are needed to achieve the same precision.

Using tables for sample size

Not every researcher wants to plug numbers into a calculator every time. For decades, published tables have allowed researchers to look up the recommended sample size based on population size and margin of error.

The Krejcie and Morgan table

One of the most widely cited references is the table developed by Krejcie and Morgan in 1970, published in Educational and Psychological Measurement. The table is built on the same underlying chi-square formula used in proportion-based sampling, with p = 0.5 (for maximum sample size) and a 5% margin of error at 95% confidence.

A few useful reference points from the table:

For a population of 100, you need a sample of about 80.
For 500 people, the recommended sample is around 217.
For 1,000 people, you need roughly 278.
For 2,800 people, the suggested sample is 338, as noted in a practical illustration of the table’s use.
For 10,000 people, the sample size is approximately 370.
For very large populations (100,000 or more), the figure stabilises around 384.

The pattern is clear: as the population grows, the required sample size grows too, but at a diminishing rate. Once the population crosses the 100,000 mark, the sample size barely changes. This is why national surveys like the WHO STEPS surveys typically aim for a few hundred to a few thousand participants per stratum, not millions.

Other widely used tables

The classic manual Sample Size Determination in Health Studies by Lwanga and Lemeshow, published by the World Health Organization, provides ready-to-use tables for different study designs-prevalence surveys, case-control studies, cohort studies, and randomised trials. Many Indian public health institutions and ICMR-funded studies continue to reference this manual when justifying their sample size in protocols.

Factors influencing sample size

Whether you use a formula or a table, three key factors drive your final number.

Confidence level

The confidence level reflects how sure you want to be that the true population value lies within your computed interval. The most common choice is 95%, which gives a Z-value of 1.96. Raising the confidence to 99% (Z = 2.576) substantially increases the required sample, because you are demanding more certainty.

For exploratory studies, 90% confidence may be acceptable. For clinical trials with serious health consequences, 99% is often used. The choice should reflect the cost of being wrong.

Population variance

The more diverse or heterogeneous your population, the larger your sample must be. In the proportion-based formula, variance is captured by p ร— q. The product is highest when p = 0.5, which is why this value is used as a default when the true proportion is unknown-it gives the maximum (most conservative) sample size.

If a pilot study or prior literature suggests, for example, that 20% of women in your district experience postpartum depression, you can use p = 0.2 and q = 0.8, which lowers the required sample. A step-by-step guide on sample size determination notes that conducting a small pilot study is often the best way to refine estimates of response rates, variances, and effect sizes before the main study.

Degree of accuracy (margin of error)

The margin of error, d, defines how close you want your sample estimate to be to the true population value. A 5% margin is standard in social and health research, but it is not a universal rule. For rare conditions like maternal mortality, a 2-3% margin may be required to detect meaningful patterns, which dramatically increases the sample size since d appears squared in the denominator.

Cutting the margin of error in half quadruples the required sample size-a useful rule of thumb when planning your budget.

Other practical considerations

Beyond the three core factors, real-world studies must account for:

Non-response and dropout: If you expect 20% of participants to refuse or drop out, inflate your sample size accordingly. A common practice is to divide your calculated n by (1 โˆ’ expected non-response rate).
Study design: Case-control studies, cohort studies, and randomised trials each have their own formulas. The choice of statistical test (chi-square, t-test, regression) also affects required sample size.
Design effect: Cluster sampling, used in surveys like the National Family Health Survey, requires multiplying the simple-random-sample size by a design effect (often between 1.5 and 2) to account for within-cluster similarity. The NFHS methodology demonstrates how complex multi-stage designs handle this in practice.

Bringing it all together

Determining the right sample size is part science, part judgement. The formulas give you a defensible starting point, but choosing the inputs-confidence level, expected proportion, acceptable margin, anticipated non-response-requires careful thinking about your specific research question and population. A well-justified sample size strengthens the credibility of your study and convinces ethics committees, peer reviewers, and funders that your work is both rigorous and respectful of the people who agree to participate.

When in doubt, start with a pilot study, consult published tables for sanity checks, and remember that no calculation can rescue a poorly designed study. The formula simply tells you how many people to ask; the rest depends on asking the right questions of the right people in the right way.

What do you think? If you were designing a study on adolescent nutrition in your home district, would you prioritise a tighter margin of error or a higher confidence level-and what trade-offs would you be willing to make to keep the study feasible?

How useful was this post?

Click on a star to rate it!

Average rating 5 / 5. Vote count: 1

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://pmc.ncbi.nlm.nih.gov/articles/PMC10000262/
  2. https://www.sopact.com/use-case/survey-sample-size-calculator
  3. https://calculator.academy/cochrans-sample-size-calculator/
  4. https://kenpro.org/sample-size-determination-using-krejcie-and-morgan-table/
  5. https://qhaireenizzati.wordpress.com/2017/10/05/sample-size-determination-using-krejcie-and-morgan-table/
  6. https://www.who.int/data/data-collection-tools/who-steps-noncommunicable-disease-risk-factor-surveillance/manuals
  7. https://pmc.ncbi.nlm.nih.gov/articles/PMC8075593/
  8. https://rchiips.org/nfhs/

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Research Methodology in Population and Family Health Studies

1 Social Science Research- An Overview

  1. The Meaning and Concept of Social Science Research
  2. The Differences between Natural and Social Science Research
  3. Approaches to Social Science Research
  4. Types of Social Science Research

2 Components of Social Science Research

  1. Concept
  2. Objectives
  3. Definition
  4. Hypothesis
  5. Variables

3 Research Designs

  1. Research Design – Meaning and Concept
  2. Functions of Research Design
  3. The Need for Research Design
  4. Features of Research Design
  5. Types of Research Design

4 Research Project Formulation

  1. Steps in the Formulation of a Research Project Proposal
  2. The Title of a Research Project
  3. Problem Statement
  4. Review of Literature
  5. Objectives of Research
  6. Methodology
  7. Work Schedule/Time Frame
  8. Budget
  9. Dissemination Strategy

5 Measurement

  1. Measurement โ€” Meaning and Concept
  2. Importance of Measurement
  3. Measurement Postulates
  4. Kinds of Measurement
  5. Admissible Statistical Tests for Measurement
  6. Criteria for Judging the Measuring Instruments
  7. Sources of Errors in Measurement

6 Scales and Tests

  1. Scales: Meaning and Techniques
  2. Types of Rating Scales
  3. Uses and Guidelines for Construction of Rating Scales
  4. Rating Errors
  5. Tests
  6. Types of Objective Test Questions
  7. Test Construction

7 Reliability and Validity

  1. Reliability
  2. Methods of Determining the Reliability
  3. Validity
  4. Types of Validity
  5. Reliability or Validity – Which is More Important?

8 Sampling

  1. Sampling: Meaning and Concept
  2. Types of Sampling
  3. Sample Design Process
  4. Errors in Sampling
  5. Determination of Sample Size

9 Quantitative Data Collection Methods and Devices

  1. Primary Data Collection: Meaning and Methods
  2. Questionnaire Method of Data Collection
  3. Interview Schedule
  4. Secondary Methods of Data Collection

10 Qualitative Data Collection Methods and Devices

  1. Qualitative Data – Meaning and Concept
  2. Methods and Techniques of Qualitative Data Collection
  3. Features of Qualitative and Quantitative Research

11 Data Sources- Primary and Secondary

  1. Sources of Data
  2. Process of Sourcing Data
  3. Qualities of Data Source
  4. Data Sources for Agriculture
  5. Data Sources for Infrastructure
  6. Data Sources for Service Sector
  7. Global Data Sources

12 Use of ICT in Data Collection and Processing

  1. ICT: Meaning and Attributes
  2. ICT and Development Interface
  3. ICT and Sectoral Development
  4. E-Development and its Strategies

13 Overview of Statistical Tools and Techniques

  1. The Data: Meaning and Types
  2. Frequency Distributions
  3. Measures of Central Tendency
  4. Measures of Dispersion
  5. Hypothesis Testing and Inferential Statistics
  6. Statistical Tests
  7. Correlation
  8. Regression

14 Data Processing and Analysis

  1. Data Measurement and Its Type
  2. Tabulation and Interpretation of Data
  3. Data Coding, Editing and Feeding
  4. Data Tabulation
  5. Graphical Presentation of Data

15 Report Writing

  1. Types of Report
  2. Writing the Research Report
  3. Preliminary Pages of Research Report
  4. Main Components or Chapterizing of Research Report
  5. Style and Layout of the Report

16 Dissemination of Findings

  1. Concept and Definition of Dissemination of Findings
  2. Importance of Dissemination
  3. Various Strategies of Dissemination of Findings
  4. Challenges in Dissemination of Findings
  5. Approaches for Dissemination

17 Project Cycle Management

  1. Projects: Meaning and Concept
  2. Difference between a Project and a Programme
  3. Project Preparation
  4. Project Cycle Management
  5. Project Appraisal Techniques

18 Monitoring

  1. Meaning and Scope of Monitoring
  2. Monitoring: What, Why, When and by Whom
  3. Basic Concepts and Elements in Monitoring
  4. Types of Monitoring
  5. The Techniques of Monitoring

19 Evaluation

  1. What is Evaluation?
  2. Appraisal vs. Monitoring vs. Evaluation vs. Impact Assessment
  3. Evaluation – Types and Designs
  4. Evaluation – Data Collection Methods
  5. Evaluation Approaches

20 Impact Assessment of Projects and Programmes

  1. Impact Assessment: Meaning and Importance
  2. Types of Impact Assessment
  3. Tools and Techniques used in Impact Assessment
  4. Steps in Implementing an Impact Assessment
  5. Associated Terms Related to Impact Assessment

21 Introduction to GIS and RS in Population Studies

  1. Basic Concepts of Geoinformatics
  2. Geospatial Data
  3. Overview of Applications of RS and GIS
  4. Application in Population Studies
  5. RS and GIS in Population Studies: Indian Examples