Every research project, whether it is the National Family Health Survey or a small classroom study on dietary habits, faces the same practical question: do we study everyone, or do we study a few and draw conclusions about the many? This is where the concept of sampling enters. It is the quiet engine behind almost every statistic you read about population, fertility, maternal health, or nutrition. Understanding what sampling means, and the vocabulary that surrounds it, is the first step toward reading and producing credible research in population and family health studies.
Table of Contents
- What is sampling?
- Why sampling matters in population and family health studies
- Key sampling concepts
- Population
- Sampling unit
- Sampling frame
- Sample size
- Sampling fraction
- Population parameter and sample statistic
- Sampling error and sampling bias
- Advantages of sampling
- Cost-effectiveness
- Time-saving and quick data collection
- Fewer non-sampling errors
- Greater scope and detailed study
- Feasibility in destructive or sensitive studies
- When sampling is not appropriate
What is sampling?
Sampling is the process of selecting a smaller group, called a sample, from a larger group, called the population, in order to study the characteristics of that larger group. Instead of contacting every household in a state to learn about contraceptive use, a researcher selects a manageable subset that represents the whole. The findings from this subset are then used to make estimates about the entire population.
Several well-known methodologists have defined the term with slightly different emphases. According to Mildred Parten, sampling is the process by which a relatively small number of individuals, objects, or events is selected and analyzed to find out something about the entire population from which it is drawn. Richard Levin and David Rubin, in their widely used textbook on statistics for management, describe a sample simply as a collection of some, but not all, of the elements of the population under study, used to describe that population. Boyce adds a practical dimension by defining sampling as a procedure by which some members of a given population are selected as representatives of the entire population.
The common thread in these definitions is the idea of representativeness. A good sample is not just any small group; it is a group whose characteristics mirror those of the larger population on the variables that matter for the study. A sample is a subset of individuals from a larger population, and sampling means selecting the group that you will actually collect data from in your research.
Why sampling matters in population and family health studies
Population studies often deal with millions of people. Examining every individual is rarely feasible. Large-scale exercises like the Census of India or the National Family Health Survey rely heavily on carefully designed samples to estimate indicators such as infant mortality, total fertility rate, anaemia prevalence, and institutional delivery rates. Without sampling, generating timely health statistics for policy decisions would be almost impossible.
Key sampling concepts
Sampling has its own vocabulary, and using these terms precisely is what separates a casual survey from a credible piece of research. The following concepts form the backbone of any sampling discussion.
Population
The population, sometimes called the universe, is the entire group of individuals, objects, or events that the researcher wants to study. In a maternal health study, the population might be all pregnant women in a particular district during a specific year. Populations can be finite, like all registered ASHA workers in Bihar, or infinite, like all possible outcomes of a coin toss. A useful distinction is also made between the target population, which is the ideal group the researcher wants to learn about, and the accessible population, which is the portion that can realistically be reached.
Sampling unit
The sampling unit is the basic element selected for inclusion in the sample. It can be an individual, a household, a village, a school, or even a hospital ward, depending on the study. The sampling unit is the unit that is actually selected during the sampling process, and the sampling frame is the list of these units from which the sample is drawn. In a study on adolescent nutrition, the sampling unit may be an individual girl aged 10 to 19. In a study on water and sanitation, it could be an entire household.
Sampling frame
The sampling frame is the actual list or device from which the sample is drawn. If the population is all registered voters in a city, the electoral roll becomes the sampling frame. A sampling frame is a researcher’s list or device to specify the population of interest, and in a simple random sample every unit in this frame has an equal probability of being drawn. An incomplete or outdated frame is one of the most common sources of bias in surveys. For instance, using only landline telephone directories in a survey on mobile health apps would exclude a huge section of users and distort results.
Sample size
The sample size, denoted by n, is the number of units actually studied. The size of the population is usually denoted by N. Choosing an appropriate sample size is a balance between precision and resources. Too small a sample produces unreliable estimates, while an unnecessarily large sample wastes time and money. Sample size depends on the variability of the characteristic being studied, the desired level of confidence, the acceptable margin of error, and the design of the study.
Sampling fraction
The sampling fraction is the ratio of the sample size to the population size, expressed as n divided by N. If a researcher selects 500 households from a population of 50,000 households, the sampling fraction is 500/50,000, or 1 per cent. This fraction helps in calculating weights and in adjusting estimates, especially when different subgroups are sampled at different rates, as often happens in large demographic and health surveys.
Population parameter and sample statistic
A population parameter is a numerical value that describes a characteristic of the entire population, such as the true average age at marriage among all women in India. A sample statistic is the corresponding value calculated from the sample. The whole purpose of inferential statistics is to use sample statistics to estimate population parameters, while acknowledging some degree of error. Common parameters and their sample counterparts include the population mean (ฮผ) estimated by the sample mean (xฬ), and the population standard deviation (ฯ) estimated by the sample standard deviation (s).
Sampling error and sampling bias
Two more terms are worth knowing. Sampling error is the natural difference between a sample statistic and the true population parameter that arises purely because we studied only a part of the population. It can be reduced by increasing sample size and improving design. Sampling bias, on the other hand, is a systematic error caused by flaws in the selection process, such as an incomplete sampling frame or self-selection by respondents. Selection bias is the consistent divergence of a sample value from the corresponding population value due to an improper selection process, and it often stems from an incomplete sampling frame or improper selection rules.
Advantages of sampling
Sampling is not just a compromise made when a full census is impossible. In many situations, it is actively preferred over studying the entire population. The reasons are practical, statistical, and ethical.
Cost-effectiveness
Studying a full population is expensive. Field staff have to be hired and trained, transport arranged, instruments printed, and data entered for every single respondent. Sampling reduces these costs dramatically. Sampling saves money by allowing researchers to gather the same answers from a sample that they would receive from the population, and is almost always more cost-effective than a full census. For a small NGO working on adolescent health in a district, a well-designed sample of a few hundred adolescents can yield insights that would otherwise require crores of rupees to collect through a census.
Time-saving and quick data collection
Decisions in public health often cannot wait. During a disease outbreak, a rapid sample survey can provide actionable estimates of case prevalence within weeks, while a full enumeration might take years. Even for routine indicators, sampling allows researchers to release timely findings. The National Family Health Survey rounds, for example, generate state and national estimates within a reasonable time frame precisely because they rely on carefully drawn samples rather than full enumeration of every household.
Fewer non-sampling errors
It may sound counterintuitive, but a smaller, well-managed study can be more accurate than a full census. With fewer units to track, supervisors can train enumerators better, monitor data quality more closely, and follow up on missing or inconsistent responses. The principal advantages of sampling as compared to complete enumeration of the population are reduced cost, greater speed, greater scope and improved accuracy, because sources of error connected with reliability of field workers, clarity of instruction, and recording mistakes can be controlled more effectively when working with manageable sample sizes. These avoidable mistakes are called non-sampling errors, and they often outweigh sampling errors in large surveys.
Greater scope and detailed study
Because resources are concentrated on a smaller group, researchers can ask more detailed questions, conduct longer interviews, and collect biomarkers or clinical measurements that would be impossible at the scale of an entire population. For example, in a sample-based maternal health study, researchers may be able to measure haemoglobin, blood pressure, and dietary intake in detail, generating richer data than a brief census-style questionnaire ever could.
Feasibility in destructive or sensitive studies
Some studies simply cannot be conducted on the entire population. In quality testing of vaccines or contraceptive pills, the testing destroys the product, so only a sample can be tested. In sensitive areas such as mental health or sexual behaviour, building trust and ensuring confidentiality is easier with a smaller, well-trained team working with a sample.
When sampling is not appropriate
Despite its many advantages, sampling is not always the right choice. When the population is very small, it may be easier and more accurate to study everyone. When the research question demands information about every single unit, such as a village-level census for planning, a complete enumeration becomes necessary. Sampling also fails when the sampling frame is so flawed that no amount of statistical adjustment can fix the resulting bias. Recognising these limits is as important as understanding the strengths.
What do you think? If you were designing a study to understand contraceptive use among young married women in your district, what would you choose as your sampling unit and sampling frame, and what challenges might you face in making the sample truly representative?
References
- https://www.scribbr.com/frequently-asked-questions/what-is-sampling/
- https://main.mohfw.gov.in/sites/default/files/NFHS-5_Phase-II_0.pdf
- https://online.stat.psu.edu/stat100/book/export/html/642
- https://www.questionpro.com/blog/sampling-frame/
- https://dhsprogram.com/pubs/pdf/DHSM4/DHS6_Sampling_Manual_Sept2012_DHSM4.pdf
- https://www.sciencedirect.com/topics/mathematics/sampling-frame
- https://www.cloudresearch.com/resources/guides/sampling/what-is-the-purpose-of-sampling-in-research/
- https://main.mohfw.gov.in/basicpage-14
- https://csr.education/csr-projects-programmes/understanding-sampling-in-research/

Leave a Reply