Behind every health policy, family planning programme, and demographic study lies one quiet foundation: data. But where does this data actually come from, and how does it travel from a household conversation in a remote village to a published statistical report? The journey is more layered than most people realise. Understanding how data is sourced, who collects it, and when it is gathered is what separates trustworthy research from misleading conclusions.
Table of Contents
- What does sourcing data really mean?
- Stages of data sourcing
- Primary sourcing: data straight from the field
- Secondary sourcing: working with what has already been gathered
- Tertiary sourcing: summaries and overviews
- How data transitions from raw collection to analysed reports
- Methods of data collection
- The census method
- The sample survey method
- Registers and registration systems
- The importance of authority: who collects the data matters
- The importance of timing: when the data was collected
- Recency
- Reference period
- Seasonality and context
- Bringing it all together
What does sourcing data really mean?
Sourcing data is the structured process of locating, gathering, and preparing information for research use. In population and family health studies, this involves more than just running a survey or downloading a report. It includes deciding what kind of source is appropriate, evaluating its credibility, and tracing the path the data has taken before reaching the researcher’s hands.
Researchers typically work with three broad categories of sources, each representing a different stage in the life of a dataset. Primary, secondary, and tertiary sources represent progressive distance from the original event or observation, and each plays a distinct role in shaping how we understand population health.
Stages of data sourcing
Think of data as moving through a pipeline. At one end, an enumerator records a birth in a sample village. At the other end, a student reads a summary in an encyclopaedia entry on demographic transition. Between these two points sit several stages of processing, analysis, and republication.
Primary sourcing: data straight from the field
Primary sources are original, firsthand records generated by the researcher or agency directly observing the phenomenon. In population studies, these include raw census schedules, individual responses to a household survey, hospital birth registers, and field notes from a community health worker. Primary sources are the original documents of an event or discovery, and they are usually considered the most credible because there is no intermediate filter between the event and its record.
The strength of primary sourcing is control. The researcher designs the questions, chooses the respondents, and decides how the data is recorded. The trade-off is cost, time, and logistical complexity. Conducting a primary survey across rural and urban households requires trained field staff, supervision, quality checks, and significant funding.
Secondary sourcing: working with what has already been gathered
Secondary sources contain data that has already been collected, processed, and often analysed by someone else. A researcher using National Family Health Survey (NFHS) tables to study maternal anaemia in West Bengal is using a secondary source, even though the underlying household interviews were originally primary data. Secondary sources are good for gaining a full overview of your topic and for accessing information the researcher cannot collect independently.
Secondary sourcing is particularly important in population studies because the scale of data needed, millions of households across diverse states, is rarely within the reach of any single research team. Reports from the Office of the Registrar General, the Ministry of Health and Family Welfare, and the National Statistical Office become indispensable.
Tertiary sourcing: summaries and overviews
Tertiary sources sit further down the pipeline. They compile, condense, or index information from primary and secondary sources without contributing new analysis. Examples include demographic encyclopaedias, statistical yearbooks, textbook chapters, and Wikipedia entries. Tertiary sources further repackage the original information used in secondary sources by indexing or condensing it.
Tertiary sources are useful in the early, exploratory stage of research when a student is trying to map the territory of a topic. They are rarely cited directly in academic work because they are too far removed from the original observation to serve as evidence.
How data transitions from raw collection to analysed reports
The movement from primary to tertiary is not just about distance. It involves real transformation. Raw field schedules are coded, cleaned, and entered into databases. Statistical software then produces frequency tables, cross-tabulations, and population estimates. These outputs are reviewed, contextualised, and published as reports. Journalists, textbook authors, and encyclopaedia editors then summarise these reports for wider audiences.
At each transition, two things happen. Some detail is lost, since aggregation hides individual variation. And some interpretation is added, because every analyst makes choices about what to highlight. A careful researcher always tries to trace data back as close to its primary source as possible, because that is where credibility lives.
Methods of data collection
Within primary sourcing, three methods dominate population and family health studies: censuses, sample surveys, and registration systems. Each answers a different kind of question and carries its own strengths and limitations.
The census method
A census attempts complete enumeration. Every person, every household, every unit in the defined population is counted. The census of India plays a vital role in the collection of accurate and reliable data from all over India and serves as the framework for almost all national-level surveys. The Census of India is one of the largest administrative exercises in the world, conducted decennially since 1872.
The advantages of a census are clear. It produces highly accurate population counts at every administrative level, from the country down to the village. It captures rare characteristics that a sample might miss. And it provides the sampling frame that every other survey relies on. Census provides the frame population for weights used in all surveys, which is why countries continue to invest in it despite the availability of large surveys.
The disadvantages are equally significant. A census is extraordinarily expensive, takes years to plan and execute, and faces real risks of non-sampling errors such as data entry mistakes, enumerator fatigue, and respondent recall issues. Because of the long gap between censuses, the data quickly becomes dated for fast-changing indicators.
The sample survey method
A sample survey collects data from a scientifically chosen subset of the population. The findings are then generalised to the whole using statistical inference. In India, the National Sample Survey Office (NSSO) and the National Family Health Survey are leading examples. The fifth round of the NFHS, conducted between 2019 and 2021, gathered data from more than six lakh households across the country.
Sample surveys are cheaper, faster, and more flexible than censuses. They can be repeated frequently, focus on specialised topics, and produce high-quality estimates when the sample is properly designed. The main limitation is sampling error, the inevitable gap between sample estimates and the true population value. Well-designed surveys measure and report this error transparently. Surveys also struggle to produce reliable estimates for very small geographic areas or rare sub-populations.
Registers and registration systems
Registers continuously record events as they happen. The Civil Registration System (CRS) records every birth and death under the Registration of Births and Deaths Act, 1969, while the Sample Registration System combines continuous enumeration with periodic retrospective surveys in selected units. Hospital records, school enrolment registers, and electoral rolls are other examples.
The strength of a register lies in its continuity. Unlike a survey or census that captures a moment in time, a register captures change as it occurs. The field investigation under the Sample Registration System consists of continuous enumeration of birth and death events by a resident part-time enumerator alongside an independent six-monthly retrospective survey by a full-time supervisor. The two streams are then matched to produce an unduplicated count, making it a self-evaluating dual record system.
The weakness of registers is coverage. Civil registration historically suffered from under-reporting in rural areas, particularly for female births and infant deaths. Even today, completeness varies across states. Registers also capture only the events they are designed for, so they cannot tell you why something happened, only that it did.
The importance of authority: who collects the data matters
Two surveys can ask the same question and produce different answers. The difference often comes down to who collected the data, with what training, and under what institutional oversight.
Authoritative sources in population and family health studies in India include the Office of the Registrar General and Census Commissioner, the Ministry of Health and Family Welfare, the International Institute for Population Sciences, and the Ministry of Statistics and Programme Implementation. These bodies follow standardised protocols, train their enumerators, audit their data, and publish detailed methodology notes. The National Sample Survey conducted by MoSPI uses a multistage stratified design with carefully defined sampling units, ensuring that the results are comparable across rounds and regions.
When data comes from an unknown source, an NGO without a published methodology, or a media report without primary attribution, the researcher faces an uncomfortable question: how was this number produced? Without that answer, the data cannot be trusted for serious analysis. Authority is not about prestige; it is about verifiable methodology.
The importance of timing: when the data was collected
Timing affects data in three distinct ways, and each matters for population and family health research.
Recency
Population characteristics change continuously. Fertility rates, urbanisation, child immunisation coverage, and disease patterns can shift substantially within a few years. The most recent NFHS round was the sixth in 2023-2024, with the data yet to be released, while the fifth round drew on fieldwork from 2019 to 2021. Using a 2011 estimate to plan a 2026 intervention can produce serious miscalculations, particularly in rapidly transitioning states.
Reference period
Every dataset has a reference period: the specific window the data describes. A survey conducted in 2024 may ask about births in the last five years, deaths in the last twelve months, or current contraceptive use. Mixing reference periods without care can produce nonsensical conclusions, such as comparing infant mortality from one year with maternal mortality calculated over a different window.
Seasonality and context
The time of year and the broader social context also matter. Surveys conducted during festival seasons, agricultural peaks, or public health emergencies can produce results that do not represent normal conditions. Researchers studying the impact of the COVID-19 lockdowns on maternal health care, for example, have to be careful not to treat 2020 data as a baseline for ordinary years.
Bringing it all together
Good data sourcing in population and family health studies is rarely about finding one perfect number. It is about understanding the chain of decisions that produced that number: who designed the study, who collected the responses, when the fieldwork happened, how the data was processed, and where the findings were published. A researcher who can trace this chain, evaluate each link, and combine sources thoughtfully is the one who produces work that decision-makers can actually rely on.
The richest evidence usually comes from triangulating multiple sources, combining a census frame, a recent sample survey, and a continuous registration system, so that the strengths of one method cover the weaknesses of another. This is how reliable population estimates, fertility trends, and mortality indicators are built in practice.
What do you think? If you were designing a study on adolescent reproductive health in your state, would you rely more on the Census, the NFHS, or the Civil Registration System, and why? And how would the timing of your data collection change the kinds of conclusions you could draw?
References
- https://www.scribbr.com/working-with-sources/primary-and-secondary-sources/
- https://libguides.umflint.edu/idinfosources/primarysecondary
- https://ohiostate.pressbooks.pub/choosingsources/chapter/primary-secondary-tertiary-sources/
- https://www.ijcmph.com/index.php/ijcmph/article/view/14898
- https://www.sunriseclassesiss.com/post/q-what-is-the-difference-between-census-and-sample-surveys-in-official-statistics-and-why-does-ind
- https://srs.census.gov.in/
- https://censusmp.gov.in/censusmp/english/srs.html
- https://www.mospi.gov.in/national-sample-survey-nss
- https://www.ijcmph.com/index.php/ijcmph/article/download/14898/8762/72588

Leave a Reply