View/Export Results
Manage Existing Surveys
Create/Copy Multiple Surveys
Collaborate with Team Members
Sign inSign in with Facebook
Sign inSign in with Google
Skip to article content

What is sampling in research?

Definitions, the 9 sampling methods compared side by side, and how to choose one you can defend.

Short answer

Sampling is the process of selecting a subset of a population, called a sample, and studying that subset to draw conclusions about the whole population. Researchers sample because measuring everyone is usually too slow, too expensive, or simply impossible. The population is the group you want to describe, the sample is the part of it you actually measure, and the sampling frame is the list you draw the sample from.

Probability sampling selects people using a random mechanism, so every unit has a known, non-zero chance of being chosen. That is what lets you attach a number to your uncertainty.

Non-probability sampling selects on convenience, quotas or judgement. It is often faster, cheaper and sometimes the only ethical option, but it cannot support a population estimate with a stated precision.

Go straight to the 9 methods compared | how to choose one | test yourself | how to write it up

Key takeaways

  1. Sampling means studying a subset: you measure a sample to learn about a population when measuring everyone is not feasible.
  2. Get five terms right first: population, sampling frame, sampling unit, sampling design, and the gap between the population you want and the one you can actually reach. Most arguments about sampling are really arguments about one of these.
  3. Only probability methods let you attach a number to your uncertainty. Convenience, quota, purposive and snowball samples can be useful and honest, but they cannot support a population estimate with a stated precision.
  4. Your frame sets your ceiling: if a group cannot appear on the list you draw from, no method and no volume of extra responses will put them in your results.
  5. Representativeness is a chain, not a label: frame, then selection, then recruitment, then response, then analysis. A random draw is one link out of five.

Sampling definition (and the 5 terms you must get right)

Sampling is the process of selecting a smaller group (a sample) from a larger group (a population) so you can learn about the population without measuring everyone. A full-population study is often impractical, so researchers sample to make the work feasible while still producing useful conclusions.

The five terms below do most of the work in any methodology chapter. Get them wrong and every later decision inherits the mistake.

  • warning
    Population: the full group you want to describe, written with boundaries and a time window. "All final-year nursing students at the three campuses, enrolled as of 1 March" is a population. "Nurses" is not.
  • warning
    Target population vs accessible population: who you want to describe versus who your recruitment and tools can actually reach. Naming the gap between the two is a strength in a write-up, not a weakness.
  • warning
    Sampling frame: the list or mechanism you draw from. A student register, a patient register, a customer database, a village household listing, a voter roll. Frame problems create coverage error, and no method repairs them.
  • warning
    Sampling unit: what actually gets selected. A person, a household, a clinic, a class, a shift, an order. In cluster and multistage designs the unit at the first stage is a group, and the group is called the primary sampling unit.
  • warning
    Sampling design: the whole plan written down. Frame, method, how many units, how they are contacted, and what you will do about the people who do not answer. "Sampling design" is what an examiner or a reviewer is actually asking for when they ask about your sampling.
Five-step sampling chain from frame to selection to recruitment to response to analysis, with the bias that can enter at each step labelled underneath.
Representativeness is produced or destroyed at five points, not one. A random draw only protects the second link.
A quick reality check

You can run a textbook random selection and still finish with a biased sample. If your frame misses a segment, they were never eligible to be drawn. If a segment does not answer, they drop out after being drawn. Treat sampling as a chain: frame, then selection, then recruitment, then response, then analysis.

For a short official definition of the same idea, see the U.S. Office of Research Integrity training note on sampling.1 If you are placing sampling inside a wider study plan, our guide to survey research design covers the steps either side of it.

Why sample at all, and when a census is the better answer

The need for sampling is practical before it is statistical. Four reasons come up again and again.

  1. Cost

    Every extra respondent costs money, staff time, or goodwill. A well-drawn sample of 600 can answer the same question as a badly run attempt at 60,000.

  2. Speed

    Decisions have deadlines. A sample gives you an answer inside the window where the answer is still useful.

  3. Feasibility

    Some populations cannot be enumerated at all: passers-by at a market, people who abandoned a service, informal traders. Others are effectively infinite, or measuring a unit destroys it.

  4. Quality

    This one surprises people. Concentrating effort on a smaller group lets you chase non-responders, train interviewers properly, and check the data. Effort per unit is usually worth more than units.

A census measures everyone instead. It is the better choice when the population is small enough to reach, when contact details already exist, or when every member has a right to be asked. An employee survey in a 200-person organisation is a census question, not a sampling question. So is a class evaluation. When the population is small, sampling buys you almost nothing and costs you the ability to report by team.

What a good sample actually looks like

Textbooks describe a good sample as "representative", which is not an operational test. These four properties are.

  • Coverage: everyone in the target population could in principle have been selected.
  • Controlled selection: who ends up in the sample is decided by a rule, not by who volunteers or who is nearest the door.
  • Enough units in the smallest group you must report: an overall number can look healthy while the group you actually care about rests on 11 responses.
  • Describable: you can write down exactly how it was drawn, and someone else could repeat it.
What your sampling choice lets you claim
If your goal isSampling implicationCommon trap
Estimate a population percentageUse a probability method, track who answers, and report precision alongside the number.Treating a convenience sample as if it described the population.
Compare subgroupsPlan the number of responses per subgroup, usually by stratifying and oversampling small groups.Reporting a subgroup whose result rests on a handful of people.
Explore experiences and mechanismsPurposive or theoretical selection is appropriate. Justify which perspectives are covered and which are not.Chasing a numeric minimum with no rationale for it.
Test an interventionSampling governs who the result generalises to. Random assignment governs whether the effect is real.Confusing random assignment with random sampling. They are different operations.

For a plain-language account of how samples support population estimates, the U.S. Census Bureau overview of survey sampling and inference is a good starting point.2

Your sampling frame sets the ceiling

A sampling frame is the concrete list or mechanism you draw from. It is the most under-discussed part of sampling and the place where most real studies quietly go wrong, because the frame decides who was even eligible to be selected.

Frames are ordinary objects. A student register, a patient register, an HR roster, a customer database. A list of villages with household counts, a voter roll, a run of orders, a list of clinics in a district. Notice that some frames list people and some list places. That distinction decides which methods are open to you.

The four things that go wrong with a frame
ProblemWhat it looks likeWhat it does to your results
UndercoveragePeople in the population are missing from the list: staff without work email, customers who bought in person, households without a registered address.Whole segments have a zero chance of selection. Extra responses cannot fix it, because the missing people were never eligible.
OvercoverageThe list contains units outside your population: former students, closed accounts, people who moved away.You waste effort and, if you do not screen them out, you measure the wrong population.
DuplicationThe same person appears more than once: two email addresses, two account numbers.Those people get a higher chance of selection than everyone else, which quietly breaks the equal-chance assumption.
StalenessThe list was accurate when it was built and is not now.Bounced contacts look like refusals, so your response rate and your bias diagnosis both come out wrong.

Audit the frame before you draw, not after. Count the rows. Compare that count with any independent total you can find, such as an HR headcount or an enrolment figure. Look for the segment you would be most embarrassed to have missed, and check that it is present. Deduplicate. Write down what the frame excludes, because that sentence belongs in your limitations either way.

Which sampling methods do not need a frame?

Convenience, quota, purposive and snowball sampling need no list at all. Cluster and multistage sampling need only a list of groups (schools, clinics, wards, villages), not a list of people. Simple random, systematic and stratified sampling all need a full list of individuals before you can start. If you have no list and cannot build one, that single fact rules out three of the nine methods on this page. It is usually the first thing to settle.

The 9 sampling methods compared

Sampling methods split into two families. Probability sampling uses a random mechanism, so every unit in the frame has a known, non-zero chance of selection. Non-probability sampling does not, so selection probabilities are unknown. Everything else follows from that one distinction.

The second column is the one most guides leave out, so it comes first here. Only a probability design supports a margin of error, a confidence interval, or the sentence "plus or minus 5 points". If a method says no there, you can still learn a great deal from it. You just cannot publish a precision figure alongside the result. Precision itself belongs to a different conversation, covered on our sample size guide.

Nine sampling methods: how each one works, what it costs you, and whether it supports a margin of error claim
Method Supports a margin of error? How it works Right choice when What it costs, and the bias it invites
Simple random sampling
Probability Needs a full list
Yes Every unit on the list has an equal chance. You draw the required number at random. You have one clean, deduplicated list and no subgroup you are obliged to report separately. CostYou must have the full list, and small groups can come out too small to report on.BiasVery little at selection. All the remaining risk moves to who actually answers.
Systematic sampling
Probability Needs a full list
Yes, if the list order is unrelated to the outcome Sort the list, pick a random starting point, then take every kth record. The list is long or arrives as a stream, and a field team needs one simple rule to follow. CostAlmost nothing, as long as the ordering has nothing to do with what you are measuring.BiasPeriodicity. If the list repeats on a cycle that lines up with your interval, you keep hitting the same kind of unit.
Stratified sampling
Probability Needs a full list
Yes, and usually tighter than simple random Split the list into groups (strata) that matter, then sample inside every group. You must report by region, faculty, department or age band, or you expect response to differ sharply between groups. CostThe grouping variable has to be on the list before you draw, and unequal rates have to be weighted at analysis.BiasLow. The real risk is choosing strata that have nothing to do with the outcome, which adds work and buys no precision.
Cluster sampling
Probability Needs a list of groups
Yes, widened for the design effect Randomly select whole groups (schools, clinics, wards, villages), then survey everyone inside the selected groups. You cannot list individuals but you can list the places they are found, and travel or set-up cost dominates the budget. CostPrecision. People inside a group resemble each other, so each extra response carries less new information.BiasLow if groups are chosen at random, but a small number of groups makes the result unstable and hard to defend.
Multistage sampling
Probability Needs a list of groups
Yes, with weights and a design effect Sample groups first, then sample units inside the selected groups, in two or more stages. District-wide or national studies where no single list of people exists but administrative units do. CostComplexity. Selection probabilities differ between units, so the analysis needs weights and somebody who can apply them.BiasLow in principle. In practice every stage is a fresh opportunity for the field team to drift from the plan.
Convenience sampling
Non-probability No list needed
No You take whoever is easiest to reach: a pop-up on a website, a class you teach, people passing a stall. Also called availability sampling. You are exploring, pilot testing a questionnaire, or you need direction this week rather than a defensible number next quarter. CostThe claim. You give up the ability to say anything about the population.BiasSelf-selection. It over-represents the reachable, the available and the already engaged.
Quota sampling
Non-probability No list needed
No. It looks balanced without being random Set targets for groups, for example 50 women and 50 men, or 40 rural and 60 urban, then recruit until every target is filled. You need a sample that at least matches the population on visible characteristics, quickly, with no list to draw from. CostControl of who fills each quota. Recruiters still choose, and they choose the approachable.BiasWithin-quota selection bias. Matching on age and gender does not make people typical on attitudes.
Purposive sampling
Non-probability No list needed
No, and it is not trying to You deliberately select cases that fit stated criteria: experts, extreme cases, typical cases, critical cases. Also called judgement sampling. Qualitative work, expert interviews, case studies, and any question where depth matters more than counting. CostNothing, provided you are clear that you were never counting.BiasThe researcher's own model of who counts as relevant.
Snowball sampling
Non-probability No list needed
No Recruit a few participants, then ask each of them to refer others who fit the criteria. The group is hidden, stigmatised, rare or unregistered: migrant workers, people who left a service, informal traders, people with a rare condition. CostTime, and control over the shape of the sample. You inherit the social network of whoever you started with.BiasHomophily. People refer people like themselves, so one corner of the population gets over-sampled.

Reading the table: the first five methods are probability designs and the last four are not. That single split, not the individual method names, is what decides whether a precision figure is honest.

Probability sampling methods, one by one

Probability sampling means every unit in the frame has a known, non-zero chance of being selected, and a random mechanism does the selecting. That is the whole basis for saying anything numerical about the population. Five designs cover almost every real study. Clinical and health research reviews cover the same family under slightly different labels, and it is worth knowing why before you meet both: each of the two cited here folds multistage sampling into cluster sampling rather than listing it separately.34

Simple random sampling

Simple random sampling gives every unit on the list exactly the same chance of selection. You number the list, generate random numbers, and take the units they land on.

Worked example. A university register holds 12,000 enrolled students and you need 400 responses. Number the register 1 to 12,000, generate 400 unique random numbers in that range with a spreadsheet or a statistics package, and invite those students. If you expect a 40 per cent response rate, draw and invite 1,000 so that 400 completed responses is realistic.

Where it fails. With 400 responses spread across a 12,000-student register, a department of 300 students contributes about 10 people. That is not enough to report on. If any subgroup has to appear in your results, simple random sampling is the wrong tool and stratified sampling is the fix.

Write down your randomisation. Software used, the seed if there was one, and the date. That single sentence is what makes the draw reproducible, and reviewers do ask for it.

Systematic sampling

Systematic sampling takes every kth unit from an ordered list after a random start. Divide the list size by the number you need to get the interval k, pick a random number between 1 and k as your starting point, then count forward.

Worked example. A district hospital records 9,000 outpatient visits in a quarter and wants 600 exit interviews. The interval is 9,000 divided by 600, so k equals 15. Pick a random start between 1 and 15, say 7, then interview visit 7, 22, 37, 52 and so on. A field team can follow that rule without a laptop, which is the real reason this method survives.

Where it fails. Periodicity. If the list has a repeating cycle that lines up with k, you sample the same slot every time. A visit list ordered by clinic day with a 7-day cycle and an interval of 7 would give you one weekday and nothing else. Check the ordering before you fix k, and shuffle the list if the order relates to what you are measuring.

Stratified sampling

Stratified sampling divides the population into non-overlapping groups called strata, then draws a random sample inside every stratum. It is the most useful upgrade available to most studies, because it guarantees that each group you must report on actually appears.

Worked example with the arithmetic shown. An organisation has 4,200 staff across six provinces and needs 600 responses.

  • Proportional allocation. A province with 1,400 staff holds 1,400 divided by 4,200, which is one third of the workforce, so it gets one third of 600, which is 200 responses. A province with 210 staff gets 210 divided by 4,200 times 600, which is 30 responses. Every province is represented in proportion to its size, and no weighting is needed at analysis.
  • Disproportionate allocation, also called oversampling. Thirty responses is too few to report on that small province separately. So you deliberately take 80 from it instead, and reduce the largest province to compensate. Now selection probabilities differ between provinces, so you must weight the data back to true proportions before quoting any organisation-wide figure. Skipping that step is the single most common stratified sampling error.

Choosing strata. Pick variables that relate to the outcome or to who is likely to respond, and that already exist on your frame. Region, site, role, tenure band, course, plan type and age band are the usual candidates. Strata that have nothing to do with the outcome add administration and buy no precision. If the grouping variable is missing from your frame, you may need to collect it inside the questionnaire instead. Our guide to demographic survey questions covers how to ask for those variables without losing respondents.

Cluster sampling

Cluster sampling randomly selects whole naturally occurring groups and then surveys everyone inside the selected groups. You need a list of the groups, not a list of the people, which is why it exists.

Worked example. A district has 180 primary schools and no central list of pupils. You randomly select 24 schools and survey every Grade 8 pupil in each one. You have made 24 school visits instead of 180, and you never needed a pupil register to start.

What you pay for it. Pupils in the same school resemble each other more than they resemble pupils across the district. So each extra pupil in a school you have already picked adds less new information than an independent pupil would. Your effective number of responses is smaller than your actual number of responses. That inflation is the design effect, and it has to be built into the plan rather than discovered at analysis.

Where it fails. Too few clusters. Six schools chosen at random can differ wildly from each other, and the result swings on which six you happened to draw. More clusters with fewer people in each is almost always better than the reverse.

Multistage sampling

Multistage sampling selects in two or more stages: groups first, then units inside the selected groups, and sometimes a third stage inside those. Almost every national household survey in the world is built this way.

Worked example. Select 120 wards from the national list of wards. Inside each selected ward, a field team lists the households on the ground, then randomly selects 20 of them. Inside each selected household, one adult is chosen at random rather than left to whoever opens the door. That is three stages, and the last one matters: letting the household decide who answers reintroduces exactly the selection bias the design was built to avoid.

What you pay for it. Different people end up with different selection probabilities, so the analysis needs weights, and the variance calculation has to account for the clustering at every stage. This is a real design that needs someone who can carry out a weighted analysis. If nobody on the project can, choose a simpler design rather than an unweighted multistage one.

Expert depth: probability proportional to size, and why big clusters get picked more often

In the school example every school had the same chance of selection, whether it held 60 pupils or 900. That gives a pupil in a small school a much higher chance of reaching the sample than a pupil in a large one. The analysis then has to correct for it.

Probability proportional to size selection fixes this at the design stage. Each cluster's chance of selection is set in proportion to how many units it holds. A 900-pupil school is then fifteen times more likely to be drawn than a 60-pupil school. Then you take the same number of pupils from every selected school. The two stages cancel out, and every pupil in the district ends up with roughly the same overall chance of selection. It is the standard approach in national surveys, and it needs a size measure for every cluster on the frame before you draw.

Stratified vs cluster sampling: the difference in one line

These two get confused constantly, because both start by dividing the population into groups. The difference is what you do next.

Stratified sampling: a few from every group Four groups of twelve units. Three units are selected from each of the four groups, so every group contributes and twelve units are selected in total. Stratified a few from every group selected not selected Cluster sampling: everyone, from the groups you pick The same four groups of twelve units. One whole group is selected and all twelve of its units are taken, so three groups contribute nobody and twelve units are selected in total. Cluster everyone, from the groups you pick selected not selected
Diagram The same population of four groups, twelve units each, sampled two ways. Both diagrams select twelve units in total. Stratified takes a few from every group, so every group is represented. Cluster takes everyone in the groups it picks, and the groups it does not pick contribute nobody, which is why a clustered sample of twelve carries less information than a stratified sample of twelve. Group counts here are illustrative, not measured.
Stratified and cluster sampling compared
AspectStratified samplingCluster sampling
Which groups you useEvery group is sampled from.Some groups are selected, the rest are not used at all.
Who you sample inside a groupA random subset of the group.Everyone in the group, or a further random subset in a multistage version.
Groups should beInternally similar, and different from each other. Region, role, course.Internally varied, each one a small copy of the population. Schools, clinics, wards.
Why you would choose itPrecision, and guaranteed subgroup representation.Cost and reachability when no list of individuals exists.
Effect on precisionUsually improves it compared with simple random at the same number of responses.Reduces it. Budget for the design effect.
What you need before you startA full list of individuals, with the grouping variable already on it.A list of groups only.
A hook that survives an exam

Stratified: all the groups, some of the people. Cluster: some of the groups, all of the people. Stratified is about accuracy, cluster is about access.

Non-probability sampling methods, one by one

Non-probability sampling means selection probabilities are unknown, because no random mechanism decided who was chosen. It is the right choice in four situations. You cannot build a usable frame. The group is hidden. Speed matters more than precision. Or the question is about meaning rather than size.

These methods are not second-class. They are answering a different question. The failure is never using them, it is using them and then writing a sentence that only a probability sample could support.

Convenience sampling

Convenience sampling, also called availability sampling, recruits whoever is easiest to reach. A pop-up on a website, a link in a newsletter, students in your own class, people passing a stall in a market.

Example. A shop places a feedback link on the receipt. Everyone who answers has bought something, was carrying a phone, had time, and had a reason to bother. Each of those filters removes a different part of the customer base.

Honest use. Pilot testing a questionnaire, generating hypotheses, finding out what language customers use, checking that a question is understood. All excellent uses. None of them require a population estimate.

Quota sampling

Quota sampling sets recruitment targets for named groups and fills them. Forty rural and sixty urban. Fifty women and fifty men. Twenty in each of five age bands.

Example. A market research team needs 500 responses matching the national split on age, gender and region, and has no list of the population. They recruit until each cell is full.

Why it is still not probability sampling. The quotas control the shape of the sample, not the selection of the people in it. Whoever recruits still chooses, and people who are easy to approach are systematically different from people who are not. A sample can match the population perfectly on age and gender and still be badly wrong about attitudes, because age and gender were never the variables doing the damage.

Purposive sampling

Purposive sampling, also called judgement sampling, deliberately selects cases that meet stated criteria. Common variants are typical-case, extreme-case, critical-case, maximum-variation, and expert selection.

Example. A study of how a new payment system is used selects eight people. Two who use it daily, two who gave up on it, two who never took it up, and two support staff who handle complaints. That spread is chosen on purpose, and the reasoning is the method.

Honest use. Most qualitative research. The standard here is not randomness. It is whether you can say why these cases and not others, and whether the range of views is wide enough to support your claims.

Snowball sampling

Snowball sampling recruits a small number of participants who then refer others who fit the criteria, and those refer more.

Example. A study of informal cross-border traders begins with four participants identified through a local association, and each one introduces two more. There is no register of informal traders, so no frame-based method is available.

What it does to your sample. Homophily. People refer people like themselves. The sample drifts toward one social network and away from isolated members of the population, who are often the ones the research most needs to hear. Starting from several unconnected seeds rather than one reduces the drift but does not remove it.

How to make a non-probability sample defensible

Most guides stop at "you cannot generalise" and leave the reader stranded, because most readers are running exactly this kind of study. Here is the part that usually goes missing.

  • Describe the recruitment, not just the method name. Where the link was posted, over what dates, who saw it, whether an incentive was offered, and how many were approached versus how many answered. This is what lets a reader judge the sample instead of guessing.
  • Compare your achieved sample against a known benchmark and publish the comparison. If your organisation is 60 per cent field staff and your respondents are 25 per cent field staff, say so in a table. A stated skew is a finding. A hidden one is a defect.
  • Know what weighting can and cannot do. Adjusting your data so the demographics match the population can correct visible imbalance. It cannot correct for people who differ from respondents on something you did not measure, and it never converts a non-probability sample into one that supports a precision claim.
  • Change the sentence, not the statistics. Instead of "38 per cent of customers are dissatisfied, plus or minus 4 points", write "38 per cent of the 620 customers who responded reported dissatisfaction. Respondents were recruited by an in-app prompt over two weeks and are not a random sample of customers." The second sentence is shorter, more useful, and cannot be attacked.

Qualitative work has its own vocabulary for this. A protocol usually asks for a rationale rather than a numeric target. It should cover the intended number and the range of views, a distinction the University of Bath guide states concisely.5

Match the method to the situation

Six situations, six methods. Decide before you open the answer. This is the fastest way to find out whether the table above actually landed.

1. You have a complete staff register of 4,200 people and must publish results for each of six provinces.

Stratified sampling. The full list exists, so probability sampling is open to you, and the obligation to report by province is exactly what stratification is for. Simple random sampling would leave the smallest province with too few responses to report.

2. You are studying people who left a service six months ago. There is no list, and the ones who left angriest are hardest to find.

Snowball sampling. No frame exists and the population is hard to reach, which rules out every probability method. Start from several unconnected seeds rather than one, and state in your limitations that referral chains bias the sample toward connected people.

3. A hospital wants patient experience data from a stream of 9,000 visits a quarter, collected by staff with no computer at the desk.

Systematic sampling. The population arrives as an ordered stream, so an interval rule is both a probability method and something a busy desk can actually execute. Check first that the visit order does not follow a weekly cycle matching your interval.

4. You need to understand why a new tool is being abandoned, and you have three weeks and no budget.

Purposive sampling. The question is about mechanism, not magnitude. Select deliberately across the range: heavy users, abandoners, non-adopters, support staff. Do not report percentages from it.

5. A district has 180 schools, no pupil register, and a travel budget that covers about 25 site visits.

Cluster sampling. A list of schools exists even though a list of pupils does not, and travel cost is the binding constraint. Widen your reported interval for the design effect, and prefer more schools with fewer pupils each over the reverse.

6. You must field 500 responses in five days that match the national age and region profile, and there is no population list.

Quota sampling. It is the fastest way to get a sample that at least matches the population on visible characteristics. It is still not a probability sample, so report the composition and skip the precision claim.

How to choose a sampling method

Most studies do not fail because someone picked the wrong named technique. They fail because the technique did not match the frame, the decision, or what the field team could realistically do. Three questions settle it almost every time.

  1. Start with the claim you need to make

    Write the sentence you want to publish, before you choose anything. "Sixty-two per cent of our customers would recommend us" and "customers who leave describe three recurring problems" are different sentences that need different designs. If nobody will treat the result as a population figure, you have far more freedom than you think.

  2. Audit the frame before you commit

    Can you actually list the population, or a close proxy? If not, can you list the groups they belong to? Answering this one question eliminates most of the nine methods immediately, which is why it goes second and not last.

  3. Name the subgroups you are obliged to report

    If a result has to appear for each region, site or role, that drives the design through strata or quotas. It also drives how many responses you need in each group. Deciding this after fieldwork is not recoverable.

  4. Check the design against the field reality

    Time, budget, contactability, language access, literacy, connectivity and incentive policy decide what is executable. A design a field team cannot follow becomes convenience sampling with a formal name attached, which is worse than choosing convenience sampling honestly. Inclusive recruitment planning reduces systematic underrepresentation.6

Where samples actually break: coverage, selection and nonresponse

Selection is the part everyone worries about and the part that goes wrong least often. The damage usually happens before selection, in the frame, or after it, in who bothers to answer. Both produce clean-looking spreadsheets and wrong conclusions.

Four errors that all get called "sampling problems"
ErrorWhat it looks likeHow to reduce it
Coverage error, a frame problemSome people cannot be selected at all: no work email, an outdated roster, customers who only buy in person, households with no connectivity.Improve or combine frames. Add modes: paper, in person, phone, SMS. Document exactly who is excluded.
Selection bias, a procedure problemThe random rule gets overridden in practice. First come first served. A manager nominates who takes part.Automate the draw. Separate the person who selects from the person who recruits. Audit every exception.
Nonresponse bias, an outcome problemSelected people do not answer, and the ones who do not answer differ from the ones who do on the thing you are measuring.Improve contact strategy and questionnaire experience, then compare respondents against a known benchmark. See response bias.
Measurement errorA flawless sample still gives you nonsense because the questions are leading, vague or double-barrelled.Test the instrument before fielding. Our guide to writing survey questions covers the common faults.

The most famous sampling failure had 2.4 million respondents

In 1936 the Literary Digest predicted Alf Landon would take roughly 55 per cent of the vote. Franklin Roosevelt won about 61 per cent.7 The magazine had mailed 10 million ballots and got about 2.4 million back, a 24 per cent response rate that no national poll today comes close to.8 The poll put Roosevelt twenty points below his actual share, on a sample of two and a half million people.

The number that should be on every sampling page. A later reanalysis computed the sampling margin of error for the Digest's state-level estimates: it ranged from 0.18 to 2.3 percentage points, with a median of 0.6.8 A precision figure of plus or minus half a point, sitting on top of an error of twenty points. Nothing demonstrates more cleanly that a margin of error measures only one of the several ways a survey can be wrong.

And the standard explanation is only half right. Everyone is taught that the Digest sampled car owners and telephone subscribers, who were rich and Republican. One empirical test of that claim used a Gallup survey from May 1937. People who owned both a car and a telephone still broke for Roosevelt, 55 to 45. Among everyone who said they received a Digest ballot, Roosevelt led 55 to 44. The frame leaned Republican, but not nearly enough to produce a Landon landslide.7

The missing half is nonresponse. Among people who said they returned their ballot, Landon led 51 to 48. Among those who did not return it, Roosevelt led 69 to 30. The rough decomposition puts about eleven points of the error on the frame and about seven on differential nonresponse.7 Weighting the returns on a question the Digest had already asked, how you voted in 1932, would have produced the correct winner using arithmetic available in 1936.8

A high response rate is not the target you think it is

Response rates for Pew Research Center's own telephone polls fell from 36 per cent in 1997 to 9 per cent by 2012, and to 6 per cent by 2018.9 Meanwhile India's National Family Health Survey, conducted face to face, selected 664,972 households and found 653,144 of them occupied. It interviewed 636,699 of those, a 97.5 per cent response rate. Read that denominator carefully, because it is the most common place a response rate goes wrong: against the number selected rather than the number occupied, the same fieldwork reads as 95.7 per cent. The survey drew its sample from 30,456 clusters across all 707 districts.10

Those two facts sit in the same decade. Mode and fieldwork effort drive response rates far more than any global claim about survey fatigue does. If a guide tells you the average response rate is some single number, it is not describing a real quantity.

Now the harder finding. A meta-analysis of 235 nonresponse-bias estimates drawn from 30 published studies found the correlation between response rate and the size of nonresponse bias was 0.33. Squared, that means the response rate explains about 11 per cent of the variation in bias.11 A follow-up across 59 studies designed specifically to measure nonresponse bias reached the same conclusion.12

The response rate is worth reporting and worth improving. It is not a certificate. What causes bias is whether the tendency to respond is tied to the thing you measure. A survey about volunteering suffers from that far more than a survey about broadband speed. Chase the correlation, not the percentage.

More responses cannot repair a broken frame

During the 2021 vaccination rollout, two very large surveys tracked United States vaccine uptake. One collected about 250,000 responses a week and overstated first-dose uptake by 17 percentage points against the official benchmark. The other collected about 75,000 every two weeks and overstated it by 14. A third survey of roughly 1,000 respondents, run to standard probability practice, produced reliable estimates.13

The same pattern shows up in ordinary survey work. Comparing three probability panels against three online opt-in sources across 28 benchmarks and 29,937 interviews, average absolute error was 5.8 percentage points for the opt-in samples and 2.6 for the probability panels. On the 25 of those benchmarks that were available for subgroups, the gap for adults aged 18 to 29 was 11.2 against 3.6.14

And a memorable diagnostic: in one opt-in survey, 12 per cent of adults under 30 said they were licensed to operate a nuclear submarine.15 The true figure is close to zero. If a panel can produce that, it can produce anything, and a larger order from the same panel simply buys more of it.

Who your frame misses, outside the US and UK

Most sampling guides quietly assume a Western frame: an email list, a landline, a national address register. Much of the world does not sample that way, and the coverage gaps have a measured shape.

  • An online survey in a low or middle income country is a sample of the connected part of it, and that part skews male. Women in low and middle income countries are 14 per cent less likely than men to use mobile internet, which works out at roughly 235 million fewer women than men online.16
  • Phone samples over-represent household heads. In World Bank phone surveys across four African countries, respondents were 52 per cent male in Uganda and 73 per cent male in Nigeria, against roughly half the adult population. In Uganda 6 per cent of phone respondents were aged 15 to 24, against 38 per cent of the adult population.17
  • What to do about it. Combine modes rather than defending one. List households on the ground where no register exists. Field in the languages your population actually reads, and pre-test the translation rather than back-translating it. Then state the residual gap in your limitations, because it will still be there.

What weighting can and cannot fix

Weighting adjusts your data so the sample composition matches known population figures. It is genuinely powerful and it is routinely oversold.

The case for it. In 2016 the final University of New Hampshire poll had Clinton ahead by 11 points in a race she won by 0.4. It had not adjusted for education. The review committee concluded that this single adjustment would have removed essentially all of the error.18 One variable, correlated with both response and outcome, doing all the work.

The case against relying on it. When more than 30,000 online opt-in interviews were scored against 24 benchmarks, average estimated bias was 8.4 percentage points unweighted. Across every weighting procedure tested, including sophisticated matching approaches, none brought average bias below 6 points.19 Quadrupling the sample size bought a further 0.2 points.

The rule that follows is simple. Weighting corrects imbalance on the variables you weight on, to the extent those variables relate to the outcome. It does nothing about the ways respondents differ from non-respondents on things you never measured, and it never converts a non-probability sample into one that supports a precision claim.

A precision figure is a floor, not a ceiling

A margin of error quantifies sampling variability, and only that. It assumes the sample was drawn at random from the target population and that the remaining bias has been removed. Coverage error, nonresponse bias, measurement error and fraudulent respondents are all outside it.

The evidence is blunt. In the final fortnight of 2020, 348 state-level presidential polls reported a margin of error, and the average reported figure was 3.9 points. Measured against that average, 55 per cent of them missed the final margin by more than one margin of error, and 21 per cent missed by more than twice it.20 These were modern polls with modern methods. Treat the margin of error as a floor on your uncertainty, never a ceiling.

The evidence behind this section, with a confidence grade for each claim
ClaimThe numberGradeSource
Response rate barely predicts nonresponse biasr = 0.33, so about 11% of the variation, across 235 estimates from 30 studiesAGroves, 200611
A very large survey can be badly wrongOverstated vaccine uptake by 17 points on about 250,000 responses a weekABradley and others, 202113
Opt-in samples are roughly twice as inaccurate as probability panels5.8 vs 2.6 percentage points average error over 28 benchmarksAMercer and Lau, 202314
Weighting improves an opt-in sample but does not rescue it8.4 points unweighted, no method below 6.0 pointsAMercer, Lau and Kennedy, 201819
A margin of error understates real polling error55% of 348 state polls in 2020 missed by more than the average reported margin of error of 3.9 pointsAAAPOR task force, 202120
The Literary Digest frame explanation is incompleteCar and telephone owners still favoured Roosevelt 55 to 45BSquire, 19887

Grade A means an official statistic, a meta-analysis, or a replicated finding. Grade B means one strong study whose conclusion has not yet been independently repeated.

Four sampling numbers we deliberately did not use

"A 50 per cent response rate is adequate, 60 is good, 70 is very good." A textbook rule of thumb with no supporting data behind it. Different authorities in the same field give 50, 60 and 80. Bias depends on the correlation between responding and the answer, not on the rate.

"Thirty is enough, because that is when the central limit theorem kicks in." There is no derivation that produces 30. For symmetric outcomes the approximation works far below it, and for skewed outcomes or rare events it fails well above it.

"The average survey response rate is 33 per cent." This circulates across vendor blogs in near-identical wording with no primary source. Mode, sponsor, population and incentive swamp any global average.

"Attention span is now eight seconds, shorter than a goldfish." Traced back to a 2015 marketing report that cited a firm which could not produce a source. The underlying research never mentions eight seconds or goldfish.

Design effect: why clustered samples need more people

Almost every guide recommends cluster sampling for its cost saving and stops there. The saving is real, and so is the bill. People inside the same school, clinic, ward or store resemble each other. They resemble each other more than they resemble people picked at random from across the population. So each extra response from a group you have already picked carries less new information.

The design effect is the number that measures this. In words: how much wider your uncertainty gets compared with drawing the same number of people independently. For a design with equal-sized groups it is

design effect = 1 + (average group size - 1) x intraclass correlation

The intraclass correlation is how alike people inside a group are on the thing you are measuring. Zero means being in the same school tells you nothing about a pupil's answer. One means everyone in a school answers identically, and the school might as well be one person. Real values are small, but small values do a lot of damage once group sizes get big. Pooled across 48 national household surveys, the average intraclass correlation is about 0.06. It varies by what you measure. Fertility items run about 0.01 to 0.03, current contraceptive use 0.03 to 0.05, and medically delivered births as high as 0.22 in the national figures.21

If you have no estimate of your own, borrow a default carefully. The DHS sampling manual falls back on 1.5 for deft when a survey has no figure of its own.22 Deft is not the same quantity as the design effect in the formula above: it is the ratio of the two standard errors, where the design effect is the ratio of the two variances. You square it, so a deft of 1.5 is a design effect of 2.25. The manual calls the value a fallback rather than a typical one, and its own worked examples use 1.40 for modern contraceptive use and 1.22 for child vaccination.

Effective number of independent responses from 960 pupils under four designs Collecting 960 responses gives 960 effective responses with no clustering, 539 at an intraclass correlation of 0.02, 325 at 0.05, and 196 at 0.10, for 24 groups of 40 pupils each. Each bar is drawn against the full 960 collected. The same figures appear in the table below this chart. No clustering: 960 effective responses Groups of 40, intraclass correlation 0.02: 539 effective responses Groups of 40, intraclass correlation 0.05: 325 effective responses Groups of 40, intraclass correlation 0.10: 196 effective responses No clustering Groups of 40, ICC 0.02 Groups of 40, ICC 0.05 Groups of 40, ICC 0.10 960 539 325 196 Effective sample, against the full 960 responses collected
Our own arithmetic Twenty-four schools with 40 pupils each gives 960 completed responses. The design effect is 1 + (40 - 1) x ICC, so at an intraclass correlation of 0.05 the design effect is 2.95 and the 960 responses carry about as much information as 325 independently chosen pupils. Assumptions: equal group sizes, no weighting, a single outcome. Method and formula stated so you can redo it with your own numbers.
The same 960 responses, four designs
DesignDesign effectEffective responsesWhat it means in practice
Independent selection, no clustering1.00960Every response is worth a full response.
24 groups of 40, ICC 0.021.78539You paid for 960 and you can defend the precision of 539.
24 groups of 40, ICC 0.052.95325Two thirds of the fieldwork bought you nothing extra in precision.
24 groups of 40, ICC 0.104.90196The design has quietly become a small study wearing a large study's budget.

Weighting adds its own component. A national survey that oversamples one region and then weights it back carries two design effects at once. A weighting effect of 1.22 and a clustering effect of 1.80 give an overall design effect of 2.20. So 5,000 interviewed households behave like 2,277 independent ones.23

The design lesson. More groups with fewer people in each beats fewer groups with more people in each, almost every time. Forty schools of 24 pupils and 24 schools of 40 pupils cost roughly the same 960 responses, but the first has a smaller design effect and a defensible result. If you are planning how many responses you need in total, work that out on our sample size guide and then multiply by the design effect. Doing it the other way round is how clustered studies end up underpowered.

Three worked examples, start to finish

Method names are easy. The decisions around them are where studies actually go wrong. These three carry a real design from population statement through to the sentence you can publish.

Example 1: patient experience across a district

  • Population. Adults who attended an outpatient consultation at any of the 46 public clinics in the district between 1 April and 30 June.
  • Frame. No district-wide patient register exists. A list of the 46 clinics does, with monthly attendance counts.
  • Method. Multistage. Select 18 clinics with probability proportional to attendance, then systematically sample every 8th exiting patient during two randomly chosen clinic days.
  • Why not something simpler. A full patient list would allow stratified sampling, but it does not exist and cannot be built inside the timeline. A convenience sample at the three largest clinics would be quicker and would describe the three largest clinics.
  • What must be reported. Selection probabilities differ by clinic size, so the analysis needs weights, and the interval needs widening for the design effect.
  • The publishable sentence. "Among adults attending public outpatient clinics in the district during the second quarter, an estimated 71 per cent rated their wait acceptable, based on 1,240 exit interviews across 18 clinics selected with probability proportional to attendance."

Example 2: a dissertation study on a student population

  • Population. Students enrolled in the Faculty of Commerce in the current academic year: 3,100 across three years of study.
  • Frame. The faculty enrolment register, which carries year of study and programme.
  • Method. Stratified random sampling by year of study, proportional allocation, target 350 completed responses.
  • The arithmetic. First year holds 1,300 students, which is 42 per cent of 3,100, so it gets 42 per cent of 350, which is 147 responses. Repeat for each year. Because allocation is proportional, no weighting is required.
  • Response planning. At an expected 35 per cent response rate you need to invite roughly 1,000 students to land 350 completed responses, and you should invite them stratum by stratum so a weak response in one year is visible while there is still time to act.
  • Limits to state. Students who have withdrawn but remain on the register are overcoverage. Students enrolled after the register was pulled are undercoverage. Both belong in the limitations paragraph.

Example 3: customer feedback for an online store

  • Population. Customers who completed an order in the last 90 days.
  • Frame. The order table. It is complete for online orders and complete for guest checkouts, but a customer with two accounts appears twice.
  • Method. Deduplicate to one row per customer, then systematic sampling of every 12th customer after a random start, running continuously rather than as one batch.
  • The trap this avoids. Inviting only newsletter subscribers is faster and is a convenience sample of your most engaged customers. It will make satisfaction look better than it is, reliably, every quarter.
  • A note on employees. The same organisation surveying its own 180 staff should not sample at all. That population is small, reachable, and everyone has a right to be asked. Run a census, and if participation is thin, treat that as a finding rather than a sampling problem.

How to write the sampling part of your methodology

This is the step most guides skip, and it is the one a supervisor, reviewer or client actually reads. A sampling write-up has to answer five questions in order: who, from what list, chosen how, how many, and what you are therefore allowed to claim.

Fill in the blanks below. It is deliberately plain: examiners are not looking for elegant prose here, they are looking for a design they can evaluate.

The target population was who, with boundaries and a time window. The sampling frame was the list or mechanism, and the date it was drawn, which contained number units. Method name sampling was used because the reason, tied to the frame and to the claim you need. Number units were selected and number completed responses were obtained, a response rate of per cent. Any weighting or adjustment applied, or a statement that none was. The frame excluded who was not eligible to be selected, so results should be read as applying to the accessible population, not the target population, if these differ.

The same study, written weakly and written well

Weak. "A random sample of students was taken and 350 responses were collected."

Nothing here can be checked. "Random" is doing work it has not earned, there is no frame, no denominator, and no statement of who was excluded.

Strong. "The target population was the 3,100 students enrolled in the Faculty of Commerce in 2026. The sampling frame was the faculty enrolment register extracted on 3 March 2026. Stratified random sampling by year of study with proportional allocation was used so that each year could be reported separately. One thousand students were invited and 350 completed the questionnaire, a response rate of 35 per cent. No weighting was applied because allocation was proportional. Students who withdrew after the extract date remained on the frame and could not be identified, which is a minor source of overcoverage."

The second version is longer, and every extra sentence is a sentence a reviewer no longer has to ask about.

The limitations paragraph for a non-probability sample

"Participants were recruited by an in-app prompt shown to active users over two weeks. This is a convenience sample, so the findings describe the people who responded and cannot be generalised to the full user base, and no precision figure is reported. Respondents were more likely than the user base to be daily users, which should be weighed when reading the results." Write that, and nobody can accuse you of overclaiming. You did not.

Sampling plan checklist

A written sampling plan prevents most avoidable errors, because it forces the decisions to happen before fieldwork rather than during analysis. Eight points, one page.

  • warning
    1. Research objective. What decision does this support, and what exactly will be reported?
  • warning
    2. Target population. Who is in scope and out of scope, including the time window.
  • warning
    3. Sampling frame. Source, date pulled, row count, known gaps, and the deduplication rule.
  • warning
    4. Method and reason. Which of the nine, and why this one given the frame and the claim.
  • warning
    5. Targets. Completed responses overall and per subgroup, plus the expected response rate you are planning against.
  • warning
    6. Recruitment plan. Modes, reminder schedule, language and accessibility provision, incentive policy.
  • warning
    7. Monitoring. What you will watch during fieldwork: response by stratum, drop-off point, item nonresponse.
  • warning
    8. Analysis notes. Planned weighting, design effect adjustment, and the limitations you will disclose whatever happens.

Once the plan is settled, the instrument decides whether the sample was worth drawing. A perfect sample cannot rescue a leading or double-barrelled question, so it is worth reading through survey question examples before you field anything.

Frequently asked questions

quizWhat is sampling in research?expand_more

Sampling is the process of selecting a subset of a population, called a sample, and studying that subset to draw conclusions about the whole population. It is used because measuring every member of a population is usually too slow, too expensive or impossible. The method used to select the sample determines what conclusions the results can support.

quizWhat is the difference between a population and a sample?expand_more

The population is the full group you want to describe. The sample is the smaller subset you actually measure. Sampling is the process that gets you from one to the other, and a census is the alternative in which you measure the whole population instead.

quizWhat is a sampling frame?expand_more

A sampling frame is the actual list or mechanism you draw your sample from: a student register, a patient register, a customer database, a household listing, a voter roll. It matters more than the method name, because anyone missing from the frame has a zero chance of being selected no matter how good your sampling technique is.

quizWhich sampling method does not require a sampling frame?expand_more

Convenience, quota, purposive and snowball sampling need no list at all. Cluster and multistage sampling need only a list of groups such as schools, clinics or villages, not a list of individuals. Simple random, systematic and stratified sampling all require a full list of individuals before you can begin.

quizWhat is the difference between probability and non-probability sampling?expand_more

In probability sampling a random mechanism does the selecting, so every unit has a known, non-zero chance of being chosen, and you can calculate how uncertain your estimates are. In non-probability sampling those chances are unknown because selection depends on convenience, quotas or the researcher's judgement. The practical consequence is that only a probability design supports a stated precision figure for a population estimate.

quizWhat is the difference between stratified and cluster sampling?expand_more

Stratified sampling divides the population into groups and takes a random sample from every group, which improves precision and guarantees each group appears. Cluster sampling selects some groups at random and then surveys everyone inside the selected groups, which saves cost when no list of individuals exists. The shorthand is: stratified uses all the groups and some of the people, cluster uses some of the groups and all of the people.

quizWhat is a primary sampling unit?expand_more

The primary sampling unit is whatever gets selected at the first stage of a multistage design. In a national household survey it is usually an area such as a ward, village or census block. Households are then selected inside it at the second stage. Naming your primary sampling unit is a standard requirement when reporting a clustered design, because the clustering it creates has to be reflected in the analysis.

quizWhat is stratified random sampling?expand_more

Stratified random sampling splits the population into non-overlapping groups called strata, then draws a random sample independently within each stratum. Proportional allocation gives each stratum a share of the sample matching its share of the population and needs no weighting. Disproportionate allocation deliberately oversamples small but important strata, and then requires weighting before any overall figure is reported.

quizIs snowball sampling valid for academic research?expand_more

Yes, when the population is hidden, rare or unregistered and no sampling frame can be built. It is a recognised approach for hard-to-reach groups. What it cannot do is support a population estimate, and its main weakness is that participants refer people similar to themselves, so the sample drifts toward one social network. Starting from several unconnected seeds reduces that drift and should be described in your methodology.

quizIs random sampling the same as random assignment?expand_more

No. Random sampling is how you select participants from a population, and it governs who your findings generalise to. Random assignment is how you allocate selected participants to conditions in an experiment, and it governs whether an observed effect can be attributed to the treatment. A study can use one without the other, and most experiments use random assignment without random sampling.

quizCan I use a non-probability sampling method in a quantitative study?expand_more

Yes, and a great deal of published quantitative research does. What changes is what you may claim. You can report counts, percentages and relationships within your sample, and you can compare groups inside it. You cannot present those percentages as estimates of the population with a stated precision. State the method, describe the recruitment, compare your respondents against a known benchmark where one exists, and write the limitation explicitly rather than hoping nobody asks.

quizWhat is the difference between sampling and a sampling distribution?expand_more

Sampling is the procedure for selecting a sample. A sampling distribution is a different idea. Imagine repeating the same sampling procedure many times. The spread of the resulting statistic, such as the mean or the percentage, across all those imagined samples is the sampling distribution. The sampling distribution is what makes confidence intervals possible, and it only behaves predictably when the sample was drawn by a probability method.

quizHow do I know if my sample is biased?expand_more

Compare your achieved sample with a known benchmark such as your roster, enrolment records or administrative data, on variables that relate to what you are measuring. Large gaps point to coverage problems, selection problems or differential nonresponse. Also compare across recruitment channels and reminder waves. Be aware that a healthy response rate is weak evidence of low bias: across 235 published estimates, response rate explained only about 11 per cent of the variation in nonresponse bias.

Related reading

Sampling is one decision inside a larger study. These guides cover the decisions either side of it.

How this guide was written, and what it does not cover

Who wrote it. Michael Hodge, who works on survey methodology, question design and data quality at SuperSurvey, and writes the practical guides in our Survey Learning Center.

What it is based on. Definitions follow the methodological review and the university research guidance cited below, which is where the technical terms are actually defined. The official research-integrity and statistical-office pages are listed as short plain-language starting points, not as the source of those definitions. Method descriptions follow the same reviews and guidance. Every number in the evidence section comes from a named primary source listed below. Each claim also carries a confidence grade, so you can see which findings are settled and which rest on a single study. The design effect figures in the chart and table are our own arithmetic. We used the standard formula, with the assumptions printed in the caption, so you can redo them or argue with them using your own cluster sizes.

What it does not cover. This guide is about sampling for surveys and observational studies. It does not cover experimental design or random assignment beyond telling them apart. It also leaves out sampling for acceptance testing, auditing and signal processing, which use the same word for different things. Respondent-driven sampling, time-location sampling and adaptive cluster sampling are real methods that specialists use for hidden populations, and they are out of scope here. The intraclass correlation figures come from household surveys in low and middle income countries. Treat them as illustrative, not as values to import into a different setting.

Where the numbers are weakest. The Literary Digest reanalysis rests on one 1937 survey and one later reanalysis, and both authors flag their own uncertainty; it is graded B for that reason. Polling error figures come from United States election polls, which are unusually well studied and not necessarily representative of survey error in general.

Disclosure. SuperSurvey sells survey software. Nothing on this page is for sale, there is no sign-up gate on any part of it, and no organisation sponsored or reviewed this guide. Where we link to our own pages it is because they cover the next decision, not because a link earns us anything.

References

  1. Office of Research Integrity, U.S. Department of Health and Human Services. Elements of Research: Sampling.
  2. U.S. Census Bureau. Sampling Estimation and Survey Inference.
  3. Elfil, M., and Negida, A. (2017). Sampling Methods in Clinical Research: an Educational Review. Emergency (Tehran), 5(1), e52. PMC5325924.
  4. Spolarich, A. E. (2023). Sampling Methods: A Guide for Researchers. Journal of Dental Hygiene, 97(4), 73-77.
  5. University of Bath. Sampling in research.
  6. UK Government (2020). A guide to inclusive social research practices.
  7. Squire, P. (1988). Why the 1936 Literary Digest Poll Failed. Public Opinion Quarterly, 52(1), 125-133. DOI 10.1086/269085.
  8. Lohr, S. L., and Brick, J. M. (2017). Roosevelt Predicted to Win: Revisiting the 1936 Literary Digest Poll. Statistics, Politics and Policy, 8(1), 65-84. DOI 10.1515/spp-2016-0006.
  9. Kennedy, C., and Hartig, H. (2019). Response rates in telephone surveys have resumed their decline. Pew Research Center.
  10. International Institute for Population Sciences and ICF (2022). National Family Health Survey (NFHS-5), 2019-21: India, Volume I.
  11. Groves, R. M. (2006). Nonresponse Rates and Nonresponse Bias in Household Surveys. Public Opinion Quarterly, 70(5), 646-675. DOI 10.1093/poq/nfl033.
  12. Groves, R. M., and Peytcheva, E. (2008). The Impact of Nonresponse Rates on Nonresponse Bias: A Meta-Analysis. Public Opinion Quarterly, 72(2), 167-189. DOI 10.1093/poq/nfn011.
  13. Bradley, V. C., Kuriwaki, S., Isakov, M., Sejdinovic, D., Meng, X.-L., and Flaxman, S. (2021). Unrepresentative big surveys significantly overestimated US vaccine uptake. Nature, 600, 695-700. DOI 10.1038/s41586-021-04198-4.
  14. Mercer, A., and Lau, A. (2023). Comparing Two Types of Online Survey Samples. Pew Research Center.
  15. Mercer, A., Kennedy, C., and Keeter, S. (2024). Online opt-in polls can produce misleading results, especially for young people and Hispanic adults. Pew Research Center.
  16. GSMA (2025). The Mobile Gender Gap Report 2025, pages 9 and 15. Figures are from the 2025 edition; GSMA has since published a 2026 edition with a narrower gap.
  17. Brubaker, J., Kilic, T., and Wollburg, P. (2021). Representativeness of individual-level data in COVID-19 phone surveys: Findings from Sub-Saharan Africa. PLOS ONE, 16(11), e0258877. DOI 10.1371/journal.pone.0258877.
  18. Kennedy, C., Blumenthal, M., Clement, S., Clinton, J. D., and others (2018). An Evaluation of the 2016 Election Polls in the United States. Public Opinion Quarterly, 82(1), 1-33. DOI 10.1093/poq/nfx047.
  19. Mercer, A., Lau, A., and Kennedy, C. (2018). For Weighting Online Opt-In Samples, What Matters Most? Pew Research Center.
  20. AAPOR Task Force on 2020 Pre-Election Polling (2021). An Evaluation of the 2020 General Election Polls. American Association for Public Opinion Research.
  21. Aliaga, A., and Ren, R. (2006). Optimal Sample Sizes for Two-stage Cluster Sampling in Demographic and Health Surveys. DHS Working Paper No. 30.
  22. ICF International (2012). Demographic and Health Survey Sampling and Household Listing Manual, page 9 and Table 1.2.
  23. Kalton, G., Brick, J. M., and Le, T. (2005). Estimating components of design effects for use in sample design. In Household Sample Surveys in Developing and Transition Countries, Chapter VI. United Nations, Series F No. 96.