Representative samples and sampling bias
A sample is representative when its selection gives a reasonable picture of the population relevant to the question.
More key points
- Sampling bias occurs when the method systematically overrepresents some groups or misses others, making a large sample misleading despite its size.
On this page9 sections
A survey can collect thousands of answers and still tell you very little about the population it claims to describe. The key question is not only how many people responded, but who had a chance to be selected and who actually responded.
Define the population first
The population is the full group the investigator wants to describe. A sample is the smaller group observed. Before judging a result, identify the population in the claim: all students at a school, all registered voters in a county, or only people who used a particular service. A survey of one group cannot automatically support a conclusion about a broader group.
How samples become biased
- Convenience sampling: asking people who are easiest to reach, such as students in one class, while missing others.
- Voluntary-response bias: inviting open responses and hearing disproportionately from people with strong opinions or extra time.
- Undercoverage: leaving part of the target population out of the sampling frame, such as an online-only survey that excludes people without reliable internet.
- Nonresponse bias: selected people who do not answer differ in a relevant way from those who do.
- Leading questions: wording pushes respondents toward a particular answer, distorting the measured response even if the selection was reasonable.
A simple random sample can reduce selection bias because each member of the defined population has an equal chance of selection. Stratified methods can help ensure important subgroups are included. Neither method fixes poor question wording, inaccurate measurement, or nonresponse by itself.
Sample size is not representativeness
Increasing the number of observations can reduce random variation when the sampling process is sound. It does not erase systematic bias. If a poll asks only people leaving a basketball game which sport the city prefers, interviewing more fans makes the sample larger but does not make it representative of the city.
Evaluate the conclusion in four steps
- Name the target population and compare it with the people who could be reached.
- Check how participants were selected and whether participation was voluntary.
- Look for missing groups, nonresponse, and wording that could shape answers.
- Limit the conclusion to what the sample and method can support; avoid claiming causation from a survey alone.
Suppose a school wants to estimate support for later start times but polls only students who ride one bus route. That group may differ from students who walk, drive, or use other routes. The result can describe respondents, but it may not represent the full student body unless the school addresses that coverage problem.
Match the sample to the population
A sample is representative when it reflects important characteristics of the population the study aims to describe. If a survey is meant to estimate opinions among all students, sampling only students who visit the library may overrepresent frequent library users. The problem is not merely sample size; it is that the selection process can systematically exclude or overrepresent certain groups.
Random sampling gives members of the target population a known chance of selection and can reduce selection bias. It does not guarantee a perfectly balanced sample, especially when the sample is small, and it is different from random assignment in an experiment. Sampling determines who is observed; assignment determines which condition participants receive. Both affect study quality, but they answer different questions.
Recognize common sources of bias
Voluntary-response bias occurs when people choose whether to participate; people with strong views may be more likely to respond. Undercoverage occurs when part of the population has little or no chance to be selected, such as a phone survey that reaches only households with one type of service. Nonresponse bias arises when selected individuals do not answer and their views differ from respondents. These problems limit generalization even if many responses are collected.
Convenience samples select whoever is easy to reach. They may be useful for a pilot or initial exploration but usually cannot support broad population estimates without additional evidence. A large convenience sample can produce a precise estimate for the people who responded while remaining biased for the broader population. Precision and representativeness are not the same.
Read the sampling method and uncertainty
When evaluating a survey, identify the target population, sampling frame, recruitment method, response rate, and timing. Ask who was excluded and whether participation could relate to the outcome. A poll conducted immediately after a controversial event may reflect short-term reactions. A survey sent to current customers cannot estimate the views of people who have never used the service.
A margin of error describes sampling variability under specified assumptions; it does not fix systematic bias from a flawed frame or nonresponse. A narrow reported margin of error should not be treated as proof that a biased sample represents everyone. Look for the method and the population to which the estimate applies.
- Name the target population and compare it with the people actually reached.
- Distinguish random sampling from random assignment.
- Look for undercoverage, voluntary response, convenience selection, and nonresponse.
- Do not confuse a large sample or narrow margin of error with lack of bias.
- Limit conclusions to the population and time period the method supports.
Use sampling methods to reduce selection problems
A simple random sample gives each member of the target population a known chance of selection. A stratified sample first divides the population into relevant groups, then samples within them to ensure each group is represented. Neither method removes every source of error, but both can reduce undercoverage compared with asking only whoever is nearby.
A representative sample reflects the population features that matter for the question; it need not reproduce every demographic proportion exactly. For a small sample, chance imbalance can still occur. Report uncertainty and avoid claiming that a sample is representative solely because it was selected randomly.
Exam takeaway
When a question asks whether evidence supports a broad claim, scrutinize the sampling method before calculating anything. A biased sample weakens generalization. A careful answer states both the limitation and the group to which the evidence most directly applies.
Common questions
Does a larger sample always reduce bias?
No. A larger sample may reduce random error, but it does not fix systematic selection, coverage, or nonresponse bias.
What makes a sample representative?
Its selection method and achieved participants reasonably reflect the target population for the characteristics relevant to the question.
Is a voluntary online poll a random sample?
Usually not. People choose whether to respond, so the result may overrepresent participants with strong views, time, or internet access.