Significant relationship in research: what it means and how to test it
Two comparisons worked through with every step of the arithmetic, and the sentence to write when you report the result.
A significant relationship means the pattern you measured would rarely turn up if the two variables were unrelated in the population you sampled. It is a statement about your data under an assumption. It is never a statement about how probable the assumption is.
Across the 109 templates in our own survey library, longer templates give open text a smaller share of their questions: r = -0.37. A correlation at least that far from zero would appear about once in 13,000 samples if the real correlation were zero, so zero is a poor explanation of what we found.
SuperSurvey Question Corpus, 109 live templates and 1,422 questions, read 11 September 2026.
Comparing two percentages? Skip to the two-proportion calculator.
What a significant relationship in research means
The direction of the conditional is the whole thing, and it is the part most write-ups get backwards.
Start from the claim you are arguing against. It is always the dull one: in the population, these two variables have nothing to do with each other. Now measure the relationship in your sample and ask a single question. If the dull claim were true, how often would a sample this size throw up a relationship at least this strong, by luck alone? That number is the p-value. Small means luck is a poor explanation.
What you must not do is turn the sentence round. The test never tells you the probability that the variables are related, because it computed everything on the assumption that they are not. The live version of this page used to say a significant relationship means the correlation "is unlikely to be zero in the population". That reads well and it is wrong. The correlation you observed is unlikely if the population correlation is zero. Same words, opposite direction, and only one of them is a thing your data can support.
"Template length and open-text share were negatively correlated, r = -0.37, n = 109, p = 7.5 x 10^-5. Longer templates devoted a smaller share of their questions to open text."
Estimate, sample size, p-value, then one plain sentence saying what moved with what. No verdict word, no causal verb, no adjective doing work the number does not support.
Three things the phrase does not carry, however often it is read that way:
- Not strength. Significance answers whether the relationship is distinguishable from nothing. How big it is comes from r itself, or from a slope. With enough responses, an r of 0.05 is significant and still worth almost nothing.
- Not cause. Two variables moving together says nothing about which one moves the other, or whether a third is moving both. That is a design question, settled by how the study was built, and no test statistic can answer it afterwards.
- Not truth. A p-value is computed inside a model that assumes your sample is what it claims to be. If the wrong people answered, the arithmetic is still perfect and the conclusion is still wrong.
Testing a significant correlation between two variables
Pearson's r goes from -1 to +1 and measures how closely two variables track a straight line. To ask whether the r you got is more than luck, convert it to a t-statistic:
t = r x sqrt(n - 2) / sqrt(1 - r^2), read against the t-distribution on n - 2 degrees of freedom.
Here is that on real data. Before looking at anything, I fixed one question: across the 109 live templates in our library, is the number of questions in a template related to the share of those questions that are open text? Then I ran it once.
-
Get r
r = -0.3701 over n = 109 templates. Rounded for reporting, r = -0.37.
-
Get the degrees of freedom
n - 2 = 107. So sqrt(107) = 10.3441.
-
Get the denominator
r squared is 0.1369, so 1 - r^2 = 0.8631 and sqrt(0.8631) = 0.9290.
-
Divide
t = -0.37 x 10.3441 / 0.9290 = -4.1197, on 107 degrees of freedom.
-
Read the tail
A t of 4.12 on 107 degrees of freedom puts 0.000075 of the two-tailed distribution beyond it. That is p = 7.5 x 10^-5, or about one sample in 13,000.
Report the interval with it. On Fisher's transformation the 95 per cent interval for r runs from -0.52 to -0.20. That is the honest summary: a negative relationship, somewhere between weak and moderate, and the data cannot pin it down more tightly than that. What such a coefficient can and cannot support, and how to design a study around one, is covered where we interpret a survey correlation in full.
Swap the outcome from the share of questions that are open text to the number of them, and the answer reverses. r = 0.07, t = 0.77 on 107 degrees of freedom, p = 0.44. Nothing significant at all.
Both results are true. Templates of 4 to 9 questions average 2.1 open-text items and templates of 16 to 36 average 2.3, so the count barely moves while the share falls mechanically as the denominator grows. Which is why you write down the variable before you run the test. Had I tried both and reported the exciting one, I would have manufactured a finding out of a definition.
The share result survives a check for one template dragging it: on ranks rather than values, Spearman's rho is -0.44 with p = 1.6 x 10^-6, and the count result stays flat at rho = -0.03. Pearson's r asks for a straight-line relationship, independent pairs, and no single point far enough out to bend the line on its own. Check the scatter before you trust the coefficient. If you want the size of the effect rather than a yes or no, fit a slope and estimate the relationship in real units: so many open-text questions fewer per extra question, rather than a bare r.
Significant difference vs significant relationship
These are the two shapes almost every survey question comes in, and they are tested differently. A difference compares groups on one outcome. A relationship asks whether two measurements move together.
| Significant difference | Significant relationship | |
|---|---|---|
| The question | Do these groups differ on this outcome? | Do these two measurements move together? |
| Your data | One grouping variable, one outcome | Two measurements on the same people |
| The dull claim being tested | The group means or shares are equal in the population | The population correlation is zero |
| What p is the probability of | A gap at least this large, if the groups really are equal | A correlation at least this far from zero, if it really is zero |
| Usual test | Two-proportion z-test, t-test, ANOVA, chi-square | Pearson r, Spearman rho, a regression slope |
| Effect size to publish | The gap in points, with its interval | r, with its interval |
| Survey example | Email invites got 52.0% top-box, in-app got 45.0% | Template length tracks open-text share at r = -0.37 |
Both are hypothesis tests and both produce a p-value, so both are read the same way. They are not interchangeable: a gap between groups is not a correlation, and reporting one as the other is a common way to overstate a finding. Either can be run on quantitative survey data.
Is the difference between two groups significant?
Say a support team sends the same one-question survey after every ticket, through two kinds of invite. Of the 400 people who answered after an email invite, 208 picked the top box. Of the 420 who answered after an in-app prompt, 189 did. That is 52.0 per cent against 45.0 per cent, a gap of 7.0 points. Is that more than sampling variation alone would produce at these sizes?
Because the outcome is a share rather than an average, the test is the two-proportion z-test. Pool the two groups to estimate the rate you would expect if the invite made no difference, work out how much a gap of this kind bounces around at these sample sizes, and divide.
-
Pool the two groups
(208 + 189) / (400 + 420) = 397 / 820 = 0.4841.
-
Build the standard error
0.4841 x 0.5159 = 0.2497. 1/400 + 1/420 = 0.004881. Multiply: 0.00121901. Square root: 0.03491.
-
Divide the gap by it
z = 0.07 / 0.03491 = 2.0049. Keep the decimals. Divide by the rounded standard error of 0.0349 instead and z comes out at 2.0057, which prints as 2.01 and carries a tail of 0.0444. That is the 0.044 this page published for years, and it is wrong by a rounding step, not by a decimal place.
-
Read the tail
A standard normal puts 0.04497 of its mass beyond 2.0049 in the two tails combined. So p = 0.045. Below 0.05, so the gap is significant at the 5 per cent level, with very little room to spare.
The z-test uses a normal curve to approximate a count, and that approximation needs enough people in every corner. Multiply each group size by the pooled rate and by one minus the pooled rate: 193.7, 206.3, 203.3 and 216.7. All four need to be at least 5. Here the smallest is 193.7, so the approximation is safe. Below 5, use an exact test instead, because the p-value stops meaning what it says.
Run the same numbers as a chi-square test of independence on the two-by-two table and you get the identical answer: chi-square is z squared, 4.0197 on one degree of freedom, p = 0.04497. They are the same test written two ways, so a significant chi-square on a two-by-two table and a significant z-test are never in conflict.
The gap did not change when we changed the sample size. The verdict did. At a quarter of the responses, 52 of 100 against 45 of 100 for the same 7.0 points, z falls to 0.9904 and p rises to 0.32. Nothing about the invites changed. Only the evidence did, which is why you plan the response count before fielding rather than after.
The same seven-point gap, tested at four sample sizes
p-value, two-tailed. Anything below 0.05 is called significant at the 5 per cent level.
If the comparison is still ahead of you, this is the moment that decides it. Create a survey to compare groups and set the size of each group from the smallest gap you need to detect. A comparison that was too small when it ran stays too small; a larger follow-up is a new study with its own p-value, not a repair of this one.
Two-proportion significance calculator
Your own numbers, same arithmetic. Enter how many people answered in each group and how many chose the answer you are tracking. It updates as you type.
Compare two percentages
Two methods, both named in the output: a two-proportion z-test on the pooled standard error for the verdict, and a Wald interval on the unpooled standard error for the size. That is the standard pairing, and because the two errors differ they can disagree close to the threshold; the calculator says so whenever it happens. The interval level follows the significance level you choose, so alpha = 0.05 gives a 95 per cent interval and the label states the level every time.
It reports the smallest expected count so you can see when the normal approximation stops being safe. It does not report observed power: post-hoc power is a fixed function of the p-value, so it adds nothing to what you can already see.
A confidence interval for the difference between two groups
The p-value answers a yes-or-no question nobody actually asked. The interval answers the one they did: how big is it?
A confidence interval is a range built so that, if you repeated the whole exercise many times, 95 per cent of the intervals you built that way would contain the true value. It is a property of the method, not a probability attached to the one interval in front of you. Loose talk about "a 95 per cent chance the truth is in here" is the same inversion that gets p-values wrong.
For a difference between two proportions, the standard error is built from each group separately rather than pooled, because you are now estimating the gap rather than testing whether it is zero:
-
Standard error of the difference
0.52 x 0.48 / 400 = 0.00062400. 0.45 x 0.55 / 420 = 0.00058929. Add them and take the square root: 0.03483, or 3.483 percentage points.
-
Multiply by the critical value
For 95 per cent, that is 1.9600, not 2. Rounding it up to 2 drags the lower bound from 0.2 down to 0.0, which is the difference between an interval that clears zero and one that appears to touch it. 1.9600 x 3.483 = 6.827 points.
-
Put it either side of the estimate
7.0 plus or minus 6.827 gives 0.2 to 13.8 percentage points.
Now read what that actually says. The email invite is ahead, and the data are consistent with a lead of anything from a fifth of a point to nearly fourteen. A finding of "significant at 0.05" hid all of that. If a two-point gap would change what the team does, this study has not answered the question, and the p-value gave no hint of it.
The two readings usually agree, and it is worth being exact about why they need not. The z-test above divided by the pooled standard error, 0.03491, because it was computed under the hypothesis that both groups share one rate. This interval used the unpooled one, 0.03483, because it estimates the gap without that assumption. A test and an interval agree exactly only when they are built from the same construction: an interval made by inverting the pooled test, or a test run on the unpooled error. This standard pairing is two approximations to that, and close to the threshold they can split.
Take 5 of 20 against 11 of 20: the pooled test gives z = -1.9365 and p = 0.0528, not significant at 0.05, while the unpooled 95 per cent interval runs from -58.9 to -1.1 points and excludes zero. Both are correct for their own method; the result is borderline and should be reported as both numbers. The interval is still the better thing to publish, because it carries the size as well as the verdict, and the verdict alone carries neither.
The estimate never moves. The interval around it does
Percentage points. Interval from the unpooled standard error of the difference.
What the p-value says about the null hypothesis
The dull claim has a name: the null hypothesis, H0. The p-value is the probability of getting a test statistic at least as extreme as yours when H0 is true. Everything else people say about it is either a shorthand or a mistake.
The mistake is always the same one, and it is worth naming precisely. P is the probability of the data given the hypothesis. What people read off it is the probability of the hypothesis given the data. Those are different quantities, and the second one cannot be computed from a p-value at all: it needs to know how plausible the hypothesis was before you collected anything. The American Statistical Association made this the second of its six principles in 2016: a p-value measures neither the probability that the hypothesis under study is true, nor the probability that the data were produced by chance alone. 4
So p = 0.045 in the invite example does not mean a 4.5 per cent chance the invites are equivalent, and it does not mean a 95.5 per cent chance the result will replicate. Replication depends on the true effect and the next study's size, neither of which is in the number.5
Alpha and the significance level
Alpha is the threshold you commit to before the data arrive. Set alpha = 0.05 and you have said: if the null is true, I accept a 5 per cent chance of calling a result significant anyway. Run twenty such tests on pure noise and about one comes back significant, by construction rather than by accident.
Nothing makes 0.05 correct. It is a convention that hardened after Fisher used it as a rough guide, and the trade-off is plain: a lower alpha buys fewer false alarms and costs you real effects you will now miss. Pick it from the cost of being wrong in each direction, write it down first, and do not move it once you have seen the p-value.
One-tailed and two-tailed tests
A two-tailed test asks whether the groups differ at all, counting extreme results on both sides. A one-tailed test asks whether one specific group is higher, and counts only that side.
The usual claim is that a one-tailed p-value is half the two-tailed one. That holds only when the effect actually runs the way you predicted. Our example had a two-tailed p of 0.04497. If you had predicted beforehand that email would beat in-app, which is the direction it went, the one-tailed p is 0.02249. If you had predicted in-app would win, the one-tailed p is not 0.02249. It is 1 minus that, 0.9775, because almost the entire distribution sits on the side you claimed. Halving is a special case, not a rule.
Choose one-tailed only when a difference in the other direction would genuinely leave your decision unchanged, and only before you look. Switching to one-tailed after seeing a two-tailed p of 0.08 is the oldest trick there is, and it doubles your false-positive rate in the direction you happened to like.
Which statistical test should I use for two groups?
Pick by the shape of the outcome, then check the conditions before you believe the p-value. The conditions are not paperwork: every one of them is an assumption the arithmetic quietly makes.
| Your outcome | Comparison | Test | Conditions |
|---|---|---|---|
| A share, such as top-box per cent | Two groups | Two-proportion z-test | Independent responses. Every expected count at least 5: n x p-pooled and n x (1 - p-pooled), in both groups |
| A share | Two groups, small counts | Fisher exact test | Independent responses. Use it whenever the z-test's expected counts fall below 5 |
| An average | Two groups | Two-sample t-test, Welch version | Independent responses, no extreme skew, roughly 30 or more per group. Welch does not assume equal spread, so prefer it |
| An average | Same people twice | Paired t-test | Pairs independent of each other, differences roughly symmetric. Pairing is the point: do not run the unpaired test |
| An average | Three or more groups | ANOVA, then a corrected follow-up | Independent responses, similar spread across groups, no extreme skew. Correct the follow-up comparisons |
| Two categorical answers | Association | Chi-square test of independence | Independent responses. Expected count of at least 5 in most cells, and no cell expected below one6 |
| Two measurements | Relationship | Pearson r with a t-test | Independent pairs, a straight-line relationship, no single point bending the line |
| Two measurements, ordinal or skewed | Relationship | Spearman rho | Independent pairs. Ranks the values first, so it needs only a consistent direction |
One condition sits above all of these and no test checks it: the responses have to come from the people you meant to ask. A tiny p-value on a skewed sample is a precise answer to the wrong question, which is the practical cost of response bias.
Statistical significance does not imply practical importance
Significance tells you an effect is detectable at your sample size. It says nothing about whether the effect is worth acting on, and with enough responses the two come apart completely.
Take a one-point move: top-box goes from 51.0 per cent to 52.0 per cent between two waves. At 400 responses per wave, p = 0.777 and nobody calls it anything. The move becomes significant at 0.05 once each wave carries about 19,191 responses: at exactly that size p = 0.0500. The point has not become more important. You have just bought enough precision to see it.
So publish the size alongside the verdict. For the invite comparison that is 7.0 percentage points, interval 0.2 to 13.8. If you want a standardised figure, Cohen's h for 52.0 per cent against 45.0 per cent is 0.14, which sits below the 0.2 that Cohen called small.7 A detectable difference, and a small one.
Before fielding, write down the smallest difference that would change what you do. Three points? Five? Then the analysis has something to answer. Without it, every significant result looks like a reason to act and every non-significant one looks like a reason to wait, which is not a decision process.
Two surveys six months apart gave different results
This is the most common way the question reaches a survey team, and reaching for a significance test first is the wrong move. A test assumes the two samples differ only by chance. Between two waves they usually differ by rather more than that.
Check these before you compute anything. Was the question worded identically, with the same scale labels in the same order? Did both waves run for the same length of time, on the same days? Did the same kinds of people answer, or did one wave go out through a channel that reaches a different crowd? Was one wave in a period the business itself disturbed? Any of these produces a gap that no test will flag as chance and that has nothing to do with the thing you are measuring.
Once the comparison is genuinely like for like, the test follows from what you are comparing, and for survey work that is usually a share. Of the 1,422 questions in our template library, 45.3 per cent are rating items and 98.9 per cent of those offer five points. The natural summary of a five-point item is the share picking the top box or the top two, so the two-proportion z-test above is the workhorse, not the t-test. Averaging a five-point scale treats the step from 4 to 5 as equal to the step from 1 to 2, which nobody has established.
Source: SuperSurvey Question Corpus, 109 live templates and 1,422 questions, read 11 September 2026.
Two more habits worth keeping. Decide which comparisons you will run before the data land, because the false-positive rate climbs with the number of tests, not with the number you report. And take the sample seriously: who answered matters more than how many, which is what sampling methods are for. The rest of the survey learning centre covers the decisions either side of this one.
Frequently asked questions
A significant chi-square value means the difference is probably not due to chance. True or false?
False as written, though most courses mark it true. The chi-square statistic measures how far the counts you saw sit from the counts you would expect under independence, and its p-value is how often a gap that large arises when the variables are independent. "Probably not due to chance" attaches a probability to the explanation rather than to the data, which is the one move the test cannot make. If the marking scheme wants true, it is treating the sentence as loose shorthand for the correct version.
Is a result significant if the probability it was due to chance is below 5 per cent?
If p is 0.06, is there a trend towards significance?
No. There is no trend and no partial credit. A threshold you set in advance was not met, and 0.06 is no more a trend than 0.04 is a proof. What you can honestly say is that the interval still includes zero, and then give the interval so a reader can see how much room is left on either side.
How many comparisons before an alpha of 0.05 stops protecting me?
Faster than people expect. Twelve independent tests on data with nothing in it give a 46 per cent chance of at least one hit, from 1 minus 0.95 to the twelfth. At twenty tests it is 64 per cent. Crosstabbing one outcome by six segments is already fifteen pairwise comparisons. Name the handful you planned, and mark everything else as exploratory rather than quietly reporting the survivors.
Can I run a significance test on a five-point rating question?
Yes, if you pick a summary the scale can support. Comparing the share who picked the top one or two boxes uses only the order of the options, and that goes straight into a two-proportion test. Comparing mean scores assumes the gaps between points are equal, which is an extra assumption most teams make without noticing. If you want the whole distribution rather than a top-box cut, a Mann-Whitney test on ranks avoids the assumption entirely.
Does a bigger sample make my result more likely to be correct?
More precise, not more correct. Extra responses shrink random error, so the interval narrows around whatever your sample is measuring. If that sample leans a particular way, more of it just tightens the estimate around the wrong number, and the p-value gets smaller as it does so. Volume shrinks random error. It does nothing at all about who chose to answer.
A report gives t = 2.000 with df = 20 and p greater than 0.05. Should the null hypothesis be rejected?
No, by the conventional standard. The exact two-tailed value for t = 2.000 on 20 degrees of freedom is 0.0593, just above the threshold, and 2.086 is what you would have needed. Degrees of freedom matter to whoever reads the study because they set that bar: the same t of 2.000 gives p = 0.0734 on 10 degrees of freedom and p = 0.0482 on 100. Only the last of those clears 0.05, and the statistic never moved.
Method and limits
Every figure on this page is computed in the file that builds it, not typed in, and the build refuses to run if any of them changes. Tail probabilities come from the incomplete gamma and incomplete beta functions at double precision, checked against closed forms where one exists: the Cauchy form for a t-distribution on one degree of freedom, the algebraic forms on two and four, and exp(-x/2) for chi-square on two. Nothing is read off a table.
The correlation uses the live SuperSurvey template library, parsed template by template on 10 September 2026: 109 live templates and 1,422 questions. Counted was the number of questions in each template and how many of them take a free-text answer. Excluded were drafts and redirects. Both versions of the outcome were specified before the test ran and both are reported above.
Two limits matter. This is our own published library, so it shows what one survey vendor considers good practice, not what the world's surveys do, so no result here should be read as a fact about surveys in general. And the share result is partly structural: template length is in the denominator of the outcome, so some of the negative correlation is arithmetic rather than editorial judgement. That is exactly why the count version is printed next to it.
The invite comparison, the sample-size ladder and the one-point move are worked examples with stated inputs, not measurements. The arithmetic is real and reproducible from the numbers shown; the scenarios are illustrations.
References
- Centers for Disease Control and Prevention, National Center for Health Statistics. Statistical significance. In Health, United States: Sources and Definitions. cdc.gov
- National Institutes of Health, National Center for Advancing Translational Sciences. Statistical significance. NCATS Toolkit glossary. toolkit.ncats.nih.gov
- US Department of Education, National Center for Education Statistics. Statistical significance and sample size. National Assessment of Educational Progress. nces.ed.gov
- Wasserstein, R. L. and Lazar, N. A. (2016). The ASA statement on p-values: context, process, and purpose. The American Statistician, 70(2), 129-133. doi.org/10.1080/00031305.2016.1154108
- Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N. and Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. doi.org/10.1007/s10654-016-0149-3
- Cochran, W. G. (1954). Some methods for strengthening the common chi-square tests. Biometrics, 10(4), 417. doi.org/10.2307/3001616
- Cohen, J. (1992). A power primer. Psychological Bulletin, 112(1), 155-159. doi.org/10.1037/0033-2909.112.1.155
- Hoenig, J. M. and Heisey, D. M. (2001). The abuse of power: the pervasive fallacy of power calculations for data analysis. The American Statistician, 55(1), 19-24. doi.org/10.1198/000313001300339897
- Cox, D. R. (2020). Statistical significance. Annual Review of Statistics and Its Application, 7, 1-10. doi.org/10.1146/annurev-statistics-031219-041051
- SuperSurvey Question Corpus, 109 live templates and 1,422 questions, read 11 September 2026.
Michael Hodge: Survey methodologist and editor, SuperSurvey. Bachelor of Science (Psychology), University of Wollongong, with coursework in psychometrics and research methods. Designing surveys since 2003. About the author and how these guides are reviewed
What to read next
Two comparable groups, the same question, enough responses to tell them apart. That is a survey you can run this afternoon.
Start from a survey templateFree to send, no card. Change the wording, keep the question identical for both groups.