Correlation survey: what an r of 0.34 between two questions means
You crossed two questions from the same survey and got a number. Here is how to design that survey properly, how to read the number, how much of it to believe, and what to write down.
A correlation survey is a nonexperimental study: you measure two variables on the same people, assign nothing and manipulate nothing, and ask how far the two sets of answers move together. The usual measure is Pearson's r, which runs from -1 to +1 and describes linear association only: how closely the paired answers sit to one straight line. At +1 or -1 every pair sits exactly on that line. At 0 there is no straight-line trend, which is not the same as no relationship, because a curved one also scores near zero.
An r of 0.34 means the two answers lean the same way in most people and disagree in plenty of others. Squared, it accounts for about 12 per cent of the movement in the second answer. It is a lean, not a rule.
What a correlation survey is
It is an ordinary survey read a particular way. One row per respondent, one column per question, and both answers in a row belong to the same person. That last part is the whole design. If the role-clarity answers came from one export and the intent-to-stay answers from another, and nothing joins them at the respondent level, there is no correlation to compute.
Nothing is manipulated. You did not assign half the company to a clear role. You measured what was already there and asked how the two columns line up. The formal definition in the standard open textbook is a study "in which the researcher measures two variables and assesses the statistical relationship (i.e., the correlation) between them with little or no effort to control extraneous variables".1
Pearson's r then summarises one specific thing about the two columns: how close the points sit to a single straight line, and whether that line rises or falls. It says nothing about which answer came first, nothing about how many points of one answer a point of the other is worth, and nothing about any relationship that is not a straight line. That last gap is easy to miss, so here is the smallest example that shows it.
| X | -2 | -1 | 0 | 1 | 2 |
|---|---|---|---|---|---|
| Y = X squared | 4 | 1 | 0 | 1 | 4 |
Y is completely determined by X in that table and r still comes out at zero, because the left half of the curve falls while the right half rises and the two cancel. A real survey rarely produces anything that tidy, but a milder version is common: satisfaction that rises with tenure for the first few years and then falls. A scatter plot catches it in seconds. r never will.
r is not a percentage. An r of 0.34 is not 34 per cent of anything. The 12 per cent is r squared, 0.34 times 0.34, and it is the share of the variation the two answers hold in common. The other 88 per cent of what moves intent to stay sits outside these two columns entirely.
Designing a correlation survey: three research questions
Decide these three things before a single question is written: who is in, which two measures are paired, and when each is taken.
A usable correlation question names a population, two variables with the instrument behind each, and a measurement timing, and it uses a verb that claims no direction. If the draft says "affects", "improves" or "drives", it is an experimental question and a survey will not answer it. Three that pass:
-
Employee, cross-sectional
Among all permanent staff at one company, is role clarity related to intent to stay?
Population: every permanent employee, invited by email. Paired variables: role clarity as the average of four 1-5 agreement items; intent to stay as one 0-10 item. Timing: both blocks in the same questionnaire, separated by unrelated questions, during one two-week fielding window. This is the example the rest of this page works through. The role-clarity block can be lifted from the employee engagement template, which is an editable starting point and not a validated correlation instrument, so treat its items as a draft to cut down.
-
Customer, two time points
Among customers who bought in the last quarter, is satisfaction at delivery related to repurchase intention three months later?
Population: every customer with a completed order in the quarter. Paired variables: a 1-5 satisfaction item asked in the post-delivery email; a 0-10 likelihood-to-buy-again item asked ninety days later. Timing: two surveys, joined on the order number, so the pairing survives even though the answers are months apart. Only customers who answer both count.
-
Training, survey plus a record
Among staff who completed the onboarding course this year, is self-rated confidence at course end related to the assessment score on file?
Population: everyone who finished the course. Paired variables: one 1-5 confidence item from the end-of-course survey; the assessment score pulled from the learning system, not asked. Timing: the survey on the last day of the course, the score whenever it was recorded, joined on staff ID. Taking one measure from a record instead of self-report is the cheapest protection against two answers sharing a mood.
Whatever the timing, the requirement is the same: paired observations linked to the same unit, whether that unit is a person, an order or a staff ID. Same day or ninety days apart is a design choice, not a rule. The join column is what makes the pair real, and reverse-coded items left uncorrected will flip the sign of a real relationship, so prepare paired responses before the coefficient is computed.
What counts as a strong survey correlation
The 0.10 / 0.30 / 0.50 labels most people quote come from 1988. The field has since counted what it actually publishes.
Cohen proposed 0.10, 0.30 and 0.50 for small, medium and large in 1988. Bosco and colleagues went and checked: they pulled 147,328 correlations out of the Journal of Applied Psychology and Personnel Psychology across thirty years and found the three cut points sit at roughly the 33rd, 73rd and 90th percentile of what actually gets published.2 A "medium" effect is in fact the top quarter.
The median published correlation is 0.16. Half of everything ever printed falls between 0.07 and 0.32.
The median correlation in applied psychology is 0.16, not 0.30
Bars are the 25th to 75th percentile. Counts: 147,328 correlations in total, 1,717 of them pairing an attitude with an intention, 295 pairing a job attitude with employees actually leaving.
Which row you compare yourself against matters more than the overall figure. Role clarity is an attitude and intent to stay is an intention, so the middle row is yours: median 0.27, middle half from 0.15 to 0.42. An r of 0.34 there is a little above typical and entirely unremarkable. It is not a discovery and it is not noise.
The bottom row is the one worth staring at. Job attitudes against employees actually leaving run at a median of 0.13. Intentions correlate with attitudes roughly twice as well as behaviour does, which is why a strong result on intent to stay is not a strong result on turnover. Gignac and Szodorai reach the same place from a different pile of papers: across 708 meta-analytically derived correlations the quartiles are 0.11, 0.19 and 0.29, and fewer than 3 per cent of published correlations are as large as 0.50.3
When to trust a survey correlation
The 0.34 is an estimate from the people who answered, not the number for the population. How far it could reasonably be off depends almost entirely on how many complete pairs went into it. Report the interval next to the coefficient and the argument mostly ends there.
At 60 replies, an r of 0.34 could really be 0.09 or 0.55
At 60 replies the interval runs from 0.09 to 0.55. That range covers everything from "barely there" to one of the strongest results in the applied literature, so the honest summary of a 0.34 on 60 people is that you do not yet know. At 412 it narrows to 0.25 to 0.42, which is tight enough to act on. Going from 412 to 1,000 buys you another 0.03 at each end and rarely earns its cost.
If you are planning rather than explaining afterwards, Bosco's figures give the target directly: about 105 responses for an eighty per cent chance of detecting a typical attitude-to-intention relationship, and about 462 for a typical job-attitude-to-turnover one.2 That gap is the same gap as in the chart above, seen from the other side. Working out how many responses you need before you field is the cheapest thing on this page.
A p-value answers a narrower question than most dashboards imply: whether a correlation this far from zero would be surprising if the true one were exactly zero. On 412 responses an r of 0.34 gives t = 7.32 and p below 0.001, which rules out zero and says nothing about whether 0.34 is worth acting on. That distinction is the whole of statistical significance.
Choosing the correlation for your survey questions
Pearson is the default in every survey tool, and it assumes both columns are numbers on a scale with even spacing. Survey answers frequently are not. The fix is to match the coefficient to what the two questions actually produce.
| Your two questions | Example pair | Use | Why |
|---|---|---|---|
| Two averaged multi-item scores | A four-item role-clarity average against a four-item manager-support average | Pearson r | Averaging several items produces something close enough to an even-spaced scale to treat as one |
| Two single rating items | One five-point agreement item against another | Spearman rho | Works on rank order only, so it never assumes the step from 4 to 5 equals the step from 1 to 2 |
| One yes/no, one scored | Left in the last year, yes or no, against an engagement score | Point-biserial | Pearson with one column coded 0 and 1; the arithmetic is identical and the reading is "average gap between the two groups" |
| Two yes/no | Completed the training against promoted this year | Phi | Correlation of two two-way splits, read straight off a two-by-two table |
The middle row is the common case and the one most reports get wrong. Of the 1,422 questions in our own published template library, 644 are rating items, 45.3 per cent, and 637 of those 644 use five points. A five-point item has five possible answers, so correlating two of them with Pearson is asking a formula built for continuous measurement to work on a five-rung ladder. Compute both, which costs one extra column, and where they disagree, Spearman is the one to trust. Whether you are correlating one item or a combined score changes the row you are in, so understand single items and combined scales before choosing.
Some questions cannot go in at all. Open text accounts for 221 of the 1,422 items in the library, 15.5 per cent, and until somebody reads those answers and codes them into categories or scores there is no column to correlate. Plan for that before you promise a cross-tab of everything.
A correlation study example, worked in Excel
Take a correlation of 0.34 on 412 complete responses. Both numbers are ordinary for this kind of pair. Everything below follows from those two by arithmetic, so your own r and your own N go in exactly the same places and the formulas do not change. The steps assume Pearson; Spearman gets its own block after them.
-
Get the two answers into two columns
One row per respondent, role clarity in column B, intent to stay in column C, header in row 1. Both cells in a row must come from the same person. If you are averaging a four-item block into one score, do that in its own column first, and check the four items are all pointing the same way before you average them.
=AVERAGE(D2:G2)
-
Count the pairs, not the responses
Excel silently drops a row when either cell is blank, so the N behind your r is often smaller than your response count. Get the real one, and report that number rather than the number of people who opened the survey.
=SUMPRODUCT((ISNUMBER(B2:B413))*(ISNUMBER(C2:C413)))
-
Compute Pearson's r
One function. It returns 0.34 for the worked example, and it ignores any row with a blank in either column, which is why the count in step 2 matters.
=CORREL(B2:B413,C2:C413)
-
Put an interval around it
Lower bound gives 0.25, upper bound gives 0.42. Swap in your own coefficient and your own pair count and it still works. This is the single most useful line in a correlation report and almost nobody includes it.
=FISHERINV(FISHER(0.34)-1.96/SQRT(412-3))
=FISHERINV(FISHER(0.34)+1.96/SQRT(412-3))
-
Get the p-value if someone will ask for it
t comes out at 7.32 on 410 degrees of freedom and p is below 0.001. The ABS() is not decoration: Excel's two-tailed T.DIST.2T rejects a negative t, so without it a negative correlation returns an error instead of a p-value. Wrapped, r = -0.34 gives exactly the same p as r = +0.34, which is what two-tailed means.
=0.34*SQRT(412-2)/SQRT(1-0.34^2)
=T.DIST.2T(ABS(0.34*SQRT(412-2)/SQRT(1-0.34^2)),412-2)
-
Plot it before you believe it
Insert a scatter chart of the two columns and look at it for ten seconds. You are checking three things: a curve, which Pearson reads as near-zero; a couple of far-out points doing all the work; and separate clumps, which usually means a department or a region is driving the whole result.
If your pair calls for Spearman
Excel has no Spearman function, so you rank both columns and run CORREL on the ranks. The trap is blanks. Rank column B on its own and column C on its own and a respondent who skipped C still gets a rank in B, so the two rank columns no longer describe the same people and the coefficient is wrong. Filter to complete pairs first, then rank. In Excel 365, one line pulls the complete rows into columns E to G:
=FILTER(A2:C13,(ISNUMBER(B2:B13))*(ISNUMBER(C2:C13)))
Rank the two filtered columns, then correlate the ranks. In older Excel, sort by each column and delete the rows with a blank in either before ranking; the principle is the same.
=RANK.AVG(F2,$F$2:$F$12,1)
=RANK.AVG(G2,$G$2:$G$12,1)
=CORREL(H2:H12,I2:I12)
The interval and p-value formulas in steps 4 and 5 are Pearson tools. For Spearman, report rho and the number of complete pairs; if a p-value is demanded, the same t formula with rho in place of r is the usual approximation once you have more than about thirty pairs.
Illustrative paired dataset, twelve respondents
| Respondent | Clarity 1-5 | Stay 0-10 |
|---|---|---|
| R01 | 4.25 | 8 |
| R02 | 3.00 | 5 |
| R03 | 4.75 | 9 |
| R04 | 2.50 | 4 |
| R05 | 3.75 | 6 |
| R06 | 4.00 | 8 |
| R07 | 3.50 | |
| R08 | 2.25 | 5 |
| R09 | 4.50 | 7 |
| R10 | 3.25 | 6 |
| R11 | 1.75 | 3 |
| R12 | 3.75 | 9 |
In that dataset R07 has no intent-to-stay answer, so the filter keeps 11 of the 12 rows and both rank columns hold the same 11 respondents. Paste it into a blank sheet with the header in row 1 and every formula on this page runs as written, with the ranges changed from 413 to 13.
Writing the correlation statement
Four things belong in the sentence: direction, coefficient, interval and the number of pairs. Anything you write without them will be quoted back at you with the caveats stripped off. The template carries the worked example's figures and says so in its first line; that line stays until your own numbers replace it.
Illustrative write-up, numbers from the worked example
Illustrative figures: replace r, the interval, N and p with your own results before this leaves your desk.
Role clarity and intent to stay were positively related (r = 0.34, 95% CI 0.25 to 0.42, N = 412, p < 0.001).
Employees who rated their role as clearer also reported higher intent to stay.
The two answers share about 12 per cent of their variation, so most of what moves intent to stay is not measured here.
Both answers came from the same questionnaire in the same fielding window, so some of the relationship may be shared method rather than a real link.
This is one snapshot. It cannot show that clearer roles keep people: manager support could lift both, and people who already plan to stay may read their role as clearer.
Drives, impacts, leads to, boosts, causes. Each one promises an experiment you did not run. "Related to", "associated with" and "goes with" carry the same information and survive a challenge.
What a correlation study cannot tell you
Four failure modes account for nearly every survey correlation that falls apart under questioning. They are worth naming individually, because the fix is different in each case.
A third thing moves both columns
Manager support plausibly raises how clear a role feels and how likely someone is to stay, which would produce a relationship between the two without either touching the other. The textbook example is Kanner and colleagues on daily hassles and physical symptoms: the strong positive relationship they found "is also consistent with the idea that symptoms cause hassles or that some third variable (e.g., neuroticism) causes both".1 Measuring the suspected third variable and holding it constant is what regression is for: it lets you adjust for measured variables, though never for the ones you did not measure.
You cannot tell which came first
A snapshot has no time order in it. Lau lists three basic correlational designs, cohort, cross-sectional and case-control, and almost every survey correlation is the cross-sectional one, measured at a single moment.4 Asking the second question three months after the first establishes measurement order, which is worth having, but it does not by itself establish causation: whatever was true of the person before the first measurement is still in both answers.
Both answers came from the same person, in the same mood
When two measures share a respondent, a format and a moment, part of what they share is the respondent's general positivity rather than anything about the constructs. Podsakoff and colleagues catalogue the mechanisms and the remedies at length.5 Three of those remedies cost nothing: separate the two blocks in the questionnaire, change the answer format between them, and pull one variable from records instead of self-report, such as tenure from the HR system rather than a question about tenure. Acquiescence, the habit of agreeing with whatever is asked, inflates agreement-scale pairs specifically, which is one of the patterns under response bias.
You looked at every pair until one worked
The median template in our library asks 13 questions. Cross every question with every other and that is 78 pairs. At the usual 0.05 threshold you would expect about four of them to come back "significant" with no relationship present at all. A twenty-question survey gives 190 pairs and roughly ten false positives. Dashboards that let you cross anything with anything are correlation-mining engines, and the only defence is to write down which pair you care about before you look.
Name the one relationship you are testing, in writing, before the survey goes out. Everything else you find is a candidate for the next survey, not a result from this one.
Questions people ask
Is a correlation survey an experiment?
No, and the dividing line is assignment. In an experiment you decide who gets what and the coin toss makes the groups comparable on everything you did not think of. In a survey you take people as they arrive. That single difference is why an experiment can support a claim about cause and a survey cannot, however large its numbers.
What is a negative correlation in a survey?
Workload and work-life balance, typically. People who report the heaviest weeks rate their balance lowest, so as one column climbs the other falls and the coefficient carries a minus sign. A minus tells you the direction, never the strength: -0.40 is a firmer finding than +0.15. Watch for the fake ones, where a reverse-coded item was never flipped back and the minus sign is a data-entry artefact.
Can I correlate an open-text answer?
Only after you turn it into a number. Code each comment into categories and you get a yes/no column per category, which pairs with a score through point-biserial. Score each comment for sentiment on a fixed scale and you get something Pearson will accept. Either way somebody has to read the comments first, and two coders should agree on a sample before the numbers count.
Why did my correlation shrink when I split by department?
Usually range restriction. Within one department people are more alike on both questions than across the whole company, and a coefficient needs spread to register anything: squeeze the spread and it falls towards zero even where the relationship is real. Occasionally the reverse happens and every department shows a relationship the company-wide figure hid. Both are worth reporting; neither means the first number was wrong.
How we counted
The library figures on this page come from parsing every question in the SuperSurvey template library on 10 September 2026: 109 live templates and 1,422 questions, of which 644 are rating items, 516 single-select, 221 open text and 41 multi-select. Only live template rows were counted; drafts and retired rows were left out. The 78-pair figure is 13 questions taken two at a time, using the library's median template length.
Those counts describe what one survey vendor publishes as good practice. They are not a sample of the world's surveys, and no claim here rests on them being one.
The intervals in the second chart are the Fisher z transformation applied to r = 0.34 at each pair count, which is the same arithmetic as the Excel formulas above. The p-value is the two-tailed t test for a correlation on N minus 2 degrees of freedom. The X-squared table and the twelve-row dataset are computed in the page build, not typed, and the build refuses to run if r for the curve is not zero or if the p-values for +0.34 and -0.34 differ. The benchmark figures in the first chart are read from Table 2 of Bosco and colleagues and are reproduced, not recalculated.
References
- Price, P. C., Jhangiani, R., and Chiang, I-C. A. (2015). Correlational research. In Research Methods in Psychology, 2nd Canadian Edition. BCcampus Open Education. opentextbc.ca/researchmethods/chapter/correlational-research
- Bosco, F. A., Aguinis, H., Singh, K., Field, J. G., and Pierce, C. A. (2015). Correlational effect size benchmarks. Journal of Applied Psychology, 100(2), 431-449. doi.org/10.1037/a0038047
- Gignac, G. E., and Szodorai, E. T. (2016). Effect size guidelines for individual differences researchers. Personality and Individual Differences, 102, 74-78. doi.org/10.1016/j.paid.2016.06.069
- Lau, F. (2017). Chapter 12: Methods for correlational studies. In Lau, F., and Kuziemsky, C. (eds.), Handbook of eHealth Evaluation: An Evidence-Based Approach. University of Victoria.
- Podsakoff, P. M., MacKenzie, S. B., Lee, J-Y., and Podsakoff, N. P. (2003). Common method biases in behavioral research: a critical review of the literature and recommended remedies. Journal of Applied Psychology, 88(5), 879-903. doi.org/10.1037/0021-9010.88.5.879
- SuperSurvey Question Corpus, 109 live templates and 1,422 questions, read 11 September 2026.
Michael Hodge: Survey methodologist and editor, SuperSurvey. Bachelor of Science (Psychology), University of Wollongong, with coursework in psychometrics and research methods. Designing surveys since 2003. About the author and how these guides are reviewed
What to read next
The rest of the survey learning centre covers the decisions on either side of this one.
Set up the correlation survey
Both questions in one form, or two forms joined on the same respondent, order or staff ID. Paired observations linked to the same unit are the only data shape a correlation can be computed from, and the timing is yours to fit to the design.
Create a survey with paired questionsFree to send, no card.