Scholars Gate
Find a TutorOnline Tutoring
GuidesHow It WorksBecome a Tutor
Log inGet Started
All guides
StatisticsA-Level 9 min read

Correlation & Linear Regression

Drag the points and watch the best-fit line and correlation coefficient r respond. See exactly what “least squares” minimises — and why correlation isn’t causation.

The ScholarsGate Statistics Team·Updated 08 Jul 2026

On this page

  • What correlation measures
  • See it: fit the line
  • What least squares minimises
  • Reading r (and its traps)
  • Common mistakes
  • Practice
  • FAQ

Scatter two variables against each other and a story often appears: taller people tend to be heavier, more revision tends to mean higher marks. Correlation measures how tight that story is, and regression draws the best straight line through it.

What correlation measures

Correlation captures the strength and direction of a straight-line relationship between two variables. We summarise it in a single number, the correlation coefficient r, which always lands between −1 and +1.

The sign tells you the direction: positive r means y tends to rise as x rises; negative r means y tends to fall. The size tells you the strength: the closer |r| is to 1, the more tightly the points hug a straight line. An r of exactly ±1 means the points lie perfectly on a line; an r of 0 means no linear pattern at all.

r lives between −1 and +1

Direction is the sign, strength is the magnitude. r = +0.9 is a strong positive relationship, r = −0.9 an equally strong negative one, and r = 0.1 barely any linear link at all.

r measures how “line-like” a cloud of points is — not how steep the line is.

See it: fit a line

Drag the points below and watch r respond in real time. Pull the cloud into a tight upward band and r races towards +1; scatter them into a shapeless blob and r collapses towards 0. Notice too that the gold line always pivots through the mean point (x̄, ȳ).

InteractiveFit the regression line
Loading interactive…
Drag any point, or click empty space to add one. The gold line is the least-squares fit; r updates live.
Text description ↓Hide text description ↑

An interactive scatter plot. Dragging points (or adding new ones) updates the least-squares regression line, its equation ŷ = mx + c, and the correlation coefficient r. The line always passes through the mean point (x̄, ȳ). Points forming a tight upward band give r near +1; a shapeless cloud gives r near 0.

The least-squares line

Of all the straight lines you could draw through a scatter, which is “best”? The regression line of y on x is defined as the one that makes the total squared vertical distance from the points to the line as small as possible. Those vertical gaps are called residuals— the error between each real y and the line’s prediction ŷ.

Why square the residuals?

Two reasons. First, squaring makes every gap positive, so points above and below the line can’t cancel out and fake a good fit. Second, squaring punishes big misses far more than small ones, so the line is pulled hard away from any large errors. Minimising the sum of squared residuals gives one unique best line.

A useful fact falls out of the algebra: the least-squares line always passes through the mean point (x̄, ȳ). So even before you compute a gradient, you know one point the line must go through — a handy check.

Reading r — and its traps

Interpreting r sensibly is where marks are won and lost. Values near ±1 indicate a strong linear relationship, values near 0 a weak one or none. But r is a summary, and summaries hide things.

A single outliercan drag both the line and r a long way, flattering or wrecking an otherwise clear pattern — drag one point far off in the widget above and watch r lurch. And most important of all: a strong r tells you two variables move together, not that one causes the other. Correlation is not causation.

Common mistakes

Assuming correlation proves causation

A strong r never, on its own, shows that x causes y. The link could run the other way, be pure coincidence, or — most often — be driven by a hidden third variable affecting both. Only a controlled experiment can establish cause.

Extrapolating beyond the data

The regression line is only trustworthy across the range of x you actually measured. Predicting far outside that range (extrapolation) assumes the straight-line trend continues forever, which it rarely does — a growth line fitted to a toddler would soon predict a three-metre adult.

Practice

Your turn

Across a summer, a town’s daily ice-cream sales correlate strongly with the number of drownings that day. Explain why this does not mean ice cream causes drownings.

Show the answer ↓Hide the answer ↑

The correlation is real but not causal. A hidden third variable — hot weather— drives both: heat pushes ice-cream sales up and sends more people swimming, so more drownings occur. Neither variable causes the other; both respond to temperature. A textbook example that correlation is not causation.

Where next?

The natural sequel is r², the coefficient of determination: square the correlation and you get the fraction of the variation in y that the line actually explains.

Key takeaways
  • Correlation measures the strength and direction of a linear relationship; r always lies between −1 and +1.
  • The least-squares regression line of y on x minimises the sum of squared vertical residuals.
  • Squaring the residuals stops positive and negative gaps cancelling and penalises large misses.
  • The regression line always passes through the mean point (x̄, ȳ).
  • A strong r shows association, not cause — beware outliers, extrapolation and hidden third variables.

Frequently asked questions

What does the correlation coefficient r tell you?+
The correlation coefficient r measures the strength and direction of a linear relationship, from −1 (perfect negative) through 0 (no linear relationship) to +1 (perfect positive). Values near ±0.8 or beyond indicate a strong linear trend.
What does the least-squares regression line minimise?+
The least-squares line minimises the sum of the squared vertical distances (residuals) between each data point and the line. Squaring keeps positive and negative gaps from cancelling and penalises large misses more heavily.
Why is correlation not causation?+
A strong correlation only shows two variables move together — it does not prove one causes the other. A hidden third factor, coincidence, or reverse causation can all produce correlation without a causal link.
Does the regression line always pass through the mean point?+
Yes. The least-squares regression line of y on x always passes through the mean point (x̄, ȳ), which is why it is a useful anchor when drawing or checking the line by hand.
SS

The ScholarsGate Statistics Team

Oxbridge & Russell Group maths & statistics tutors

Written and reviewed by ScholarsGate tutors who teach A-Level and undergraduate statistics. Every explainer is checked against the AQA, Edexcel and OCR specifications.

Keep exploring

A-Level Maths & Statistics tutorsThe normal distribution & z-scoresAll interactive guides

Want a tutor to walk you through it?

Book a DBS-checked A-Level Statistics tutor for a 1-on-1 lesson — online or in person.

Find a A-Level Statistics tutor
Scholars GateTutoring & Learning

ScholarsGate (Scholars Gate) is a premium UK tutoring marketplace connecting ambitious students with expert tutors across every subject, level, and admissions pathway.

hello@scholarsgate.co.uk
+44 7544 736438
Bath, United Kingdom

Platform

  • Find a Tutor
  • Online Tutoring
  • How It Works
  • Become a Tutor
  • Interactive Guides
  • Blog
  • FAQ

Subjects

  • Mathematics & Computing
  • Natural Sciences
  • Humanities
  • Languages
  • All Subjects

Admissions

  • Medicine & UCAT
  • Maths (TMUA, STEP)
  • Law & LNAT
  • Statements & Interviews
  • All Admissions
4.9/5 average tutor rating
All tutors DBS verified
Secure payments via Stripe
500+ expert tutors

© 2026 ScholarsGate Ltd. All rights reserved.

Privacy PolicyTerms of ServiceCookie PolicySafeguarding