Differentiation from First Principles
Shrink h towards zero and watch a chord collapse onto the tangent. Where the derivative really comes from — the limit definition, made visual.
Every differentiation shortcut you will ever use rests on one idea: the gradient of a curve is the limit of the gradient of a chord as the chord shrinks to nothing. First principlesis where that idea is made precise — and where the derivative earns its definition.
Gradients that change as you move
A straight line has a single gradient everywhere. A curve does not: it might be climbing steeply at one point and almost flat a little further along. So “the gradient of the curve” only means something at a particular point.
To pin it down we approximate. Pick your point at x, then a nearby point a small distance h to the right, at x + h. The straight line joining the two — a chord— has a gradient we can measure with ordinary rise over run. As we shrink h, that chord swings closer and closer to the true steepness at x.
Zoom in far enough and any smooth curve looks straight. The derivative is the slope of that “infinite zoom”.
Drag the point: the chord becomes the tangent
Move the point along the curve below. The gold line is the tangent— the line that just grazes the curve at that point — and its slope is the derivative there.
First principles builds this tangent as the limiting position of a chord: start with a chord from x to x + h, then let h shrink towards zero and watch the chord settle onto the tangent.
Text description ↓Hide text description ↑
A curve with a draggable point and its tangent line. The slope of the tangent at each point is the derivative there. First principles builds this tangent as the limit of a chord between x and x + h as h shrinks to zero.
The definition as a limit
The gradient of the chord from x to x + h is the change in height divided by the change in x:
gradient of chord = [f(x + h) − f(x)] / h.
This ratio is called a difference quotient. The derivative is the value it approaches as h shrinks to zero:
Notice that we cannot simply put h = 0 in the quotient — that would be 0/0, which is meaningless. The whole trick is to simplify the fraction so that h cancels, and only then let h → 0.
Worked examples
Let f(x) = x². Substitute into the definition:
[f(x + h) − f(x)] / h = [(x + h)² − x²] / h.
Expand (x + h)² = x² + 2xh + h², so the x² terms cancel and the top becomes 2xh + h²:
= [2xh + h²] / h.
Every term on top carries a factor of h, so cancel one h from top and bottom:
= 2x + h.
Now — and only now — let h → 0. The stray h vanishes, leaving f’(x) = 2x. That is the familiar rule for differentiating x², derived from scratch.
Common mistakes
Practice
Differentiate f(x) = 3x from first principles.
Show the answer ↓Hide the answer ↑
[f(x + h) − f(x)] / h = [3(x + h) − 3x] / h = [3x + 3h − 3x] / h = 3h / h = 3. There is no h left to send to zero, so f’(x) = 3— a straight line has a constant gradient, exactly as expected.
Frequently asked questions
What does differentiation from first principles mean?+
What is the limit definition of the derivative?+
How do you differentiate x² from first principles?+
Why can’t we just set h = 0?+
The ScholarsGate Maths Team
Oxbridge & Russell Group maths tutors
Written and reviewed by ScholarsGate tutors who teach A-Level and undergraduate mathematics. Every explainer is checked for accuracy against the AQA, Edexcel and OCR specifications.
Keep exploring
Want a tutor to walk you through it?
Book a DBS-checked A-Level Mathematics tutor for a 1-on-1 lesson — online or in person.
Find a A-Level Mathematics tutor