This deck covers the gradient and directional derivatives, steepest ascent, tangent planes and normal lines, relative and absolute extrema with the Second Partials Test, and constrained optimization by Lagrange multipliers. It targets the classic traps: using a direction vector that is not a unit vector, misreading the D-test, forgetting the boundary, and dropping the constraint equation.
Subject: Calculus III · 112 slides · symbolic lesson
Open the interactive version of this deck · Homework for this lesson
Objectives
By the end of this deck you will be able to:
1. Compute the gradient of a function and use it to find a directional derivative.
2. Explain why the gradient points in the direction of steepest ascent and is perpendicular to level curves and surfaces.
3. Write the tangent plane and normal line to a surface.
4. Find and classify critical points with the Second Partials Test, and locate absolute extrema on a closed region.
5. Solve constrained optimization problems with Lagrange multipliers, including two constraints.
Warm-up
Discussion prompt
Before we open Week 7 - Gradient, Extrema & Lagrange Multipliers: without looking back, what was the main idea of Week 6 - Partial Derivatives & Chain Rules, and what could you do by the end of it that you could not do before?
Hint: One sentence for the idea, one for the skill. If the second one is blank, that is the part to revisit.
Answer:
That deck covers partial derivatives - hold the other variable constant, and read the result as a slope - then higher-order and mixed partials with Clairaut's theorem, the total differential and linear approximation, the multivariable chain rule with tree diagrams, and implicit differentiation. It targets the classic traps: forgetting to hold a variable constant, confusing the order of a mixed partial, and dropping a term by using the single-variable chain rule where there are several paths.
Section
Concept
For a function of two variables, the gradient collects both first partial derivatives into a single vector.
\[ \nabla f(x,y) = \left\langle f_x(x,y),\; f_y(x,y) \right\rangle \]
gradient — The vector of first partial derivatives, written grad f or with the nabla symbol. It is vector-valued: at each input point it returns a vector.
Counterexample
Discussion prompt
For a function of two variables, the gradient collects both first partial derivatives into a single vector.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
For three variables the idea is identical: one slot per variable.
\[ \nabla f(x,y,z) = \left\langle f_x,\; f_y,\; f_z \right\rangle \]
The gradient always has as many components as the function has input variables.
Analogy
Discussion prompt
Explain The gradient in three variables by analogy to something with no Calculus III in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
For three variables the idea is identical: one slot per variable.
Ranking
Put in order
Put the moves of Worked example: computing a gradient into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. This gives the first component of the gradient.
Worked example
Find the gradient of the function below, then evaluate it at the point one, two.
\[ f(x,y) = x^2 y + 3y \]
Differentiate with respect to x, holding y constant
Why: This gives the first component of the gradient.
\[ f_x = 2xy \]
Differentiate with respect to y, holding x constant
Why: The first term contributes x-squared; the term 3y contributes 3.
\[ f_y = x^2 + 3 \]
Assemble the gradient
Why: Stack the two partials into one vector.
\[ \nabla f = \left\langle 2xy,\; x^2+3 \right\rangle \]
Evaluate at the point one, two
Why: Substitute x equals 1 and y equals 2 into each component.
\[ \nabla f(1,2) = \left\langle 4,\; 4 \right\rangle \]
Verify by recomputing each component at the point
Why: The x-partial is two times 1 times 2, which is 4; the y-partial is 1 plus 3, which is 4. Both match.
\[ \nabla f(1,2) = \langle 4,\, 4\rangle \]
Picture it
Animation
Shows: Each line of the worked example "computing a gradient", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The x-partial is two times 1 times 2, which is 4; the y-partial is 1 plus 3, which is 4. Both match.
Intuition
Do not draw the gradient on the surface itself. It lives in the input plane, the floor beneath the graph.
At every point of the floor the gradient is an arrow. Its direction tells you which way to walk to climb fastest, and its length tells you how steep that climb is.
That single picture explains almost everything else in this deck.
Explain it
Discussion prompt
Explain Picture the gradient as an arrow on the floor to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
Do not draw the gradient on the surface itself. It lives in the input plane, the floor beneath the graph.
Concept
A partial derivative measures the slope as you move parallel to an axis. But you can walk in any direction.
The directional derivative measures the instantaneous rate of change of the function as you move away from a point in a chosen direction.
Intuition
Imagine standing on a hill. Face east and the ground might rise gently. Turn to face uphill and it rises steeply. Face along the slope and it stays level.
Same spot, different directions, different slopes. The directional derivative is exactly that slope, once you fix which way you are facing.
Concept
To get the slope in a direction, dot the gradient with a unit vector pointing that way.
\[ D_{\mathbf{u}} f = \nabla f \cdot \mathbf{u}, \qquad |\mathbf{u}| = 1 \]
The direction vector must have length one. If you are handed a direction that is not a unit vector, divide it by its own length first.
\[ \mathbf{u} = \frac{\mathbf{v}}{|\mathbf{v}|} \]
Concept
The two partial derivatives are just the directional derivatives along the coordinate axes.
\[ D_{\mathbf{i}} f = \nabla f \cdot \langle 1,0\rangle = f_x, \qquad D_{\mathbf{j}} f = \nabla f \cdot \langle 0,1\rangle = f_y \]
So the directional derivative is the general idea and the partials are two special cases of it.
Step zero
Discussion prompt
Worked example: a directional derivative — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Compute the gradient
Answer:
Worked example
Find the directional derivative of the function at the point one, two, in the direction of the vector three, four.
\[ f(x,y) = x^2 + 3xy, \quad \mathbf{v} = \langle 3, 4\rangle \]
Compute the gradient
Why: The x-partial is 2x plus 3y; the y-partial is 3x.
\[ \nabla f = \langle 2x+3y,\; 3x\rangle \]
Evaluate the gradient at one, two
Why: Substitute x equals 1 and y equals 2.
\[ \nabla f(1,2) = \langle 8,\; 3\rangle \]
Normalize the direction vector
Why: Its length is 5, so divide each component by 5 to get a unit vector.
\[ \mathbf{u} = \left\langle \tfrac{3}{5},\; \tfrac{4}{5}\right\rangle \]
Dot the gradient with the unit vector
Why: This is the directional derivative formula in action.
\[ D_{\mathbf{u}} f = 8\cdot\tfrac{3}{5} + 3\cdot\tfrac{4}{5} = \tfrac{24}{5}+\tfrac{12}{5} = \tfrac{36}{5} \]
Verify the direction really is a unit vector
Why: Three-fifths squared plus four-fifths squared is nine twenty-fifths plus sixteen twenty-fifths, which equals 1. So the slope 36 over 5 is trustworthy.
\[ \left(\tfrac{3}{5}\right)^2 + \left(\tfrac{4}{5}\right)^2 = 1 \]
Picture it
Animation
Shows: Each line of the worked example "a directional derivative", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Three-fifths squared plus four-fifths squared is nine twenty-fifths plus sixteen twenty-fifths, which equals 1. So the slope 36 over 5 is trustworthy.
Trap
Dotting the gradient with the raw direction vector instead of a unit vector.
\[ D = \langle 8,3\rangle \cdot \langle 3,4\rangle = 24 + 12 = 36 \]
This answer is five times too large, because the direction vector had length five.
Normalize first, then dot.
\[ \mathbf{u} = \left\langle \tfrac35, \tfrac45\right\rangle, \quad D_{\mathbf{u}} f = \langle 8,3\rangle\cdot \mathbf{u} = \tfrac{36}{5} \]
The correct slope is 36 over 5, which is 7.2.
Explain it to yourself
Discussion prompt
In Pattern: computing a directional derivative this move is made:
2. Turn the direction into a unit vector
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
Divide the direction vector by its length. Skip this and every answer is scaled wrong.
Pattern
1. Compute the gradient and evaluate it at the point
Why: This captures how the function changes in the coordinate directions.
2. Turn the direction into a unit vector
Why: Divide the direction vector by its length. Skip this and every answer is scaled wrong.
3. Dot the gradient with the unit vector
Why: The dot product blends the coordinate rates into the rate along your chosen direction.
Elimination
Eliminate the wrong options
What is the directional derivative of f at the point (1,2) in the direction of the vector (4,3)?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: The gradient is (2xy, x-squared) = (4, 1) at (1,2). The direction (4,3) has length 5, so the unit vector is (4/5, 3/5). The dot product is 16/5 + 3/5 = 19/5.
Check
Work it out, then choose the value.
\[ f(x,y) = x^2 y, \quad \text{point } (1,2), \quad \mathbf{v} = \langle 4, 3\rangle \]
Check your understanding
What is the directional derivative of f at the point (1,2) in the direction of the vector (4,3)?
Answer: A
Why: The gradient is (2xy, x-squared) = (4, 1) at (1,2). The direction (4,3) has length 5, so the unit vector is (4/5, 3/5). The dot product is 16/5 + 3/5 = 19/5.
Concept
Among all directions you could face, one gives the largest directional derivative. That direction is the gradient itself.
Walk along the gradient and the function increases as fast as possible. Walk exactly opposite and it decreases as fast as possible.
Intuition
The directional derivative is the gradient dotted with a unit vector. A dot product with a fixed vector is largest when the two point the same way.
\[ D_{\mathbf{u}} f = \nabla f \cdot \mathbf{u} = |\nabla f|\,\cos\theta \]
The cosine is largest, equal to one, when the angle is zero, meaning your direction matches the gradient. That is why the gradient is the steepest direction.
Concept
Because the cosine tops out at one, the greatest possible rate of increase equals the length of the gradient.
\[ \max D_{\mathbf{u}} f = |\nabla f|, \qquad \min D_{\mathbf{u}} f = -|\nabla f| \]
And if you walk perpendicular to the gradient, the rate of change is zero: you are moving along a level curve.
Estimation
Predict first
At the point one, two, find the direction of steepest increase and the maximum rate of increase.
Commit before you compute: what does Worked example: steepest ascent and its rate come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the rate along u equals the gradient length
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Dotting the gradient with u gives two over root five plus eight over root five, which is ten over root five, equal to two root five, matching the magnitude.
Worked example
At the point one, two, find the direction of steepest increase and the maximum rate of increase.
\[ f(x,y) = x^2 + y^2 \]
Compute the gradient
Why: Each partial is twice its variable.
\[ \nabla f = \langle 2x,\; 2y\rangle \]
Evaluate at one, two
Why: This vector already points in the steepest-ascent direction.
\[ \nabla f(1,2) = \langle 2,\; 4\rangle \]
The maximum rate is the length of the gradient
Why: Steepest slope equals the magnitude of the gradient.
\[ |\nabla f(1,2)| = \sqrt{2^2+4^2} = \sqrt{20} = 2\sqrt{5} \]
Give the direction as a unit vector
Why: Divide the gradient by its length.
\[ \mathbf{u} = \left\langle \tfrac{1}{\sqrt5},\; \tfrac{2}{\sqrt5}\right\rangle \]
Verify the rate along u equals the gradient length
Why: Dotting the gradient with u gives two over root five plus eight over root five, which is ten over root five, equal to two root five, matching the magnitude.
\[ \nabla f \cdot \mathbf{u} = \tfrac{2}{\sqrt5} + \tfrac{8}{\sqrt5} = \tfrac{10}{\sqrt5} = 2\sqrt5 \]
Picture it
Animation
Shows: Each line of the worked example "steepest ascent and its rate", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Dotting the gradient with u gives two over root five plus eight over root five, which is ten over root five, equal to two root five, matching the magnitude.
Check
Find the largest possible rate of increase at the given point.
\[ f(x,y) = x^2 + y^2, \quad \text{point } (3,4) \]
Check your understanding
What is the maximum rate of change of f at the point (3,4)?
Answer: A
Why: The gradient is (2x, 2y) = (6, 8) at (3,4). The maximum rate of change equals the magnitude of the gradient, which is the square root of 36 plus 64, equal to the square root of 100, which is 10.
Section
Concept
A level curve is the set of input points where the function equals a fixed value. Along it, the function does not change.
Moving along a level curve gives a directional derivative of zero. Since the gradient dotted with that direction is zero, the gradient must be perpendicular to the level curve.
\[ \nabla f \perp \{\, f(x,y) = c \,\} \]
Concept
The same reasoning works one dimension up. A level surface is where a three-variable function equals a constant.
\[ \nabla F \perp \{\, F(x,y,z) = c \,\} \]
This is the key that unlocks tangent planes: the gradient of F gives a normal vector to the surface.
Picture it
Figure (svg): Three concentric circles as level curves with a red gradient arrow pointing radially outward, crossing them at right angles.
Discussion prompt
Read the picture before the words. What is this showing, and what is the one thing it is built to make obvious? Commit to an answer, then read on.
Hint: Name the parts, then say what changes between them — and if nothing changes, say what is being held still.
Answer:
On a topographic map, the closed loops are level curves of elevation. The steepest path, where water runs downhill, always crosses the contour lines at right angles.
Intuition
Figure (svg): Three concentric circles as level curves with a red gradient arrow pointing radially outward, crossing them at right angles.
On a topographic map, the closed loops are level curves of elevation. The steepest path, where water runs downhill, always crosses the contour lines at right angles.
That right angle is the gradient being perpendicular to the level curves. Closely spaced contours mean a long gradient and a steep slope.
Concept
Write the surface as a level surface: move everything to one side so the surface is where a function equals zero.
\[ F(x,y,z) = 0 \]
The gradient of F at a point on the surface is normal to the surface. A plane through that point with that normal is the tangent plane.
\[ F_x(P)\,(x-x_0) + F_y(P)\,(y-y_0) + F_z(P)\,(z-z_0) = 0 \]
Concept
The normal line passes through the point and runs straight along the gradient, perpendicular to the tangent plane.
\[ \langle x,y,z\rangle = \langle x_0,y_0,z_0\rangle + t\,\nabla F(P) \]
In parametric form each coordinate starts at the point and moves in proportion to the matching gradient component.
Intuition
Zoom in on a smooth surface far enough and it looks flat, like standing on the curved Earth. That local flat sheet is the tangent plane.
Near the point of contact the tangent plane and the surface are almost the same, which is why the plane is the linear approximation to the surface there.
Step zero
Discussion prompt
Worked example: tangent plane and normal line — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Write the surface as F equals zero
Answer:
Worked example
Find the tangent plane and the normal line to the sphere at the point one, two, two.
\[ x^2 + y^2 + z^2 = 9 \]
Write the surface as F equals zero
Why: Subtract 9 so the sphere becomes a level surface of F.
\[ F(x,y,z) = x^2+y^2+z^2-9 \]
Compute the gradient of F
Why: It will serve as the normal vector to the surface.
\[ \nabla F = \langle 2x,\; 2y,\; 2z\rangle \]
Evaluate at one, two, two
Why: This is the normal vector at the point of tangency.
\[ \nabla F(1,2,2) = \langle 2,\; 4,\; 4\rangle \]
Write the tangent plane
Why: Use point-normal form, then divide by 2 to simplify.
\[ 2(x-1)+4(y-2)+4(z-2)=0 \;\Rightarrow\; x+2y+2z = 9 \]
Write the normal line
Why: Start at the point and follow the gradient direction.
\[ x=1+2t,\quad y=2+4t,\quad z=2+4t \]
Verify the point lies on the surface and the plane
Why: One plus four plus four equals 9, so the point is on the sphere; and one plus two times 2 plus two times 2 equals 9, so it satisfies the plane.
\[ 1^2+2^2+2^2 = 9 \quad\text{and}\quad 1+2(2)+2(2)=9 \]
Picture it
Animation
Shows: Each line of the worked example "tangent plane and normal line", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: One plus four plus four equals 9, so the point is on the sphere; and one plus two times 2 plus two times 2 equals 9, so it satisfies the plane.
Explain it to yourself
Discussion prompt
In Pattern: tangent plane and normal line this move is made:
1. Rewrite the surface as F(x,y,z) = 0
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
Even a graph z equals g(x,y) becomes F equals g(x,y) minus z, set to zero.
Pattern
1. Rewrite the surface as F(x,y,z) = 0
Why: Even a graph z equals g(x,y) becomes F equals g(x,y) minus z, set to zero.
2. Compute the gradient of F and evaluate at the point
Why: This gradient is the normal vector to the surface.
3. Plane: use point-normal form with that normal
Why: Normal components times coordinate differences, summed to zero.
4. Line: start at the point and add t times the normal
Why: The normal line is parallel to the gradient.
Prediction
Predict first
Which equation is the tangent plane to the surface z = x-squared + y-squared at the point (1,1,2)?
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: z = 2x + 2y - 2
Why: For a graph the plane is z = f(a,b) + f_x(a,b)(x - a) + f_y(a,b)(y - b). Here f_x = 2x = 2, f_y = 2y = 2, and f(1,1) = 2, giving z = 2 + 2(x - 1) + 2(y - 1) = 2x + 2y - 2.
Check
Use the tangent-plane formula for a graph.
\[ z = x^2 + y^2, \quad \text{point } (1,1,2) \]
Check your understanding
Which equation is the tangent plane to the surface z = x-squared + y-squared at the point (1,1,2)?
Answer: A
Why: For a graph the plane is z = f(a,b) + f_x(a,b)(x - a) + f_y(a,b)(y - b). Here f_x = 2x = 2, f_y = 2y = 2, and f(1,1) = 2, giving z = 2 + 2(x - 1) + 2(y - 1) = 2x + 2y - 2.
Section
Concept
A relative maximum is a point higher than every nearby point; a relative minimum is lower than every nearby point.
relative extremum — A point where the function value is at least as large (a maximum) or at least as small (a minimum) as at all nearby points in the domain.
Concept
At a smooth peak or valley the tangent plane is flat, so both partial derivatives are zero. Those inputs are called critical points.
\[ \nabla f = \langle f_x, f_y\rangle = \langle 0, 0\rangle \]
Every relative extremum of a differentiable function occurs at a critical point. But not every critical point is an extremum.
Intuition
Stand exactly on top of a hill. Look in any direction and, for an instant, the ground is level. No direction goes up and none goes down.
That is the gradient being the zero vector: every directional derivative is zero. The same is true at the bottom of a bowl.
Sorting
Sort into buckets
These are the pieces of Week 7 - Gradient, Extrema & Lagrange Multipliers, out of order. Put each one back under the part of the lesson it belongs to.
Concept
Finding a critical point is not enough. The Second Partials Test uses second derivatives to decide what kind of point it is.
\[ D = f_{xx}\,f_{yy} - \left(f_{xy}\right)^2 \]
Compute this quantity at the critical point, then read the verdict from its sign together with the sign of the second partial in x.
Intuition
In one variable a positive second derivative means the curve holds water like a cup, a minimum. The Second Partials Test extends that idea to surfaces.
The mixed term guards against saddle shapes, where the surface curves up one way and down another. When the mixed curvature wins, the discriminant goes negative and you have a saddle.
Concept
Evaluate the discriminant at the critical point, then use this table.
| Sign of D | Extra condition | Conclusion |
|---|---|---|
| D > 0 | f_xx > 0 | relative minimum |
| D > 0 | f_xx < 0 | relative maximum |
| D < 0 | none needed | saddle point |
| D = 0 | none | test is inconclusive |
Comparison
Comparison matrix
From Reading the Second Partials Test: refill the Extra condition column from what you know. The rest of the table is as it appeared.
| Sign of D | Extra condition | Conclusion |
|---|---|---|
| D > 0 | f_xx > 0 | relative minimum |
| D > 0 | f_xx < 0 | relative maximum |
| D < 0 | none needed | saddle point |
| D = 0 | none | test is inconclusive |
Estimation
Predict first
Find and classify all critical points of the function.
Commit before you compute: what does Worked example: find and classify critical points come out to? A rough magnitude and the right form is enough — the point is to have something concrete to be wrong about.
Correct: Verify the minimum value
Why: A prediction you can defend turns the computation into a check rather than a leap of faith — and an answer that contradicts it is caught on the spot. Evaluating gives f at (2,0) equal to 8 minus 24, which is negative 16, the relative minimum value.
Worked example
Find and classify all critical points of the function.
\[ f(x,y) = x^3 - 12x + y^2 \]
Set both first partials to zero
Why: Critical points solve grad f equals the zero vector.
\[ f_x = 3x^2 - 12 = 0, \quad f_y = 2y = 0 \]
Solve the system
Why: The first equation gives x-squared equals 4, so x is 2 or negative 2; the second gives y equals 0.
\[ (2,0) \quad\text{and}\quad (-2,0) \]
Compute the second partials
Why: These feed the discriminant.
\[ f_{xx}=6x,\quad f_{yy}=2,\quad f_{xy}=0 \]
Form the discriminant
Why: The discriminant is f_xx times f_yy minus the mixed partial squared, which simplifies to 12x.
\[ D = (6x)(2) - 0^2 = 12x \]
Classify the first point
Why: At (2,0) the discriminant is 24, positive, and f_xx is 12, positive, so this is a relative minimum.
\[ D(2,0) = 24 > 0,\quad f_{xx} = 12 > 0 \]
Classify the second point
Why: At (negative 2, 0) the discriminant is negative 24, so this is a saddle point.
\[ D(-2,0) = -24 < 0 \]
Verify the minimum value
Why: Evaluating gives f at (2,0) equal to 8 minus 24, which is negative 16, the relative minimum value.
\[ f(2,0) = 2^3 - 12(2) = -16 \]
Picture it
Animation
Shows: Each line of the worked example "find and classify critical points", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Evaluating gives f at (2,0) equal to 8 minus 24, which is negative 16, the relative minimum value.
Trap
At the point negative two, zero, seeing that the x second partial is negative and calling it a relative maximum.
\[ f_{xx}(-2,0) = -12 < 0 \;\Rightarrow\; \text{max?} \]
This skips the discriminant. The sign of the second partial only matters after you already know the discriminant is positive.
Check the discriminant first.
\[ D(-2,0) = 12(-2) = -24 < 0 \]
Because the discriminant is negative, the point is a saddle. The negative second partial is irrelevant here.
Break the constraint
Discussion prompt
The rule this trap just fixed:
Because the discriminant is negative, the point is a saddle. The negative second partial is irrelevant here.
Now break it on purpose. Build a case that violates it and follow the consequences until something visibly fails. Where does the failure first show up — and would you have noticed it if you had not been looking?
Hint: The dangerous rules are the ones whose violation still produces an answer. If yours fails loudly, try to find one that fails quietly.
Pattern
1. Solve grad f equals zero for all critical points
Why: Both first partials must vanish at the same input.
2. Compute the three second partials and form D
Why: The discriminant is f_xx times f_yy minus the mixed partial squared.
3. If D is positive, use f_xx: positive means minimum, negative means maximum
Why: A positive discriminant means a genuine bowl or dome.
4. If D is negative it is a saddle; if D is zero the test is inconclusive
Why: Negative D means opposing curvatures; zero D needs another method.
Prediction
Predict first
Using the Second Partials Test, the critical point with these second partials is a:
Answer it in your own words, now, with nothing to choose from. The options are on the next slide — and picking the right one off a list is an easier skill than producing it.
Correct: relative maximum
Why: The discriminant is f_xx times f_yy minus the mixed partial squared, which is (-2)(-8) minus 3-squared, equal to 16 minus 9, which is 7. Since D is positive and f_xx is negative, the point is a relative maximum.
Check
At a critical point the second partials are given below. Classify the point.
\[ f_{xx}=-2,\quad f_{yy}=-8,\quad f_{xy}=3 \]
Check your understanding
Using the Second Partials Test, the critical point with these second partials is a:
Answer: A
Why: The discriminant is f_xx times f_yy minus the mixed partial squared, which is (-2)(-8) minus 3-squared, equal to 16 minus 9, which is 7. Since D is positive and f_xx is negative, the point is a relative maximum.
Concept
At a saddle point the function has a critical point that is neither a peak nor a valley. It rises as you leave in one direction and falls as you leave in another.
Think of a mountain pass or a horse saddle: the lowest point along the ridge but the highest point along the trail that crosses it.
Ranking
Put in order
Put the moves of Worked example: a second classification into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. The second equation gives x equals negative 2y; substituting into the first gives negative 3y minus 3 equals 0, so y is negative 1 and x is 2.
Worked example
Find and classify the critical point of the function.
\[ f(x,y) = x^2 + xy + y^2 - 3x \]
Set the first partials to zero
Why: This is the critical-point condition.
\[ f_x = 2x + y - 3 = 0,\quad f_y = x + 2y = 0 \]
Solve the linear system
Why: The second equation gives x equals negative 2y; substituting into the first gives negative 3y minus 3 equals 0, so y is negative 1 and x is 2.
\[ (x,y) = (2,-1) \]
Compute the second partials and the discriminant
Why: Here f_xx is 2, f_yy is 2, and f_xy is 1.
\[ D = (2)(2) - 1^2 = 3 > 0 \]
Classify
Why: D is positive and f_xx is 2, positive, so this is a relative minimum.
\[ f_{xx} = 2 > 0 \;\Rightarrow\; \text{relative minimum} \]
Verify the minimum value
Why: Evaluating gives 4 minus 2 plus 1 minus 6, which is negative 3.
\[ f(2,-1) = 4 - 2 + 1 - 6 = -3 \]
Picture it
Animation
Shows: Each line of the worked example "a second classification", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Evaluating gives 4 minus 2 plus 1 minus 6, which is negative 3.
Concept
On a closed and bounded region a continuous function is guaranteed to reach both an absolute maximum and an absolute minimum somewhere.
Those extreme values occur either at a critical point inside the region or somewhere on its boundary.
Intuition
Picture a metal plate heated in some pattern. The hottest and coldest spots are either in the interior, at a flat critical point, or pushed out onto the rim.
Miss the boundary and you can miss the true extreme entirely. Both places must be searched.
Explain it to yourself
Discussion prompt
In Pattern: absolute extrema on a region this move is made:
2. Find the extreme values of f on each boundary piece
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
Restrict f to each edge, reducing it to a one-variable problem, and include the corners.
Pattern
1. Find interior critical points and evaluate f there
Why: These are the candidate values coming from the inside.
2. Find the extreme values of f on each boundary piece
Why: Restrict f to each edge, reducing it to a one-variable problem, and include the corners.
3. Compare every candidate value
Why: The largest is the absolute maximum; the smallest is the absolute minimum.
Step zero
Discussion prompt
Worked example: absolute extrema on a rectangle — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Find interior critical points
Answer:
Worked example
Find the absolute maximum and minimum of the function on the rectangle where x runs from 0 to 3 and y runs from negative 2 to 2.
\[ f(x,y) = x^2 - 2x + y^2 \]
Find interior critical points
Why: Set both partials to zero: 2x minus 2 equals 0 gives x equal to 1, and 2y equals 0 gives y equal to 0.
\[ (1,0), \qquad f(1,0) = -1 \]
Check the left and right edges
Why: On x equal to 0, f is y-squared, ranging 0 to 4; on x equal to 3, f is 3 plus y-squared, ranging 3 to 7.
\[ x=3,\; y=\pm 2:\quad f = 3 + 4 = 7 \]
Check the top and bottom edges
Why: On y equal to plus or minus 2, f is x-squared minus 2x plus 4, whose own critical point at x equal to 1 gives 3; the corners give 7.
\[ y=\pm 2,\; x=1:\quad f = 1 - 2 + 4 = 3 \]
Collect all candidate values
Why: The interior gives negative 1; the edges and corners give values from 0 up to 7.
\[ \{-1,\; 0,\; 3,\; 4,\; 7\} \]
Verify by comparing the candidates
Why: The largest candidate, 7, occurs at (3,2) and (3, negative 2); the smallest, negative 1, occurs at (1,0).
\[ \max = 7,\qquad \min = -1 \]
Picture it
Animation
Shows: Each line of the worked example "absolute extrema on a rectangle", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: On x equal to 0, f is y-squared, ranging 0 to 4; on x equal to 3, f is 3 plus y-squared, ranging 3 to 7.
Trap
Finding only the interior critical point and stopping there.
\[ (1,0):\; f = -1 \;\Rightarrow\; \text{that is the answer?} \]
This happens to catch the minimum but completely misses the maximum, which lives on the boundary.
Search the boundary too.
\[ \text{corner } (3,\pm 2):\; f = 7 \]
The absolute maximum is 7 on the boundary. The Extreme Value Theorem guarantees it exists, but only the boundary reveals it.
Check
Find the absolute maximum on the closed square with x and y each from 0 to 2.
\[ f(x,y) = xy, \quad 0 \le x \le 2,\; 0 \le y \le 2 \]
Check your understanding
What is the absolute maximum value of f on the square?
Answer: A
Why: The only interior critical point is (0,0) where f is 0. The product xy is largest at the corner (2,2), where it equals 4, so the absolute maximum is 4.
Section
Concept
Real problems ask you to make something as large or small as possible: maximum volume, minimum cost, least material.
The recipe is familiar: write the quantity to optimize as a function, use any given relationship to reduce the number of variables, then find and classify the critical points.
Hypothesis
Predict first
Worked example: least material for an open box is about to be worked. State your hypothesis first: which rule or definition decides this one, and what is the first move it forces? Then watch whether the example agrees with you.
Correct: Reduce to two variables using the volume
Why: Solve the volume relation for z and substitute into the surface area.
A hypothesis you wrote down is falsifiable; a vague sense of how it will go is not. If the example opens somewhere else, that gap is the thing worth chasing.
Worked example
An open-top rectangular box must hold a volume of 32. Find the dimensions that use the least material.
\[ V = xyz = 32,\qquad S = xy + 2xz + 2yz \]
Reduce to two variables using the volume
Why: Solve the volume relation for z and substitute into the surface area.
\[ z = \frac{32}{xy}, \qquad S = xy + \frac{64}{y} + \frac{64}{x} \]
Set the partial derivatives to zero
Why: A minimum of S occurs at a critical point.
\[ S_x = y - \frac{64}{x^2} = 0,\qquad S_y = x - \frac{64}{y^2} = 0 \]
Solve the system
Why: Symmetry forces x equal to y; then x-cubed equals 64, so x is 4.
\[ x = y = 4, \qquad z = \frac{32}{16} = 2 \]
Confirm it is a minimum
Why: The second partials give a discriminant of 2 times 2 minus 1, which is 3, positive, with S_xx positive, so it is a minimum.
\[ S = 16 + 16 + 16 = 48 \]
Verify the volume constraint
Why: Four times 4 times 2 equals 32, matching the required volume, and the least surface area is 48.
\[ 4\cdot 4\cdot 2 = 32 \]
Picture it
Animation
Shows: Each line of the worked example "least material for an open box", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Four times 4 times 2 equals 32, matching the required volume, and the least surface area is 48.
Concept
Sometimes you cannot choose the inputs freely: they must satisfy an equation, called a constraint.
You want the largest or smallest value of one function while a second function is pinned to a fixed value.
\[ \text{optimize } f(x,y) \quad \text{subject to} \quad g(x,y) = 0 \]
Intuition
The constraint is a curve, like a fence you must stay on. As you walk along it, the value of f rises and falls.
At the highest point on the fence, f stops increasing along the fence. There the level curve of f just grazes the constraint, touching it tangentially.
Concept
Where the level curve of f is tangent to the constraint curve, the two curves share the same perpendicular direction.
Since gradients are perpendicular to their own level curves, the gradient of f and the gradient of g must point along the same line.
Concept
Pointing along the same line means one gradient is a scalar multiple of the other.
\[ \nabla f = \lambda\,\nabla g \]
Lagrange multiplier — The scalar, written lambda, that relates the gradient of the objective to the gradient of the constraint at an optimum.
Matching
Match the pairs
Match each term to the definition this lesson gave it — not the one you would guess from the word.
Why: These are the working definitions of gradient, relative extremum, Lagrange multiplier as Week 7 - Gradient, Extrema & Lagrange Multipliers uses them. Pairing them correctly is the test of whether you could state each one with the slide switched off.
Intuition
The multiplier is not just bookkeeping. It measures how fast the best value would change if the constraint level were nudged.
A large multiplier means the constraint is expensive: relaxing it a little would buy a big improvement.
Concept
The gradient condition alone is not enough. It gives directions but not the actual point.
You must solve the gradient equations together with the original constraint. The constraint is what pins the solution onto the fence.
\[ \nabla f = \lambda\,\nabla g \quad\text{and}\quad g(x,y) = 0 \]
Step zero
Discussion prompt
Worked example: Lagrange with one constraint — before any calculation: what is the plan? Name the moves in order, in plain English, without doing the arithmetic.
Hint: It starts with: Write the constraint as g equals zero
Answer:
Worked example
Minimize the function subject to the constraint. Geometrically this is the squared distance from the origin to a line.
\[ f(x,y) = x^2 + y^2 \quad\text{subject to}\quad x + 2y = 5 \]
Write the constraint as g equals zero
Why: Move the constant across so the constraint reads g equals zero.
\[ g(x,y) = x + 2y - 5 \]
Compute both gradients
Why: Gradient of f is twice the position; gradient of g is the coefficient vector.
\[ \nabla f = \langle 2x, 2y\rangle,\qquad \nabla g = \langle 1, 2\rangle \]
Set gradient of f equal to lambda times gradient of g
Why: This gives one scalar equation per variable.
\[ 2x = \lambda,\qquad 2y = 2\lambda \]
Relate x and y
Why: Substituting lambda equal to 2x into the second equation gives 2y equal to 4x, so y equals 2x.
\[ y = 2x \]
Use the constraint to solve
Why: Substitute into x plus 2y equals 5: x plus 4x equals 5, so x is 1 and y is 2.
\[ x = 1,\qquad y = 2 \]
Verify the point and value
Why: The point one, two satisfies 1 plus 4 equals 5, so it is on the line; and f equals 1 plus 4 equals 5, the minimum squared distance.
\[ f(1,2) = 1 + 4 = 5 \]
Picture it
Animation
Shows: Each line of the worked example "Lagrange with one constraint", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: The point one, two satisfies 1 plus 4 equals 5, so it is on the line; and f equals 1 plus 4 equals 5, the minimum squared distance.
Trap
Solving only grad f equals the zero vector and reporting the origin.
\[ \nabla f = \langle 2x, 2y\rangle = \langle 0,0\rangle \;\Rightarrow\; (0,0) \]
But the origin does not lie on the line x plus 2y equals 5, so it is not even an allowed point.
Keep the constraint: solve the gradient condition together with g equals zero.
\[ y = 2x,\quad x + 2y = 5 \;\Rightarrow\; (1,2) \]
The real minimizer is one, two, a point that actually sits on the constraint.
Two truths and a lie
Sort into buckets
Some of these hold up and some are the exact mistakes this lesson is built to prevent. Sort them.
Explain it to yourself
Discussion prompt
In Pattern: the Lagrange method this move is made:
3. Solve those equations together with g equals zero
Why is that legal? Name the rule or definition it rests on before you read on.
Hint: If you can only say "because that is what you do", the rule is the thing to go and find.
Answer:
Never drop the constraint; it selects the actual points on the fence.
Pattern
1. Write the constraint as g equals zero
Why: Move everything to one side.
2. Set grad f equal to lambda times grad g
Why: This produces one scalar equation per variable.
3. Solve those equations together with g equals zero
Why: Never drop the constraint; it selects the actual points on the fence.
4. Evaluate f at every solution and compare
Why: The largest and smallest values are the constrained extrema.
Real world
Discussion prompt
Outside this lesson: where does Week 7 - Gradient, Extrema & Lagrange Multipliers actually turn up? Name one concrete situation — a job, a piece of software someone ships, a decision somebody has to make — and say which part of Pattern: the Lagrange method is doing the work in it.
Hint: Vague is the failure mode here. "Engineering" is not a situation; "deciding whether this build is fast enough to ship" is.
Answer:
That deck covers the gradient and directional derivatives, steepest ascent, tangent planes and normal lines, relative and absolute extrema with the Second Partials Test, and constrained optimization by Lagrange multipliers. It targets the classic traps: using a direction vector that is not a unit vector, misreading the D-test, forgetting the boundary, and dropping the constraint equation.
Check
Minimize using Lagrange multipliers.
\[ f(x,y) = x^2 + y^2 \quad\text{subject to}\quad x + y = 4 \]
Check your understanding
What is the minimum value of f = x-squared + y-squared subject to x + y = 4?
Answer: A
Why: Setting the gradient (2x, 2y) equal to lambda times (1,1) forces x equal to y. The constraint x plus y equals 4 then gives x equal to y equal to 2, so f equals 4 plus 4, which is 8.
Concept
With two constraints the feasible set is where both are satisfied at once, usually a curve in space where two surfaces meet.
Now the gradient of f must be a combination of both constraint gradients, each with its own multiplier.
\[ \nabla f = \lambda\,\nabla g + \mu\,\nabla h \]
Explain it
Discussion prompt
Explain Two constraints, two multipliers to a student a year behind you. No notation, no jargon they have not met — and it still has to be true.
Hint: If your explanation needs a symbol they have never seen, you are describing the notation rather than the idea.
Answer:
With two constraints the feasible set is where both are satisfied at once, usually a curve in space where two surfaces meet.
Intuition
One constraint is a surface; a second constraint is another surface. Together they meet in a curve, and you optimize while walking along that curve.
The extra multiplier is simply the price of the second fence you must stay on.
Analogy
Discussion prompt
Explain The intersection curve by analogy to something with no Calculus III in it at all — a queue, a recipe, a map, a bank balance, whatever fits. Then say where your analogy breaks.
Hint: An analogy that never breaks is not an analogy, it is the same idea wearing a hat. Find the seam — that is the part that is actually new.
Answer:
One constraint is a surface; a second constraint is another surface. Together they meet in a curve, and you optimize while walking along that curve.
Ranking
Put in order
Put the moves of Worked example: Lagrange with two constraints into the order they have to happen.
Why: These are the moves of the worked example in the order it makes them, and each one is set up by the one before it. Let g be x plus y plus z minus 6, and h be x minus y minus 2.
Worked example
Find the point on the intersection of the two planes closest to the origin, by minimizing the squared distance.
\[ f = x^2+y^2+z^2,\quad x+y+z=6,\quad x - y = 2 \]
Name the constraints and their gradients
Why: Let g be x plus y plus z minus 6, and h be x minus y minus 2.
\[ \nabla g = \langle 1,1,1\rangle,\qquad \nabla h = \langle 1,-1,0\rangle \]
Write the Lagrange condition
Why: Gradient of f equals lambda times grad g plus mu times grad h.
\[ \langle 2x,2y,2z\rangle = \lambda\langle 1,1,1\rangle + \mu\langle 1,-1,0\rangle \]
Read off the three component equations
Why: One equation per coordinate.
\[ 2x = \lambda + \mu,\quad 2y = \lambda - \mu,\quad 2z = \lambda \]
Combine to relate the variables
Why: Adding the first two gives x plus y equal to lambda; and 2z equal to lambda gives z equal to half of x plus y.
\[ x + y = \lambda,\qquad z = \tfrac{x+y}{2} \]
Use both constraints
Why: Substituting z into x plus y plus z equals 6 gives three halves of (x plus y) equal to 6, so x plus y equals 4; with x minus y equals 2 this gives x equal 3, y equal 1, z equal 2.
\[ x + y = 4,\; x - y = 2 \;\Rightarrow\; x=3,\; y=1,\; z=2 \]
Verify both constraints and the value
Why: Check: 3 plus 1 plus 2 equals 6 and 3 minus 1 equals 2; the minimum squared distance is 9 plus 1 plus 4, which is 14.
\[ f(3,1,2) = 9 + 1 + 4 = 14 \]
Picture it
Animation
Shows: Each line of the worked example "Lagrange with two constraints", appearing one at a time.
The same working the example does, in the order a tutor would write it.
Takeaway: Check: 3 plus 1 plus 2 equals 6 and 3 minus 1 equals 2; the minimum squared distance is 9 plus 1 plus 4, which is 14.
Concept
If the constraint is easy to solve for one variable, plain substitution can be faster.
But when the constraint is tangled, or there are several variables and constraints, Lagrange multipliers keep the work symmetric and organized.
Counterexample
Discussion prompt
If the constraint is easy to solve for one variable, plain substitution can be faster.
That is stated as though it always holds. Do one of two things: produce a case where it fails, or say precisely what rules such a case out. "It just does" is not on the menu.
Hint: Hunt at the extremes first — zero, one, negative, empty, equal. If every extreme survives, the reason they survive is the proof.
Concept
Look back and notice one object appeared in every part of this deck: the gradient.
| Task | Role of the gradient |
|---|---|
| Directional derivative | dot it with a unit vector |
| Steepest ascent | it is the steepest direction; its length is the rate |
| Tangent plane | it is the normal vector to the surface |
| Extrema and Lagrange | set it to zero, or parallel to a constraint gradient |
Comparison
Comparison matrix
From The gradient did all the work: refill the Role of the gradient column from what you know. The rest of the table is as it appeared.
| Task | Role of the gradient |
|---|---|
| Directional derivative | dot it with a unit vector |
| Steepest ascent | it is the steepest direction; its length is the rate |
| Tangent plane | it is the normal vector to the surface |
| Extrema and Lagrange | set it to zero, or parallel to a constraint gradient |
Elimination
Eliminate the wrong options
To locate the extrema of f subject to the constraint g = 0, the Lagrange method requires you to solve which system?
3 of these 4 are wrong. Strike them one at a time, and say what rules each one out before you strike the next. The survivor is the answer.
Survives elimination: A
Why: Lagrange requires the gradient equation, grad f equals lambda times grad g, together with the original constraint g equals 0. The constraint is essential to pin down the actual points and the value of lambda.
Check
Recall exactly which equations the method needs.
Check your understanding
To locate the extrema of f subject to the constraint g = 0, the Lagrange method requires you to solve which system?
Answer: A
Why: Lagrange requires the gradient equation, grad f equals lambda times grad g, together with the original constraint g equals 0. The constraint is essential to pin down the actual points and the value of lambda.
Connect it up
Draw it
One page, no notation unless you need it: draw how these connect — Part 1 - The Gradient and Directional Derivatives · Part 2 - Tangent Planes and Normal Lines · Part 3 - Relative and Absolute Extrema · Part 4 - Applied and Constrained Optimization. Put an arrow wherever one of them is what makes another possible, and label the arrow with why.
Recap
The gradient collects the partial derivatives and points in the direction of fastest increase; its length is the greatest rate of change.
Dot the gradient with a unit vector for a directional derivative, and use it as the normal vector to build tangent planes and normal lines.
Find critical points where the gradient is zero, classify them with the Second Partials Test, and always check the boundary for absolute extrema.
For constrained problems, set the gradient of the objective parallel to the constraint gradient and solve it together with the constraint itself.
Want this taught 1-on-1? Alexander tutors Calculus III — $55/session, free consultation.