The Response Curve
An influence function is the tangent, at $\varepsilon=0$, of the curve $\theta^\star(\varepsilon)$ traced by the minimizer as one observation's weight changes. This page draws the whole curve, marks how fast it is traversed, and lets you bend the observations so that the curve bends too.
Companion pages: Influence Functions, the Kaczmarz Way (lines and projections) and The Price of Influence (the general derivation and the costs).
1. With linear observations the curve is a straight line, traversed at a slowing pace
Take the setting of the companion page: parameters $x\in\mathbb R^2$, observations $a_i^\top x=b_i$ with unit normals, loss $f(x)=\tfrac12\sum_i(a_i^\top x-b_i)^2$, Hessian $H=\sum_i a_ia_i^\top$, and a new observation $(a,\beta)$ added with weight $\varepsilon$. Sherman–Morrison gives the minimizer in closed form: $$x(\varepsilon)=x^\star-\underbrace{\frac{\varepsilon}{1+\varepsilon q}}_{s(\varepsilon)}\;r_{\text{new}}\,H^{-1}a,\qquad q=a^\top H^{-1}a,\quad r_{\text{new}}=a^\top x^\star-\beta .$$ Everything about the shape of the response curve is in this line.
- It is straight. $x(\varepsilon)$ moves along the fixed direction $H^{-1}a$ for every $\varepsilon$. The influence function gets the direction exactly right; only the distance travelled is approximated.
- It is traversed with a Möbius speed. The distance along the line is proportional to $s(\varepsilon)=\varepsilon/(1+\varepsilon q)$, not to $\varepsilon$. The first-order prediction uses $s\approx\varepsilon$, so it overshoots by the factor $1+\varepsilon q$. Think of $s$ as the effective weight of the new observation once the rest of the data has pushed back.
- It has an endpoint. As $\varepsilon\to\infty$, $s\to1/q$ and $x(\infty)$ lands exactly on the new line: the old least-squares point projected onto the new constraint in the metric $H$. Doubling a large $\varepsilon$ barely moves anything.
- It has a pole. As $\varepsilon\downarrow-1/q$, $s\to-\infty$ and the minimizer runs off to infinity in the direction $-H^{-1}a$. At $\varepsilon=-1/q$ the augmented Hessian $H+\varepsilon aa^\top$ is singular; below it the problem is unbounded. Removing an ordinary observation is $\varepsilon=-1$, which is safe because for an observation already in the fit $q$ is its leverage, always below one.
So in parameter space the response curve is a half-line from the pole to the endpoint, the point $x^\star$ sitting somewhere along it. In the space $(x,\varepsilon)$, with the weight as a third coordinate, it is a planar hyperbola: asymptotically vertical above $x(\infty)$, asymptotically horizontal at the pole.
2. Playground
The response curve is drawn in purple with tick marks at $\varepsilon=-\tfrac12,\ \tfrac12,\ 1,\ 2,\ 4,\ 8,\ 32$ and an arrowhead at the $\varepsilon\to\infty$ endpoint. The influence-function tangent is green, with hollow markers at the same $\varepsilon$ values so you can see the first-order prediction pull away from the truth. Drag the scrubber to move along both.
Drag the circle handles to move observations (each handle sits on its observation; the second handle sets its direction). Press and drag on empty space to draw a new (red) observation.
The same curve as a function of $\varepsilon$
Parameter space hides the parameterization. Below, the chosen quantity is plotted against $\varepsilon$ from just above the pole to $\varepsilon=8$, exact in purple and first-order in green (a straight line, since it is a derivative). The dashed vertical line is the pole; the dashed horizontal line is the value at $\varepsilon\to\infty$. For lines the exact curve is a hyperbola with these two asymptotes.
3. What bending the observations changes
With $\kappa>0$ each residual is $r_i(x)=1/\kappa-\lVert x-c_i\rVert$, so $\nabla r_i$ is the unit vector pointing from $x$ toward the center $c_i$ and $\nabla^2 r_i=-\kappa_i(I-uu^\top)$ with $\kappa_i=1/\lVert x-c_i\rVert$: the residual curves along the circle. The Hessian of the total loss picks up a second term, $$\nabla^2 f=\sum_i\Big(\underbrace{\nabla r_i\nabla r_i^\top}_{\text{Gauss–Newton }G}+\underbrace{r_i\,\nabla^2 r_i}_{\text{dropped by }G}\Big).$$ Three things become visible in the playground.
- The curve bends. The direction $x(\varepsilon)$ moves in depends on where it currently is, because the normals $\nabla r_i$ rotate as $x$ moves around the circles. The influence function is still the tangent at $\varepsilon=0$, but it is now wrong in direction as well as in length. For small $\kappa$ the bend is slight; as $\kappa$ grows it becomes a clearly curved arc.
- Two tangents. The exact tangent uses the full Hessian; the Gauss‑Newton tangent uses $G$. They differ by the $r_i\nabla^2 r_i$ terms, which are large exactly when residuals are large, so the two tangents agree for a well-fitting model and disagree for a poorly-fitting one. Practical influence functions use $G$ (it is positive semidefinite and factorizes); this is the price.
- Damping shortens the tangent. Replacing $G$ by $G+\lambda I$ shrinks the step, most in the flat directions. The damped tangent is a worse tangent but often a better predictor of $x(\varepsilon)$ at finite $\varepsilon$, because the true curve also slows down. In the linear case a damping $\lambda$ mimics the Möbius factor $1/(1+\varepsilon q)$ for one particular $\varepsilon$.
The negative side of the curve is computed by continuation from $x^\star$ and stops where the augmented Hessian at the minimizer loses positive definiteness. That is the pole; for circles it can arrive sooner or later than $-1/q$, and just before it the minimizer accelerates away.