Influence Functions, the Kaczmarz Way
An interactive tutorial on how a new observation changes the residuals of the observations you already have, using the simplest possible model: least-squares in two dimensions, where every observation is a line and the solver is Kaczmarz's method of projections.
1. Observations are lines, the loss is a sum of rank-1 quadratics
Our model has two parameters $x=(x_1,x_2)\in\mathbb R^2$. Each observation $i$ says "the parameters should satisfy a linear equation": $$a_i^\top x = b_i,\qquad a_i\in\mathbb R^2,\ \|a_i\|=1 .$$ In the plane this is a line with unit normal $a_i$ and offset $b_i$. The residual of observation $i$ at parameters $x$ is $$r_i(x) = a_i^\top x - b_i ,$$ which, because $a_i$ is a unit vector, is exactly the signed distance from $x$ to line $i$. A residual is our "prediction error" on that observation.
With more than two lines that do not pass through a common point, the system $Ax=b$ is inconsistent: no $x$ has all residuals zero. Least squares picks the $x$ with the smallest total squared residual, $$f(x)=\tfrac12\sum_{i=1}^n \big(a_i^\top x-b_i\big)^2=\sum_i f_i(x),\qquad \nabla^2 f_i = a_i a_i^\top .$$ Each term $f_i$ is a quadratic whose Hessian $a_i a_i^\top$ has rank one: it only curves in the direction $a_i$, perpendicular to the line, and is flat along the line. The full Hessian $$H=\sum_i a_i a_i^\top = A^\top A$$ is invertible as soon as two lines are not parallel, and the least-squares solution is $$x^\star = H^{-1}A^\top b,\qquad\text{equivalently}\qquad \sum_i a_i\, r_i(x^\star)=0 .$$
2. Kaczmarz: solve by projecting onto one line at a time
Kaczmarz's method visits the lines one at a time. At line $i$ it moves the current iterate straight toward the line, along the normal $a_i$: $$x_{k+1} = x_k + \lambda\,\frac{b_i - a_i^\top x_k}{\|a_i\|^2}\,a_i = x_k - \lambda\, r_i(x_k)\,a_i .$$ With relaxation $\lambda=1$ this is the orthogonal projection of $x_k$ onto line $i$. Note that it is also one gradient step on the rank-1 quadratic $f_i$ with step size $\lambda$: $\nabla f_i(x)=a_i r_i(x)$. So Kaczmarz is SGD on the sum of rank-1 quadratics, with one observation per step.
If the lines all met at a point, the projections would converge to it. When the system is inconsistent, cyclic projections settle into a limit cycle: a little loop that visits each line in turn. As $\lambda\to0$ the loop shrinks onto $x^\star$ (Censor, Eggermont & Gordon, 1983). You can watch this in the playground below, "Kaczmarz" panel: run it at $\lambda=1$, then at $\lambda=0.1$.
3. The question: what does one more line do to everyone else?
Suppose we add a new observation, a line $(a,\beta)$, with a strength $\varepsilon\ge0$ (a weight on its squared residual): $$f_\varepsilon(x) = f(x) + \frac{\varepsilon}{2}\,(a^\top x-\beta)^2 .$$ $\varepsilon=0$ is the old problem, $\varepsilon=1$ treats the new line like any other, $\varepsilon\to\infty$ forces the solution onto the line. The new minimizer $x(\varepsilon)$ moves, and every old residual $r_i(x(\varepsilon))$ changes with it. An influence function is the derivative of that change at $\varepsilon=0$. In this linear model we can also get the exact change for any $\varepsilon$, so you can see how good the first-order prediction is.
4. Playground
A new (red, dashed) line is already drawn. Move the strength slider and watch $x^\star$ slide toward it while the residual segments on lines 1 to 4 change. Draw your own new line by pressing and dragging on empty canvas, or drag the small circles to move any line.
Press and drag on empty space to draw the new (red) line. Drag circle handles to move any line. The gray square is the Kaczmarz start point.
5. Walkthrough with the numbers on screen
Every quantity below is recomputed live from the playground. Drag a line or move the slider and watch it change. Vectors are written as rows to save space.
Collect the observations
Rows of $A$ are the unit normals $a_i^\top$, entries of $b$ are the offsets.
Form the Hessian as a sum of rank-1 pieces
$H=\sum_i a_i a_i^\top$. Each term has rank one; the sum is invertible because the lines are not all parallel.
Solve for the least-squares point
$x^\star = H^{-1}A^\top b$. This is the black dot. Kaczmarz with small $\lambda$ circles around it.
Read off the current residuals
$r_i = a_i^\top x^\star - b_i$, the signed distance from $x^\star$ to each line (the dashed segments). These are the "predictions" we want to track.
Measure the new line's own residual
Before we fit anything, how wrong is the current solution on the new observation? $r_{\text{new}} = a^\top x^\star - \beta$. If this is zero, the new line passes through $x^\star$ and nothing will change no matter the strength.
Compute the influence direction
Solve one linear system, $v = H^{-1}a$. This is the direction $x^\star$ will move in. Compare with the Kaczmarz direction, which is simply $a$: the influence direction is the same request, preconditioned by the curvature of all the other observations. Also record the scalar $q = a^\top H^{-1} a$, a leverage-like number that says how much the existing data already pins down the direction $a$.
Move the solution
The exact new minimizer (derived in section 6 with the Sherman–Morrison identity) and its first-order approximation are $$x(\varepsilon) = x^\star - \frac{\varepsilon\, r_{\text{new}}}{1+\varepsilon q}\,v ,\qquad x(\varepsilon)\approx x^\star - \varepsilon\, r_{\text{new}}\, v .$$ Both move along the same straight line through $x^\star$; the exact one slows down as $\varepsilon$ grows and stops on the new line, the linear one keeps going.
Update every old residual
Residuals are linear in $x$, so $\Delta r_i = a_i^\top\big(x(\varepsilon)-x^\star\big)$: $$\Delta r_i(\varepsilon) = -\frac{\varepsilon\, r_{\text{new}}}{1+\varepsilon q}\; a_i^\top H^{-1} a ,\qquad \Delta r_i \approx -\varepsilon\, r_{\text{new}}\; \underbrace{a_i^\top H^{-1} a}_{k_i} .$$ The number $k_i = a_i^\top H^{-1} a$ is the whole story of "which old observations does the new one talk to": it is large when line $i$ and the new line are nearly parallel (their normals agree) and the data is thin in that direction.
6. The math behind the steps
Exact update: a rank-1 change to the Hessian
Setting the gradient of $f_\varepsilon$ to zero gives the normal equations of the augmented problem: $$(H+\varepsilon\, a a^\top)\,x = A^\top b + \varepsilon\,\beta\, a .$$ The Hessian changed by a rank-1 term, so its inverse changes by a rank-1 term too (Sherman–Morrison): $$(H+\varepsilon a a^\top)^{-1} = H^{-1} - \frac{\varepsilon\, H^{-1}a\,a^\top H^{-1}}{1+\varepsilon\, a^\top H^{-1} a}.$$ Multiply out, use $H^{-1}A^\top b = x^\star$ and write $v=H^{-1}a$, $q=a^\top v$, $r_{\text{new}}=a^\top x^\star-\beta$: $$x(\varepsilon) = x^\star + \varepsilon\beta v - \frac{\varepsilon\, v\,(a^\top x^\star + \varepsilon\beta q)}{1+\varepsilon q} = x^\star - \frac{\varepsilon\, r_{\text{new}}}{1+\varepsilon q}\, v .$$ This is exact for every $\varepsilon\ge0$, which is what makes the linear model a good place to build intuition: the "influence function" is not an approximation to something unknown, it is the tangent of a curve we can draw.
The influence function is the derivative at $\varepsilon=0$
Differentiating, $$\frac{d x}{d\varepsilon}\Big|_{\varepsilon=0} = -\,r_{\text{new}}\, H^{-1} a = -\,H^{-1}\nabla f_{\text{new}}(x^\star),$$ which is the classic formula (Cook 1977; Koh & Liang 2017): the parameter change is minus the inverse Hessian times the gradient of the new example's loss. For the residual on observation $i$ and its loss $f_i=\tfrac12 r_i^2$, $$\frac{d r_i}{d\varepsilon}\Big|_0 = -\,r_{\text{new}}\, a_i^\top H^{-1} a,\qquad \frac{d f_i}{d\varepsilon}\Big|_0 = -\,r_i\, r_{\text{new}}\, a_i^\top H^{-1} a = -\nabla f_i(x^\star)^\top H^{-1}\nabla f_{\text{new}}(x^\star).$$ The loss influence is a bilinear form between the two gradients, with the inverse Hessian as the metric. In the plane it factors into three interpretable numbers: how wrong the model already is on $i$ ($r_i$), how wrong it is on the newcomer ($r_{\text{new}}$), and how aligned the two lines are in the geometry set by all the data ($a_i^\top H^{-1}a$).
Two ways to "project onto the new line"
Let the strength go to infinity. Then $\varepsilon/(1+\varepsilon q)\to 1/q$ and $$x(\infty) = x^\star - \frac{r_{\text{new}}}{q}\,H^{-1}a ,\qquad a^\top x(\infty) = a^\top x^\star - r_{\text{new}} = \beta .$$ So $x(\infty)$ lies exactly on the new line: it is the minimizer of the old loss subject to the new equation. Compare with a single Kaczmarz step from $x^\star$: $$x_{\text{K}} = x^\star - r_{\text{new}}\, a .$$ Both land on the new line, and both are projections of $x^\star$ onto it. Kaczmarz projects orthogonally in the Euclidean metric, moving along $a$. The least-squares update projects orthogonally in the metric defined by $H$, moving along $H^{-1}a$. It is the projection that disturbs the old residuals as little as possible in total, which is why it is generally not the shortest path. The "K" marker in the playground shows the Kaczmarz landing point; the end of the red path shows the $H$-projection. They coincide only when $a$ is an eigenvector of $H$.
How the strength slider maps onto Kaczmarz
Kaczmarz with relaxation $\lambda$ is SGD with step $\lambda$ on $\sum_i f_i$. A weight $\varepsilon$ on the new observation means its gradient is counted $\varepsilon$ times, so the playground visits the new line $\lceil\varepsilon\rceil$ times per sweep, with the last visit's relaxation scaled by the fractional part. For small $\lambda$ the limit cycle of this schedule closes in on $x(\varepsilon)$, the same point the closed-form update produces. For $\lambda=1$ the cycle is larger, but its "center of mass" still shifts in the direction $H^{-1}a$ rather than along $a$: the other lines push back on each projection, and the aggregate of all those pushbacks is exactly what the inverse Hessian computes in one shot.
What carries over beyond two dimensions
- Every formula above holds for $x\in\mathbb R^d$ unchanged, with $H=A^\top A$ a $d\times d$ matrix.
- For a general (nonlinear) model with per-example losses $\ell_i(\theta)$, replace $a_i$ by $\nabla\ell_i(\theta^\star)$ and $H$ by the Hessian of the total loss. The exact rank-1 update no longer exists, but the derivative formula $-\nabla\ell_i^\top H^{-1}\nabla\ell_{\text{new}}$ is exactly what practical influence-function methods compute, usually via an iterative solve for $H^{-1}\nabla\ell_{\text{new}}$.
- The two-projection picture survives too: an SGD step on the new example moves along its gradient, while the influence step moves along the inverse-Hessian-preconditioned gradient. The difference between them is the reaction of all the other data.
Companion page: The Response Curve. The whole curve $x(\varepsilon)$, not just its tangent: straight with a Möbius speed for lines, with a pole at $\varepsilon=-1/q$ and an endpoint at $\varepsilon\to\infty$; bend the observations into circles and watch it curve.
Companion page: The Price of Influence. The same derivation for a general model, why the inverse Hessian is not optional, and what every step costs at Llama-3-8B scale, with a live cost calculator.