How SCAFFOLD Reduces Client Drift in Federated Learning
Federated learning gains a drift-correction mechanism for non-uniform client data.

Federated learning trains one shared model across many clients without ever pulling raw data into a central server, which is the whole point of the design and the reason it gets used in hospitals, phones, and banks. FedAvg is the algorithm most people reach for first, and it works reasonably well when every client's data looks statistically similar to every other client's, the classic IID assumption from statistics textbooks. Real deployments almost never look like that. A hospital in one city sees a different patient population than a hospital across the country, a phone keyboard app sees different words depending on who's typing, and a sensor network sees different readings depending on where it's bolted down.
That's not a minor caveat. It means the algorithm that most federated systems are built on carries a structural weakness the moment real-world heterogeneity enters the picture, and no amount of additional rounds fixes it on its own. The gap comes down to one mechanical fact: FedAvg's aggregation step, the part where the server averages everyone's updates, has no way to detect or undo the directional bias that built up during local training. It just averages. If the inputs are already pointed in the wrong direction, averaging them faithfully preserves that error rather than correcting it.
What client drift is, mechanically
Picture two clients in a federated image classifier. One only ever sees photos of cats, the other only ever sees photos of dogs. Each one runs several local gradient steps before sending anything back to the server, and each one is, quite reasonably, walking toward the best model for the data it has. With IID data, this isn't a problem, because a local batch is basically a smaller, noisier version of the global data, so the local gradient direction approximates the global gradient direction on average. With non-IID data, that approximation breaks down. The cat-only client's gradient points toward "the cat model," and the dog-only client's gradient points toward "the dog model," and neither of those is the direction the global objective actually needs.
This bias is systematic, which makes this a genuinely hard problem rather than a cosmetic one. Random noise, the kind that comes from small batch sizes or stochastic sampling, tends to cancel out when you average across many clients. Directional bias from heterogeneous data does not cancel; averaging drifted updates does not undo the drift, it just blends everyone's individual errors into a compromise that may not serve anyone. Running more local steps before communicating (a common trick to save bandwidth) compounds the problem, because each additional local step walks the client model further from the shared destination and closer to its own private one. Averaging the endpoints of several wrong journeys does not somehow produce the right destination. The research brief, grounded in Karimireddy et al., holds that the heterogeneity (non-IID-ness) in the client's data results in a "drift" in the local updates, resulting in poor performance.
How SCAFFOLD was designed to solve client drift
SCAFFOLD, short for Stochastic Controlled Averaging for Federated Learning, was built specifically to answer that mechanical failure.
The conceptual move the authors made is a genuine departure from how earlier fixes approached the same problem. Regularization-based methods try to limit how far a local model is allowed to wander from the global one, essentially putting a leash on the client. SCAFFOLD does something different: it estimates the drift directly and corrects the local gradient step itself, rather than just constraining how far the consequences of that drift are allowed to travel. That's a shift from damage control to root-cause correction.
It also isn't invented from nothing. Under a specific and fairly narrow set of conditions, one local step per round, full client participation, no added noise, SCAFFOLD reduces exactly to SAGA, one of the foundational variance-reduction algorithms in stochastic optimization. That lineage matters to anyone skeptical of a new federated learning method showing up with big promises. SCAFFOLD isn't a bespoke hack tuned for one benchmark; it's an application of a technique with an established mathematical pedigree, transplanted into the federated setting where the "variance" being reduced is really the divergence between client and server objectives. The work was published at ICML 2020, Proceedings of the 37th International Conference on Machine Learning, pp.. 5132–5143 (arXiv:1910.06378).
The control variate mechanism: how SCAFFOLD corrects each local update
The mechanism itself rests on a pair of control variates, and the asymmetry between them is the entire trick. One lives on the server and tracks an estimate of the global update direction. One lives on each client and tracks that specific client's own update direction. They are not the same object playing two roles; they are two separate trackers doing two separate jobs, and both are needed for the correction to work.
At the start of a round, the server sends out the current global model along with its own control variate. During the K local steps that follow, each client doesn't just apply its raw gradient. It subtracts its own control variate, which removes the estimated local drift, then adds the server's control variate, which substitutes in an estimate of where the global objective actually wants to go. Subtract the bias, add the correction. That's the whole move, repeated at every local step, so the correction compounds across the local training phase rather than getting applied once at the end.
After local training wraps up, each client sends back two things: the updated model, and an update to its own local control variate. The server aggregates both. Crucially, clients hold onto their control variates between rounds. They are stateful. A client's drift estimate gets sharper the more it participates, rather than resetting to zero every round. The net effect described in the original paper is elegant to state even if it's intricate to prove: each client's local update gets steered toward the direction it would have taken had it seen the global data distribution, all without that client ever actually seeing another client's data.
What SCAFFOLD's theoretical guarantees say
The headline result from the 2020 paper is a strong one: SCAFFOLD provably needs significantly fewer communication rounds than FedAvg to converge, and its convergence is not affected by how heterogeneous the client data happens to be. That second part is the load-bearing claim. It means the algorithm's guarantees don't quietly degrade as data gets weirder across clients, which is exactly the failure mode FedAvg cannot escape.
There's a sharper result buried in the quadratic case specifically. When the objective function is quadratic, SCAFFOLD can exploit similarity between clients' data to cut communication further still, and the paper claims this as the first result to actually quantify when and how local steps help distributed optimization outperform large-batch SGD. That's a fairly notable theoretical contribution on its own, independent of the federated learning application, because it puts a number on something practitioners had long assumed but hadn't proven.
None of this means the theory is finished. A 2025 analysis from Mangold and Moulines, presented at IEEE FLTA, took a closer look at the quadratic setting and found that SCAFFOLD retains a higher-order bias structurally similar to FedAvg's, and that bias does not shrink as the number of clients grows. Should that worry anyone relying on SCAFFOLD today? Not in the sense of a fatal flaw. It's a refinement of the picture. But it does flag a real gap between the intuition (drift gets corrected) and the asymptotic reality (drift gets corrected to first order, with a residual that persists at scale). The 2020 result established that the correction works in a meaningful, provable sense. The 2025 result sharpens how far that correction goes, which is what good follow-up theory is supposed to do.
How SCAFFOLD performs in practice against FedAvg and FedProx
Theory is one thing; benchmarks are another, and SCAFFOLD has been run through a fair number of them. Evaluations have used both simulated heterogeneous splits and real datasets built around naturally non-IID structure: EMNIST, where handwriting style varies by user, Sent140, where sentiment varies by the Twitter account posting it, and Shakespeare, where dialogue style varies by character. These are datasets where heterogeneity is baked into who generated the data.
Across a systematic review published in Frontiers in Computer Science in 2025, SCAFFOLD comes out ahead of both FedAvg and FedProx consistently, and the gap widens specifically in the most heterogeneous settings. There's a nice confirming detail buried in how it behaves as local steps increase. FedAvg tends to get worse with more local steps under non-IID data, because more steps just means more drift accumulated before correction ever happens. SCAFFOLD does the opposite: its performance improves as local steps increase, which is close to a direct empirical signature that the drift correction mechanism is doing its job. If SCAFFOLD were merely a regularization trick in disguise, this pattern wouldn't hold.
Where does it lead most clearly? Feature distribution skew, cases where the inputs themselves differ structurally across clients, is where the gap is widest. Where does it barely matter? Mild non-IID cases, where the improvement over FedAvg shrinks to something close to negligible. And speed of convergence isn't the same thing as speed on the clock. On CIFAR-10 under non-IID splits, SCAFFOLD reaches 75% accuracy in fewer communication rounds than FedAvg, but the wall-clock time to get there runs more than 1.5 times longer, because each of those rounds costs more to communicate. Fewer rounds and faster training are not synonyms here, and conflating them is a mistake practitioners genuinely make.
There's also a result that resists the "SCAFFOLD always wins" narrative outright. On a clinical heart disease prediction task, SCAFFOLD came in slightly behind FedProx. Small margin, but real, and it complicates a clean story. No single drift-correction method dominates every task type, and that's a useful thing to know before assuming SCAFFOLD is the default right answer.
The real costs SCAFFOLD introduces: communication, memory, and statefulness
Every one of those gains comes attached to a bill. The fewer-rounds argument can offset that in principle, but the CIFAR-10 wall-clock result already shows it doesn't always play out that way in practice. Whether the tradeoff nets out favorably depends on the bandwidth available and how expensive each round actually is to run, which is a deployment-specific question, not a universal one.
Memory is the next line item. Both the server and every participating client have to store and continuously update their own control variate, round after round. On a laptop cluster that's trivial.
Clients must persist their local control variate between rounds of participation, which conflicts with systems that treat clients as stateless, a common assumption in large-scale federated learning. Systems built around the idea that clients are stateless, drop in, contribute, disappear, come back fresh next time, cannot apply SCAFFOLD directly. That assumption is common in large-scale FL for good reason: it simplifies client churn, device turnover, and dropped connections. SCAFFOLD asks for the opposite. And this connects to a related practical wall: where only a subset of clients participates per round and that subset is chosen randomly, SCAFFOLD's correction mechanism stops functioning effectively, since the control variate estimates depend on some continuity of participation. That's a genuine limitation for the exact large-scale, high-churn deployments federated learning is often built for.
One more gap deserves its own beat rather than a passing mention. SCAFFOLD, like FedAvg and FedProx before it, has no built-in privacy-preserving mechanism. The gradient updates moving between client and server remain potentially exploitable regardless of how well the drift correction works. That distinction matters enormously for anyone deploying in healthcare or finance: correcting client drift and protecting client privacy are separate engineering problems with separate solutions, and solving one does not touch the other. Drift correction is about getting the model to converge correctly. Privacy protection requires its own deliberate, ground-up architecture, built in from the start rather than bolted on after the optimization algorithm is chosen. SCAFFOLD effectively doubles the per-round communication cost compared to FedAvg, because both model updates and control variate updates must be transmitted, a finding corroborated by multiple independent studies, including arXiv:2111.14345, which reports approximately 2× cost for SCAFFOLD and FedNova versus FedAvg.
How SCAFFOLD compares to FedProx and FedNova as competing drift corrections
SCAFFOLD isn't the only answer to client drift, and it should be mapped against the field rather than treated as the last word. FedProx takes the regularization route mentioned earlier: it adds a proximal term, an L2 penalty, to the local objective that punishes a client's model for wandering too far from the global one. That limits how much damage drift can do, but it doesn't touch the direction of the drift at all, only its magnitude. A client can still be pulled the wrong way; it just can't be pulled as far.
SCAFFOLD, by contrast, tracks the drift's direction explicitly through its control variates and subtracts it out. That's a more principled fix in the sense that it addresses the actual cause rather than dampening the symptom, but it comes at the price of statefulness and roughly double the communication load already discussed. FedNova takes yet another angle, normalizing local updates according to how many local steps each client actually performed, which targets a different facet of heterogeneity entirely (the number of steps taken rather than the direction those steps point). FedNova's communication cost is in the same rough neighborhood as SCAFFOLD's, around double FedAvg's per round.
FedProx is simpler, stateless, and easy to bolt onto an existing FedAvg deployment. SCAFFOLD is more theoretically grounded about correcting direction but demands persistent client state and more bandwidth. FedNova solves a related but distinct problem around step-count normalization. On strongly non-IID, label-skewed splits like certain MNIST partitions, these more elaborate methods don't necessarily beat plain FedAvg by much, and sometimes the more effective lever is simply running more communication rounds rather than swapping algorithms. The clinical heart disease numbers make the point concretely: SCAFFOLD scored 84.2% against FedProx's 85.0%, a margin small enough to remind anyone comparing these methods that the right choice depends on the task, not just the degree of heterogeneity.
Where the research community is taking SCAFFOLD next
SCAFFOLD published in 2020, but it hasn't sat still since. It's been extended to handle random communication intervals by Mishchenko and colleagues in 2022, adapted to finite-sum problems with variance reduction by Jiang and colleagues in 2024, and pushed into federated compositional optimization by Zhang and colleagues, also in 2024.
The most recent development is ST-GT, a December 2025 paper titled "Beyond Scaffold: A Unified Spatio-Temporal Gradient Tracking Method," which reframes SCAFFOLD through the lens of gradient tracking and proposes a unified algorithm for distributed stochastic optimization over graphs that change over time. That paper contains a formal containment result: when the number of clients sampled in a round equals the total number of clients, SCAFFOLD becomes exactly equivalent to ST-GT. In other words, SCAFFOLD turns out to be a special case, full participation, of a broader framework rather than a standalone method sitting off to the side.
A systematic review out of Murang'a University of Technology, led by Muthii and colleagues, searched seven databases across a decade of literature from 2016 to 2026, turning up 1,847 records that narrowed down to 33 included studies. That review counts the most prevalent approach as variance reduction via gradient estimation, covering 11 studies, or 34%. That concentration tells its own story: the field's center of gravity is still where SCAFFOLD (Stochastic Controlled Averaging for Federated Learning), introduced by Sai Praneeth Karimireddy, Satyen Kale, Mehryar Mohri, Sashank Reddi, Sebastian Stich, and Ananda Theertha Suresh, started. The correction mechanism at the center of SCAFFOLD hasn't been replaced. It's being sharpened, generalized, and folded into bigger frameworks, which is usually what happens to an idea that got the fundamentals right the first time.
Sources
- Scaffold Extensions for Client Drift Mitigation in Federated Learning: A Synthesis of Approaches, Limitations, and Future Directions | Zenodo
- Frontiers | Deep federated learning: a systematic review of methods, applications, and challenges
- Beyond Scaffold: A Unified Spatio-Temporal Gradient Tracking Method ⋆
- Beyond Client Averaging: A Client-Independent Second-Order Stationary-Bias Component in Stochastic SCAFFOLD
- SCAFFOLD: Stochastic Controlled Averaging for Federated Learning
- A Comparative Analysis of FedAvg, FedProx, and Scaffold in Gait-Based Activity Recognition by Evaluating Accuracy, Privacy, and Explainability


