Est.
FeaturesLong read

Where Federated Learning Is Actually Being Deployed

Regulation forced healthcare and finance to adopt federated learning.

Senior Writer · · 11 min read
Cover illustration for “Where Federated Learning Is Actually Being Deployed”
Features · October 1, 2026 · 11 min read · 2,420 words

Federated learning is moving into production because regulation has made centralizing sensitive data across institutions legally untenable. HIPAA and GDPR make it nearly impossible to pool patient records from multiple hospitals into a single training set, and that legal barrier, not any architectural preference for distributed systems, is the primary reason healthcare has turned to federated approaches. Finance runs into the same wall from a different direction: moving transaction records across borders to train one unified model is often prohibited outright or so operationally burdensome that it defeats the purpose.

Federated learning sidesteps the problem structurally. The model itself travels to wherever the data sits, each participating institution or device trains locally and sends back only mathematical updates, a central server aggregates those updates into an improved shared model, and the raw data never leaves its original location. A reasonable objection follows immediately: if the goal is just privacy, why not anonymize the data and centralize it the old way? Anonymization increasingly fails to satisfy GDPR's re-identification risk standard, and cross-border transfer restrictions apply regardless of whether data has been pseudonymized. Data that meets the irreversibility bar set out in GDPR's Recital 26 falls outside the regulation's scope entirely, but very little real-world data clears that bar. Federated learning, rather than anonymization, has become the practical workaround as a result.

The gap between FL research volume and actual clinical adoption

The volume of published federated learning research has grown rapidly, but that growth has not translated into a matching rate of clinical use. Real-world clinical deployment stands at only 5.2%. That figure sets the frame for everything that follows: this is an account of a technology in narrow, constrained production, not one in broad use.

The 2026 Annual Review of Biomedical Engineering attributes part of this gap to a structural feature of AI research generally: most studies remain academic and rarely make the transition into clinical practice, largely because researchers lack access to diverse, real-world datasets. Federated learning is proposed as the structural fix for exactly that access problem, since it lets models learn from data that never has to leave its home institution. Realizing that promise requires infrastructure that is scalable and interoperable across institutions, plus workflows aligned with regulatory review, and most institutions have not yet built either one. That is the actual work of this piece: not cataloguing the research literature, but mapping where deployments are confirmed and operating, what specific constraint made each one possible, and what is still holding the rest back.

Healthcare: the most active sector, led by medical imaging

Healthcare has pushed federated learning further from research into practice than any other sector, and even its confirmed deployments show the organizational friction that explains why the field-wide adoption rate remains low. Medical imaging leads within healthcare FL, accounting for the largest share of published studies. Standardized formats like DICOM and clearly defined clinical use cases make imaging a natural fit, and radiology, internal medicine, ophthalmology, and oncology are the specialties where most of that work concentrates. Electronic health records, genomics, and wearable-derived data all appear in the literature too, but as named modalities alongside imaging rather than as separately documented deployments at this stage.

A multi-hospital pediatric radiotherapy project for organ-at-risk segmentation ran across two clients (one at DKFZ, one at UMCU) and a single central server, using NVFlare. Institutional firewalls at both hospitals blocked direct communication between the sites, so model updates were exchanged through SURFdrive, a secure cloud storage platform, rather than a purpose-built federated pipeline. The workaround exposes how real deployments must negotiate institutional IT constraints as seriously as they negotiate algorithmic ones, and a solution built on existing cloud storage is a more honest picture of production federated learning than any clean reference architecture in a paper.

Drug discovery offers a second, larger-scale example. The MELLODDY project brought together ten pharmaceutical companies, including Bayer, GSK, and Novartis, alongside technical partners NVIDIA, Owkin, KU Leuven, and Kubermatic. The consortium trained across a combined dataset of more than 2.6 billion confidential experimental activity data points, covering millions of physical small molecules and tens of thousands of assays, without any single partner ever exposing its proprietary compound data to the others. The collaborative models that came out of this were better at categorizing molecules by pharmacological or toxicological activity, showed a wider applicability domain, and were better at estimating activity values than models any single company could have trained alone.

Regulatory pressure is now pushing in a direction that makes multi-site federated learning less of a technical option and more of a compliance requirement. The FDA is beginning to require validation studies across multiple sites, with analyses proving that a model's relevance holds across diverse populations, a response to the fact that most healthcare models today are trained on data from only three states. Single-site training will increasingly fail to clear that evidentiary bar. When it does, the sector's forcing function shifts from a legal barrier to data pooling into a legal requirement for the kind of multi-site collaboration federated learning already enables. The hardest problem in practice is not model architecture. Getting "blood pressure" or any other clinical field to mean the same thing across every participating institution's records is the harmonization work that holds all of this together.

Mobile: the most mature FL deployment at scale

If healthcare shows federated learning working across a handful of institutions, Google Gboard shows it working across millions of individual devices, and it remains the most mature, largest-scale deployment of the technology in existence. Google first introduced federated learning in 2016 as a method for communication-efficient, on-device training from decentralized data. Gboard now applies it to next-word prediction, smart compose, on-the-fly rescoring, and emoji suggestion, with millions of Android devices collaborating on shared models while users' private text never leaves their own phones. Every next-word prediction neural network model in Gboard carries a differential privacy guarantee, and all future launches of Gboard neural network language models will require the same guarantee. That is a standing policy applied to every future launch, not a one-off technical flourish. Production systems at this scale, from Google, Apple, and Meta alike, demonstrate that federated learning can operate across millions of devices and multiple learning domains while still delivering meaningful differential privacy guarantees.

What does that scale actually prove about privacy, though? Gboard is also the clearest illustration of where federated learning's privacy protection stops. Published research has demonstrated working attacks against the very federated system used to train Gboard's next-word prediction model, which confirms that federated learning does not eliminate privacy risk so much as redistribute and reduce it. That is the strongest and most correct objection to treating FL as a privacy solution outright: it is a mitigation, not a guarantee, and gradient inversion, membership inference, and poisoning attacks remain live threats against real systems. The precise claim to make here is that federated learning combined with differential privacy and secure aggregation is meaningfully more protective than centralized training, without being unconditionally private under every threat model.

A second, less discussed limitation involves who gets represented in the model. Earlier FL systems scheduled training computation for when devices were idle, typically overnight and plugged into power, and that scheduling choice has a quiet cost: it biases the resulting model toward the data of whichever users happen to leave their phones charging at night, while other usage patterns and data distributions get underrepresented. Current production systems have not fully resolved that bias. A scheduling decision made for battery and bandwidth reasons ends up shaping whose data teaches the model, an effect with no clean parallel in centralized training.

Finance: cross-border AML and the unsolved interpretability problem

Finance shares healthcare's basic regulatory forcing function, a legal barrier to moving sensitive data across borders, but it runs into an obstacle that neither healthcare nor mobile FL faces with the same intensity: regulators and customers expect financial decisions to be explainable, and federated training compounds the opacity that complex models already carry. Banking Circle's expansion into the US market makes the forcing function concrete. Regional differences in transaction patterns and strict data-transfer constraints limited how well models trained only on European data would generalize, so the firm adopted Flower's federated learning platform to train anti-money-laundering models across regions without moving sensitive transaction data across borders.

Avoiding cross-border data transfer this way still leaves federated learning without interpretability, and explainable AI methods lack the privacy protections federated learning is built to provide, so financial institutions need both regulatory compliance and customer trust at the same time. Right now, no available approach satisfies all three requirements simultaneously. Regulators increasingly expect institutions to explain individual credit and fraud decisions to the customers affected by them. FL models resist that kind of explanation just as other complex machine learning models do, and the distributed nature of federated training adds a further layer of opacity on top of the model's own complexity. There is a genuine tension buried in the proposed fix, too: using explainable AI techniques to surface how a federated model reached a decision risks exposing patterns from the shared model parameters that the federated architecture was designed to keep private in the first place.

Layered on top of the interpretability gap is a data problem finance experiences more acutely than the other sectors covered here. Transaction patterns vary by region, customer segment, and regulatory environment, producing non-IID data across clients that degrades model performance and generalization. In fraud detection, a poorly generalized model produces consequences that appear immediately, in missed fraud or wrongly blocked transactions, while a degraded recommendation model degrades gradually.

Energy and sensor networks: a 2026 real-world deployment on Madeira Island

Energy infrastructure offers a smaller-scale but unusually well-documented test of federated learning, because it was run and instrumented in the field rather than simulated. A study published in the ECML PKDD 2025 proceedings, first online in May 2026, deployed a federated learning framework for residential solar photovoltaic power forecasting using five prosumers on Madeira Island. Raspberry Pi devices handled local data collection at each site, dedicated cloud servers ran the federated training, and the system produced 24-hour forecasts at 15-minute intervals. This is the clearest documented case of federated learning operating in energy infrastructure, and its most useful finding has nothing to do with model architecture.

Offline evaluation of the trained model showed reasonable accuracy. Live deployment did not match it: real-time performance measurably declined once the system moved from controlled evaluation into actual operation. The cause traced back to something far more mundane than an algorithmic flaw. The inference pipeline pulled weather data from a different source than the one used during training, and that mismatch alone degraded performance. This is a data pipeline consistency problem, not a federated learning problem, but it generalizes well beyond solar forecasting. Any FL deployment spanning multiple institutions or edge devices will eventually run into some version of it: data collected under different conditions, from different sensors, or through a different upstream source than the one a model was trained against will erode that model's accuracy no matter how sound the underlying algorithm is. Data harmonization in healthcare and pipeline consistency in energy are, underneath the different vocabulary, the same category of failure: distributed training amplifies the cost of inconsistent data in ways centralized training does not.

The privacy motivation behind the Madeira deployment mirrors healthcare's, too. Prosumer data, meaning individual household solar generation and consumption records, is sensitive information, and centralizing it raises data-sovereignty concerns comparable to those surrounding patient records. Grid operators already face customer data regulations that function much like GDPR in practice, giving energy providers the same legal incentive toward federated architectures that hospitals and banks have.

What the confirmed deployments share: non-IID data

Set healthcare, mobile, finance, and energy side by side: client data is non-IID in every confirmed deployment, and federated averaging, the standard aggregation algorithm underlying most production FL systems, was never designed to handle that condition well. In centralized machine learning, data can usually be treated as independent and identically distributed across the training set. In federated learning, each client's local data is shaped by its own geography, its own user behavior, its own sensor configuration, so the assumption breaks down by default, and the practical consequences are degraded performance, weaker generalization, and slower convergence.

The same underlying obstacle appears under a different name in each sector examined here: data harmonization in healthcare, regional transaction differences in finance, scheduling-driven sampling bias in mobile, and mismatched weather sources in energy. Client dropout, unreliable network connections, and the tradeoff between privacy-preserving overhead and training efficiency remain largely unresolved in production systems; an algorithm paper's benchmark results capture none of that friction. The SURFdrive workaround in the pediatric radiotherapy deployment is a specific case of this broader pattern: real federated learning systems spend as much effort navigating institutional infrastructure as they spend on the learning algorithm itself.

The infrastructure layer that now exists

Given how much friction occurs at the data and institutional level, the tools themselves may no longer be the constraint. The evidence from these deployments says no. The framework ecosystem for federated learning has matured enough that tooling is no longer the primary obstacle standing between an organization and a working deployment. The bottleneck has moved to organizational governance and regulatory alignment instead.

NVIDIA's NVFlare is now production-grade for healthcare applications, and it is the framework underlying the DKFZ/UMCU pediatric radiotherapy deployment described earlier. Flower, an open framework used by Banking Circle to train its cross-border anti-money-laundering models, shows the same maturity applied to finance. Together with the custom cloud-and-edge pipeline built for the Madeira solar forecasting study, these examples span healthcare, finance, and energy with software that already works well enough to run in production.

That maturity reframes the remaining problem. None of the deployments examined here stalled because the software could not do what it was asked to do. The pediatric radiotherapy project ran into institutional firewalls, not an NVFlare limitation. Banking Circle's AML models run into an interpretability gap that has nothing to do with Flower's engineering. The Madeira study ran into a weather-data mismatch in its federated framework. Healthcare's 5.2% clinical deployment rate reflects institutions still building the regulatory-aligned workflows and cross-organizational governance that no framework can supply on its own. The technology has cleared the bar of working. What remains is the harder, slower work of building the institutional arrangements that let federated learning be trusted, adopted, and held accountable in the places regulation made it necessary.

Sources

  1. A Real-World Deployment of Federated Learning for Residential Solar PV Power Forecasting
  2. Federated Learning in Healthcare — 2026 Use Cases + Tools
  3. Federated Learning in Healthcare: From Research to Real-World Deployment
  4. Overcoming data scarcity through multi-center federated learning for organs-at-risk segmentation in pediatric upper abdominal radiotherapy

More in Features