Federated learning: definition, how it works and use cases
Federated learning is an AI training method in which several actors train a single model without ever centralizing or sharing their data. Each one trains the model locally, on its own data; only the model updates (its parameters) travel, never the raw data. The result is an AI trained by the collective (better than what each actor would get alone) without any data ever leaving its original infrastructure.
At Mesh, this principle boils down to one sentence: the model travels, your data stays home.
What is federated learning?
Federated learning is a machine learning technique where a model's training is distributed across several participants who keep control of their data.
In the classic (so-called centralized) approach, all the data are gathered into a single warehouse, then the model is trained on them. Federated learning reverses this scheme: the model moves to the data, not the data to the model. Each participant trains a copy of the model on its local data, then returns only the result of that training, the model parameters (the "weights"), to a coordinator that combines them.
What this reversal costs has been measured: on the three public datasets of our bench, the federated model sits 0.3 of an AUC point at most from the model you would get by gathering everything in one place, against a measurement uncertainty at least six times larger. Moving the model instead of the data therefore gives up nothing measurable. → the demonstration and its figures.
The term was introduced by Google researchers in 2016-2017, with the FedAvg (Federated Averaging) algorithm, initially to train models on users' phones without uploading their personal data to a central server. Reference: H. B. McMahan, E. Moore, D. Ramage, S. Hampson, B. Agüera y Arcas, "Communication-Efficient Learning of Deep Networks from Decentralized Data", AISTATS 2017, PMLR 54, pp. 1273-1282, proceedings.mlr.press/v54/mcmahan17a.html.
Federated learning ≠ pooling the data
An essential and often misunderstood point: federated learning does not consist of gathering the data of several actors into a shared space. The data are neither copied, transferred, nor merged. Each actor keeps its own, at home, in the clear. What is shared is the learning, not the raw material.
This is the distinction Mesh phrases as follows: an AI trained by the collective, not on pooled data.
How does federated learning work?
Training unfolds in cycles called rounds. Here is the flow of a round, as it is actually implemented in the Mesh demonstration engine:
- Broadcast. The coordinator sends the current global model to each participant. At Mesh, this send is end-to-end encrypted (RSA-OAEP 2048-bit + AES-GCM 256-bit): one envelope per participant.
- Local training. Each participant trains this model on its own data, without letting it out. At the end, it computes the difference between the received model and the obtained model: that is its "update."
- Upload. Each participant returns only its model update (numbers: the weights), encrypted, to the coordinator. No line of training data is ever serialized or transmitted.
- Aggregation. The coordinator combines the received updates (see FedAvg below) to produce a new, improved global model.
- Repeat. We start over, round after round. The global model gradually accumulates knowledge that no single participant had alone.
The scheme, in words
Picture three factories, A, B and C, each in its own building, each with its machines and sensors. At the center, a coordinator. On each round:
- the coordinator posts the same copy of the current model to A, B and C (in a sealed envelope only the recipient can open);
- A, B and C open the envelope, train the model on their own readings without anything leaving their walls, then return to the coordinator a sealed envelope containing only model settings, not their readings;
- the coordinator opens the three envelopes, averages the settings, and gets a slightly better model, which it will re-post on the next round.
After several rounds, the central model "knows" things drawn from the three factories, while nothing that factory A measures has ever been seen by B or C, nor by the coordinator.
What is FedAvg?
FedAvg (Federated Averaging) is the reference aggregation algorithm of federated learning. Its principle: the new global model is the average of the local models, weighted by each participant's amount of data. An actor that trained on 10,000 examples weighs more than one with 500. Formally, if p_i is participant i's data share and w_i its model after local training, the aggregated model equals the sum of the p_i × w_i.
It is a simple, robust mechanism, well suited to convex models. On more complex models (deep neural networks), variants exist (FedProx, SCAFFOLD, FedAdam) to handle heterogeneity between participants.
What are the use cases of federated learning?
Federated learning is relevant anywhere data is both valuable and impossible to share, for regulatory, competitive or contractual reasons. A few examples:
- Industry: cross-plant predictive maintenance. Several sites (or even several companies) train a shared failure-detection model without exchanging their often-strategic production data.
- Banking: cross-institution fraud detection. Competing banks jointly improve an anti-fraud model without sharing their customers' transactions.
- Healthcare: research on patient data. Hospitals train a model on distributed cohorts, without medical records leaving each institution.
- Cybersecurity: collaborative intrusion detection. Organizations pool threat detection without exposing their logs, which are themselves sensitive.
- Mobile / edge. The original case: training on users' devices (predictive keyboard, speech recognition) without uploading personal data.
The common thread: each actor, taken alone, has a blind spot: it only measures or observes part of reality. Federation gives each one access to what it doesn't see, without asking it to expose what it does see.
Federated learning or confidential computing: what's the difference?
These are two privacy-enhancing technologies (PET), but they don't work the same way.
| Federated learning | Confidential computing | |
|---|---|---|
| Where is the data? | At each actor, never gathered | Gathered, but inside the processor's trusted execution environment (TEE) |
| What travels? | The model parameters only | The data, encrypted, to the TEE |
| Computation happens… | locally, at each actor | centrally, inside the TEE |
| Relies on… | a distributed architecture + transport encryption | a hardware guarantee from the processor (Intel SGX, AMD SEV…) |
In short: confidential computing centralizes the computation while protecting it in hardware; federated learning does not centralize the data at all. Mesh is federated learning: data are never gathered, anywhere. The two approaches are not mutually exclusive and can be combined. → see Confidential AI.
Federated learning or data clean room: what's the difference?
A data clean room is a neutral environment where several parties deposit data to cross-reference it under strict rules (common in advertising to match audiences). There, the data are indeed brought together in one place, even if access is controlled. Federated learning, by contrast, never brings the data together: it stays with its holder, and only the model travels. The clean room mostly serves a need for one-off cross-analysis; federated learning, a need to train a shared model without transfer.
Is federated learning GDPR-compliant?
Federated learning helps with GDPR compliance, without guaranteeing it automatically.
It helps because it natively applies the principle of data minimization: personal data are neither copied, transferred, nor centralized: they stay with their controller. This strongly reduces the exposure surface and the transfers, which are at the heart of GDPR requirements.
That said, final compliance depends on the implementation: nature of the data, consortium governance, legal basis, risk of re-identification via shared parameters, complementary measures (such as differential privacy). Compliance is assessed on a concrete deployment, by a lawyer, not on a technology in general. Mesh provides the architecture that makes this compliance attainable; it does not certify it on your behalf.
What are the risks and limitations of federated learning?
- The coordinator sees the individual updates. In a basic implementation (including the Mesh demonstrator), the coordinator decrypts each update before aggregating them. The countermeasure is secure aggregation, where only the combined result is readable. It is a published method, not a Mesh invention: K. Bonawitz et al., "Practical Secure Aggregation for Privacy-Preserving Machine Learning", CCS 2017, pp. 1175-1191, eprint.iacr.org/2017/281. At Mesh, it is an option.
- Information leakage via parameters. On very large models, shared parameters can leak information about the data (gradient inversion: L. Zhu, Z. Liu, S. Han, "Deep Leakage from Gradients", NeurIPS 2019, pp. 14747-14756, arxiv.org/abs/1906.08935; membership inference: R. Shokri, M. Stronati, C. Song, V. Shmatikov, IEEE S&P 2017, pp. 3-18, arxiv.org/abs/1610.05820). The risk is low on small models, high on deep networks: hence the importance of secure aggregation at scale.
- Malicious actor. Without robustness and signature mechanisms, a participant can try to "poison" the model. Countermeasures exist (robust aggregation, attestation) and are part of the hardening.
- Coordination. Each training cycle is a coordinated event between partners, not a permanent, automatic flow.
These limitations do not disqualify the approach: they define the security perimeter to harden according to the sensitivity of the use case.
FAQ: Federated learning
What is federated learning in one sentence?
It's a method where several actors jointly train a single AI model without sharing or centralizing their data: only the model's parameters travel, never the raw data.
What's the difference between federated learning and centralized learning?
In centralized learning, all the data are gathered to train the model. In federated learning, the model moves to the data: each actor trains locally and only returns the learned parameters. The cost of that choice has been measured: on the three public datasets of our bench, the gap to the centralized model is 0.3 of an AUC point at most, less than a sixth of the measurement uncertainty.
What is FedAvg?
FedAvg (Federated Averaging) is the reference aggregation algorithm: the global model is the average of the locally trained models, weighted by each participant's amount of data.
Does federated learning fully protect the data?
It strongly reduces exposure, because the raw data never leave home. But shared parameters can, on large models, leak some information: you then add secure aggregation and/or differential privacy to strengthen protection.
Can you train an AI without sharing your data thanks to federated learning?
Yes, that's precisely its purpose: each actor trains the model locally and shares only the learned parameters, never the data. → see How to train an AI without sharing your data.
Let's see it on your data
An interactive demonstration is available on request: we'll give you access.