Evaluate Privacy quality in an AI search product
Senior evaluation interview question on Privacy within Small Language Models.
Read full explanationOn 2 October 2026 Google Research detailed TEE-based federated learning for Gboard: a 3x smaller privacy budget and 3 weeks of training, not 2 months.

Google Research announced the next generation of its federated learning system on 2 October 2026. Phones upload encrypted training data, and only attested code running in Trusted Execution Environments (TEEs) can decrypt it. That code must match a policy published in a public log before any data is uploaded. Gboard already trains English and Japanese next-word prediction models on it. The paper reports a privacy budget 3 times smaller and training in 3 weeks instead of 2 months, with neutral results on key typing metrics.
The old setup needed trust in the server. Each phone computed model updates on-device and uploaded them. Secure Aggregation protected those uploads cryptographically, but Google says it was not compatible with state-of-the-art central differential privacy (DP). Nobody outside Google could check the server logic either. In Google's words, it "needed to be trusted to correctly add random noise" to the gradient sums.
The new system has four parts. First, upload: each device encrypts its training examples and ties them to an access policy, which lists the exact server programs allowed to read that data. The device checks that the policy is in a public transparency log (Sigstore's Rekor) before it sends anything. Second, keys: a Key Management System (KMS), itself a cluster of TEEs running Raft consensus, hands decryption keys only to workloads that match the policy. Third, training: a root TEE runs a Python training loop and hands parallel work to worker TEEs using Federated Language, an open-source orchestration language derived from TensorFlow Federated. Fourth, recovery: the program saves KMS-encrypted recovery state each round so it can resume after a failure without leaking extra information.
The policy contains the Python program itself plus the hashes of the root and worker binaries. Those binaries can be rebuilt from open source in the Confidential Federated Compute GitHub repository, and the TEEs run on AMD SEV-SNP and Intel TDX. Operators only see metrics and differentially private model weights. Uploaded data also has a TTL: after a set period no TEE can decrypt it, a limit the paper says the KMS enforces on a best-effort basis.
The privacy gain comes from scheduling. All uploads are collected before training starts, so the program can plan how often each device participates. For an English model at 5,000 rounds, the old system let a device appear in up to 8 rounds with at least 561 rounds between appearances. The TEE system cut that to 3 appearances at least 1,822 rounds apart. Fewer, more spread-out participations mean less noise for the same privacy budget. To reach a zCDP of 0.232, the new system needed a noise multiplier of 5.16; the old one would have needed 9.54.
Coverage improved too. For a Japanese model, the old system used 8.5M of 35.5M available devices over 38 days, about 23.9%. The new system collected 17.8M uploads in about 6 days and trained on every one. In a live Gboard A/B test with 3.5M devices per arm, the best TEE-trained English model used a zCDP of 0.215 against 0.641 for the production model and held Words Per Minute and Words Modified Ratio steady.
Google lists the limits itself. Current TEEs have known weaknesses, and side channels remain an open risk. Logic loaded at runtime stays proprietary, so all privacy-relevant code has to be hardcoded in the published program. The operator can still see which uploads each round uses, though not their contents. So far the system has trained models of up to 10M parameters; Google says larger models may need worker TEEs that use GPUs.
Interview angle: this is a strong case study for a privacy-preserving ML system design question. Start with the trade-off. Classic federated learning keeps raw data on the phone and sends updates. This design sends encrypted raw data, and the protection comes from attestation and policy-bound keys rather than data location. Then cover what the auditor can check: a published policy, reproducible builds, and DP enforced inside the program. Explain why central DP accounting depends on maxP and minSep, why a key TTL forces a choice between waiting for more uploads and leaving time to train, and why a recovery path that replays earlier rounds would leak privacy. Finish with the risks: side channels, the trusted hardware vendor, and scaling to larger models.
https://research.google/blog/toward-provably-private-learning-from-federated-data/ https://arxiv.org/abs/2609.31494 https://github.com/google-parfait/confidential-federated-compute https://github.com/google-parfait/federated-language
Deep explanations with architecture diagrams for every question below.
Senior evaluation interview question on Privacy within Small Language Models.
Read full explanationSenior implementation interview question on Privacy within Small Language Models.
Read full explanationStaff security interview question on Privacy within Small Language Models.
Read full explanationMid-Level conceptual interview question on Google Gemini within Model Providers.
Read full explanationJunior production incident interview question on Google Gemini within Model Providers.
Read full explanationJunior security interview question on Google Gemini within Model Providers.
Read full explanationMid-Level scenario interview question on Adam within Deep Learning.
Read full explanationJunior trade-off interview question on Backpropagation within Deep Learning.
Read full explanation