Differentially Private Diffusion Models Generate Useful Synthetic Images

Channel:

Google TechTalks

Subscribers:

349,000

Published on May 30, 2023 6:55:19 PM ● Video Link: https://www.youtube.com/watch?v=tUHlEuWt0ig

Duration: 40:49

342 views

A Google TechTalk, presented by Sahra Ghalebikesabi, 2023/03/08
ABSTRACT: The ability to generate privacy-preserving synthetic versions of sensitive image datasets could unlock numerous ML applications currently constrained by data availability. Due to their astonishing image generation quality, diffusion models are a prime candidate for generating high-quality synthetic data. However, recent studies have found that, by default, the outputs of some diffusion models do not preserve training data privacy. By privately fine-tuning ImageNet pre-trained diffusion models with more than 80M parameters, we obtain SOTA results on CIFAR-10 and Camelyon17 in terms of both FID and the accuracy of downstream classifiers trained on synthetic data. We decrease the SOTA FID on CIFAR-10 from 26.2 to 9.8, and increase the accuracy from 51.0% to 88.0%. On synthetic data from Camelyon17, we achieve a downstream accuracy of 91.1% which is close to the SOTA of 96.5% when training on the real data. We leverage the ability of generative models to create infinite amounts of data to maximise the downstream prediction performance, and further show how to use synthetic data for hyperparameter tuning. Our results demonstrate that diffusion models fine-tuned with differential privacy can produce useful and provably private synthetic data, even in applications with significant distribution shift between the pre-training and fine-tuning distributions.

Other Videos By Google TechTalks

2023-07-03	2023 Blockly Developer Summit Day 2-15: Thoughts on Bidirectional Text to Blocks to Text
2023-07-03	2023 Blockly Developer Summit Day 2-6: Code.org - Sprite Lab
2023-07-03	2023 Blockly Developer Summit Day 2-8: Active STEM with Unruly Splats
2023-06-29	A Constant Factor Prophet Inequality for Online Combinatorial Auctions
2023-06-21	Open Problems in Mechanistic Interpretability: A Whirlwind Tour
2023-06-11	Online Prediction in Sub-linear Space
2023-06-06	Accelerating Transformers via Kernel Density Estimation Insu Han
2023-06-06	Differentially Private Synthetic Data via Foundation Model APIs
2023-06-05	Foundation Models and Fair Use
2023-05-30	Differentially Private Online to Batch
2023-05-30	Differentially Private Diffusion Models Generate Useful Synthetic Images
2023-05-30	Improving the Privacy Utility Tradeoff in Differentially Private Machine Learning with Public Data
2023-05-30	Randomized Approach for Tight Privacy Accounting
2023-05-30	Almost Tight Error Bounds on Differentially Private Continual Counting
2023-05-30	EIFFeL: Ensuring Integrity for Federated Learning
2023-05-30	Differentially Private Diffusion Models
2023-05-15	Damian Grimling \| Sentistocks \| Sentimenti \| web3 talks \| March 9th 2023 \| MC: Blake DeBenon
2023-04-21	Branimir Rakic \| CTO & Co-Founder of OriginTrail \| web3 talks \| Feb 27th 2023 \| MC: Alex Ticamera
2023-04-15	A Nearly Tight Analysis of Greedy k-means++
2023-04-15	Introduction to Length-Constrained Expanders and Expander Decompositions
2023-04-07	Improved Feature Importance Computation for Tree Models Based on the Banzhaf Value

Channel	Latest
TobisTutorials	6 hours ago
domisumReplay: Wukong	6 hours ago
GenK.	6 hours ago
AguilaCultura	6 hours ago
Míster Seitán	6 hours ago
domisumReplay: LeBlanc	7 hours ago
地球防禦軍領袖TLEarth	7 hours ago
DAminguez	7 hours ago
Zodd The Destroyer	7 hours ago
Pwnance	7 hours ago
Rocquito	7 hours ago
Dreemtum	7 hours ago
Jacefiho	7 hours ago
Void Angel	7 hours ago
Amazon Prime Video UK & IE	7 hours ago
GuideRealm	7 hours ago
Forox	7 hours ago
Carlos Gaming Channel	7 hours ago
Jorge Santos	7 hours ago
PURNIMA SAREE	7 hours ago
Hour Loop of Everything	7 hours ago
2ofus	7 hours ago
domisumReplay: Cho'Gath	7 hours ago
Riley Ravenhost	7 hours ago
domisumReplay: Yone	7 hours ago