-
Notifications
You must be signed in to change notification settings - Fork 397
Add SlaClip adaptive clipping optimizer for DP-SGD (algorithm of ICML 2026 spotlight paper) #824
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
ZsyRock
wants to merge
2
commits into
meta-pytorch:main
Choose a base branch
from
ZsyRock:feature/slaclip-pr-skeleton
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
2 commits
Select commit
Hold shift + click to select a range
File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,241 @@ | ||
| # SlaClip: Gradient Norm Slacks can be an Indicator for Adaptive Clipping in DP-SGD | ||
|
|
||
| SlaClip is a privacy-preserving adaptive clipping method for differentially | ||
| private stochastic gradient descent (DP-SGD). It uses the norm budget left | ||
| unused by clipping—the *slack*—to release a noisy, binned estimate of the | ||
| gradient-norm cumulative distribution function (CDF). That Slack Indicator can | ||
| then drive the clipping-threshold update from the SlaClip paper or another | ||
| post-processing controller. | ||
|
|
||
| This directory is a self-contained research prototype. It does not modify | ||
| `PrivacyEngine` or any other Opacus core API. | ||
|
|
||
| ## Method | ||
|
|
||
| For a per-sample gradient `g_i`, clipping threshold `C`, and `K` CDF slots, | ||
| SlaClip defines the extended gradient from equations (6)-(8) as | ||
|
|
||
| ```text | ||
| g_i^+ = [Clip_C(g_i); s_i] | ||
| lambda = C / sqrt(K) | ||
| sqrt(K) * max(C - ||g_i||, 0) = a * lambda + b | ||
| s_i = [lambda * 1(a); b; 0], where 0 <= b < lambda. | ||
| ``` | ||
|
|
||
| This construction guarantees `||g_i^+||_2 <= C`. Under add/remove adjacency | ||
| and the fixed normalization constant `B`, the extended average query therefore | ||
| has the same `C / B` L2 sensitivity as the vanilla clipped-gradient query. With | ||
| the same sampling rule and noise multiplier, it can use the same per-step | ||
| privacy-accounting parameters as DP-SGD. | ||
|
|
||
| The first `d` coordinates are exactly the native DP-SGD gradient release. The | ||
| implementation computes the gradient and `K` slack coordinates separately to | ||
| avoid materializing a `d + K` tensor. Adding independent `N(0, (sigma*C)^2)` | ||
| noise to both parts is distributionally identical to one isotropic Gaussian | ||
| draw in `d + K` dimensions. | ||
|
|
||
| After aggregation and noise, the last `K` coordinates are divided by | ||
| `B * lambda`. The resulting `optimizer.slack_indicator` is the paper's noisy, | ||
| bin-averaged CDF estimate from equation (11). Its first coordinate describes | ||
| norms near `C`; its last coordinate describes norms near zero. | ||
|
|
||
| > **Privacy boundary:** only the noisy `slack_indicator` property is a public | ||
| > output. Per-sample slack and the unnoised slack aggregate are internal to the | ||
| > joint DP-SGD mechanism and must not be released. Computing a separate CDF | ||
| > query outside this mechanism would require its own privacy analysis and | ||
| > accounting. | ||
|
|
||
| ### Selecting the slack dimension K | ||
|
|
||
| When `num_slots=None` (the default), the optimizer selects `K` using the | ||
| 99%-confidence SNR rule in equation (36): | ||
|
|
||
| ```text | ||
| K_max = (B / (2 * 2.576 * sigma))^(2/3) | ||
| K = max(1, floor(K_max)). | ||
| ``` | ||
|
|
||
| Here, `B` is the fixed normalization constant used by Opacus | ||
| (`expected_batch_size`), not a realized Poisson batch size or a physical | ||
| microbatch size. Table 3 gives the following illustrative practical choices | ||
| for representative batch sizes when `sigma=1`: | ||
|
|
||
| | B | 128 | 256 | 512 | 1024 | 2048 | | ||
| |---:|---:|---:|---:|---:|---:| | ||
| | Illustrative practical K | 8 | 10 | 20 | 30 | 50 | | ||
|
|
||
| These rounded values show the approximate scale of `K`; they are not treated | ||
| as special cases by the implementation. The automatic rule always evaluates | ||
| equation (36) directly, giving `K = {8, 13, 21, 34, 54}` for the batch sizes | ||
| above when `sigma=1`. Pass a positive integer as `num_slots` to reproduce a | ||
| particular practical choice or run an ablation. The selected value is | ||
| available as `optimizer.K`. | ||
|
|
||
| ## Usage | ||
|
|
||
| The prototype supports two composable steps: | ||
|
|
||
| 1. `SlaClipDPOptimizer` jointly releases the DP gradient and private Slack | ||
| Indicator. With no controller, the clipping threshold remains unchanged. | ||
| 2. `SlaClipController` consumes the released indicator and applies equations | ||
| (28)-(30) to update the threshold. Since this is post-processing of a DP | ||
| release, it adds no privacy cost. | ||
|
|
||
| ### Prepare private training | ||
|
|
||
| `SlaClipPrivacyEngine` is the recommended entry point. It uses Opacus's native | ||
| model wrapping, data loader, secure RNG, and privacy accountant while selecting | ||
| `SlaClipDPOptimizer` automatically: | ||
|
|
||
| ```python | ||
| from research.slaclip import SlaClipPrivacyEngine | ||
|
|
||
|
|
||
| privacy_engine = SlaClipPrivacyEngine() | ||
| model, optimizer, train_loader = privacy_engine.make_private( | ||
| module=model, | ||
| optimizer=optimizer, | ||
| data_loader=train_loader, | ||
| noise_multiplier=1.0, | ||
| max_grad_norm=1.0, | ||
| ) | ||
| ``` | ||
|
|
||
| This entry point also supports the inherited | ||
| `PrivacyEngine.make_private_with_epsilon()` API. It intentionally rejects | ||
| distributed training, non-flat clipping, and ghost clipping, which the current | ||
| research optimizer does not support. | ||
|
|
||
| ### Step 1 only: obtain private CDF information | ||
|
|
||
| The default `SlaClipPrivacyEngine()` has no controller, so it keeps `C` fixed | ||
| while obtaining the noisy Slack Indicator: | ||
|
|
||
| ```python | ||
| for images, targets in train_loader: | ||
| optimizer.zero_grad() | ||
| loss = criterion(model(images), targets) | ||
| loss.backward() | ||
| optimizer.step() | ||
|
|
||
| private_cdf = optimizer.slack_indicator | ||
| # Use private_cdf only in DP-safe post-processing. | ||
| ``` | ||
|
|
||
| Because Gaussian noise is unbounded, individual coordinates can fall outside | ||
| `[0, 1]` or fail to be monotone. A downstream method may project or smooth the | ||
| released vector as post-processing without additional privacy cost. | ||
|
|
||
| ### Steps 1 and 2: paper SlaClip | ||
|
|
||
| Construct the entry point with the paper controller to adapt `C` after every | ||
| release, then call `make_private()` as above: | ||
|
|
||
| ```python | ||
| from research.slaclip import SlaClipController, SlaClipPrivacyEngine | ||
|
|
||
|
|
||
| privacy_engine = SlaClipPrivacyEngine( | ||
| clipping_controller=SlaClipController( | ||
| eta=0.5, | ||
| min_clipbound=0.1, | ||
| max_clipbound=50.0, | ||
| ), | ||
| ) | ||
| model, optimizer, train_loader = privacy_engine.make_private( | ||
| module=model, | ||
| optimizer=optimizer, | ||
| data_loader=train_loader, | ||
| noise_multiplier=1.0, | ||
| max_grad_norm=1.0, | ||
| ) | ||
|
|
||
| for images, targets in train_loader: | ||
| optimizer.zero_grad() | ||
| loss = criterion(model(images), targets) | ||
| loss.backward() | ||
| optimizer.step() | ||
|
|
||
| print(optimizer.current_clip) | ||
| ``` | ||
|
|
||
| The controller implements | ||
|
|
||
| ```text | ||
| r_t = clip_[0,1](slack_indicator[K] / C_t) | ||
| gamma_t = clip_[0,1](1 - (1 - r_t) / 2) | ||
| C_(t+1) = clip_[C_min,C_max]( | ||
| C_t * exp(eta * (gamma_t - slack_indicator[1])) | ||
| ). | ||
| ``` | ||
|
|
||
| The released indicator contains unbounded Gaussian noise. The controller | ||
| therefore projects `r_t` and `gamma_t` onto `[0, 1]`, as specified in the | ||
| paper, and bounds the next positive clipping threshold to | ||
| `[min_clipbound, max_clipbound]`. It intentionally does not force | ||
| `gamma_t - slack_indicator[1]` to be nonnegative: a negative value is the | ||
| feedback that decreases an overly large clipping threshold. | ||
|
|
||
| A custom callable with signature `(current_clip, slack_indicator) -> next_clip` | ||
| can replace `SlaClipController`. Such a method reuses the SlaClip Slack | ||
| Indicator but is not the paper's threshold controller and should be described | ||
| accordingly. | ||
|
|
||
| ### Parameters | ||
|
|
||
| - `num_slots`: number `K` of CDF bins. `None` automatically selects it from | ||
| equation (36). A positive integer overrides the automatic value. | ||
| Larger values increase resolution but also increase normalized indicator | ||
| noise. Pass it to `SlaClipPrivacyEngine`. | ||
| - `clipping_controller`: optional post-processing callable. `None` enables the | ||
| indicator-only mode. Pass it to `SlaClipPrivacyEngine`. | ||
| - `eta`: positive multiplicative update step size used by | ||
| `SlaClipController`. | ||
| - `min_clipbound`, `max_clipbound`: positive lower and upper bounds for the | ||
| clipping threshold. Their defaults, `0.1` and `50.0`, match Opacus | ||
| `AdaClipDPOptimizer` and the SlaClip experiment configuration. | ||
| - All other optimizer arguments have the same meaning as in Opacus | ||
| `DPOptimizer`. | ||
|
|
||
| ## Limitations | ||
|
|
||
| - `SlaClipPrivacyEngine` supports the standard, non-distributed | ||
| `DPOptimizer` path. SlaClip is not registered as a clipping mode on the core | ||
| Opacus `PrivacyEngine`. | ||
| - Use it from the repository root through `research.slaclip`; research modules | ||
| are not part of the installed Opacus public API. | ||
| - The privacy argument assumes the same sampling rule, normalization constant, | ||
| clipping threshold, noise multiplier, and accountant parameters for the | ||
| gradient and slack parts of the joint release. | ||
| - As research code, it is not covered by Opacus public API compatibility | ||
| guarantees. | ||
|
|
||
| ## Tests | ||
|
|
||
| From the repository root: | ||
|
|
||
| ```bash | ||
| python -m pytest research/slaclip -q | ||
| ``` | ||
|
|
||
| The tests cover the `SlaClipPrivacyEngine` entry point and accountant hook, | ||
| automatic `K` selection, equations (7)-(8), the extended-gradient norm bound, | ||
| exact agreement of the first `d` coordinates with native `DPOptimizer`, | ||
| indicator-only operation, the paper controller, and empty Poisson batches. | ||
|
|
||
| ## Citation | ||
|
|
||
| ```bibtex | ||
| @inproceedings{zou2026slaclip, | ||
| title={{SlaClip}: Gradient Norm Slacks Can Be an Indicator for Adaptive | ||
| Clipping in {DP-SGD}}, | ||
| author={Zou, Shuyan and Wang, Shaowei and Zhu, Zhanxing and Li, Jin and | ||
| Dong, Changyu and Sassone, Vladimiro and Wu, Han}, | ||
| booktitle={Proceedings of the 43rd International Conference on Machine | ||
| Learning}, | ||
| year={2026} | ||
| } | ||
| ``` | ||
|
|
||
| The authors' reference implementation is available at | ||
| [ZsyRock/SlaClip](https://github.com/ZsyRock/SlaClip). |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,18 @@ | ||
| # Copyright (c) 2026 SlaClip authors. | ||
| # | ||
| # Licensed under the Apache License, Version 2.0 (the "License"); | ||
| # you may not use this file except in compliance with the License. | ||
| # You may obtain a copy of the License at | ||
| # | ||
| # http://www.apache.org/licenses/LICENSE-2.0 | ||
| # | ||
| # Unless required by applicable law or agreed to in writing, software | ||
| # distributed under the License is distributed on an "AS IS" BASIS, | ||
| # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| # See the License for the specific language governing permissions and | ||
| # limitations under the License. | ||
|
|
||
| """Research package for the SlaClip prototype.""" | ||
|
|
||
| from .privacy_engine import SlaClipPrivacyEngine # noqa: F401 | ||
| from .slaclipoptimizer import SlaClipController, SlaClipDPOptimizer # noqa: F401 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,92 @@ | ||
| #!/usr/bin/env python3 | ||
| # Copyright (c) 2026 SlaClip authors. | ||
| # | ||
| # Licensed under the Apache License, Version 2.0 (the "License"); | ||
| # you may not use this file except in compliance with the License. | ||
| # You may obtain a copy of the License at | ||
| # | ||
| # http://www.apache.org/licenses/LICENSE-2.0 | ||
| # | ||
| # Unless required by applicable law or agreed to in writing, software | ||
| # distributed under the License is distributed on an "AS IS" BASIS, | ||
| # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. | ||
| # See the License for the specific language governing permissions and | ||
| # limitations under the License. | ||
|
|
||
| """PrivacyEngine entry point for the SlaClip research prototype.""" | ||
|
|
||
| from __future__ import annotations | ||
|
|
||
| from typing import Callable, List, Optional, Union | ||
|
|
||
| import torch | ||
| from opacus.optimizers import DPOptimizer | ||
| from opacus.privacy_engine import PrivacyEngine | ||
| from torch import optim | ||
|
|
||
| from .slaclipoptimizer import SlaClipDPOptimizer | ||
|
|
||
|
|
||
| class SlaClipPrivacyEngine(PrivacyEngine): | ||
| """Prepare Opacus training objects with :class:`SlaClipDPOptimizer`. | ||
|
|
||
| Passing no controller enables the indicator-only mode. Passing the paper's | ||
| ``SlaClipController`` enables the complete SlaClip adaptation rule. | ||
| """ | ||
|
|
||
| def __init__( | ||
| self, | ||
| *, | ||
| accountant: str = "prv", | ||
| secure_mode: bool = False, | ||
| num_slots: Optional[int] = None, | ||
| clipping_controller: Optional[Callable[[float, torch.Tensor], float]] = None, | ||
| ): | ||
| super().__init__(accountant=accountant, secure_mode=secure_mode) | ||
| self.num_slots = num_slots | ||
| self.clipping_controller = clipping_controller | ||
|
|
||
| def _prepare_optimizer( | ||
| self, | ||
| *, | ||
| optimizer: optim.Optimizer, | ||
| noise_multiplier: float, | ||
| max_grad_norm: Union[float, List[float]], | ||
| expected_batch_size: int, | ||
| loss_reduction: str = "mean", | ||
| distributed: bool = False, | ||
| clipping: str = "flat", | ||
| noise_generator=None, | ||
| grad_sample_mode: str = "hooks", | ||
| **kwargs, | ||
| ) -> SlaClipDPOptimizer: | ||
| if distributed: | ||
| raise ValueError("SlaClip does not currently support distributed training") | ||
| if clipping != "flat": | ||
| raise ValueError("SlaClip requires clipping='flat'") | ||
| if "ghost" in grad_sample_mode: | ||
| raise ValueError("SlaClip does not currently support ghost clipping") | ||
| if isinstance(max_grad_norm, list): | ||
| raise ValueError("SlaClip requires a scalar max_grad_norm") | ||
|
|
||
| if isinstance(optimizer, DPOptimizer): | ||
| optimizer = optimizer.original_optimizer | ||
|
|
||
| generator = None | ||
| if self.secure_mode: | ||
| generator = self.secure_rng | ||
| elif noise_generator is not None: | ||
| generator = noise_generator | ||
|
|
||
| return SlaClipDPOptimizer( | ||
| optimizer=optimizer, | ||
| noise_multiplier=noise_multiplier, | ||
| max_grad_norm=float(max_grad_norm), | ||
| expected_batch_size=expected_batch_size, | ||
| loss_reduction=loss_reduction, | ||
| generator=generator, | ||
| secure_mode=self.secure_mode, | ||
| num_slots=self.num_slots, | ||
| clipping_controller=self.clipping_controller, | ||
| **kwargs, | ||
| ) |
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.