diff --git a/deployments/apple_iconic_scenes_photo_selections_2023.yaml b/deployments/apple_iconic_scenes_photo_selections_2023.yaml index 3ab7b15..70fc4f1 100644 --- a/deployments/apple_iconic_scenes_photo_selections_2023.yaml +++ b/deployments/apple_iconic_scenes_photo_selections_2023.yaml @@ -56,7 +56,7 @@ deployment: - Pre-defined privacy/utility targets: Apple described in its blog post that "We are limited to fixed precision to provide a consistent privacy assurance with our current approach, which can’t be optimal for all locations". - Apple is committed to balancing privacy and utility: "We combined local noise addition with a technique called secure aggregation to address these concerns of balancing privacy with utility" - resources: + administrative: sources: | - Blog post: https://machinelearning.apple.com/research/scenes-differential-privacy # notes: '' # other extra information diff --git a/deployments/apple_popular_emojis.yaml b/deployments/apple_popular_emojis.yaml index ab64528..87fc44c 100644 --- a/deployments/apple_popular_emojis.yaml +++ b/deployments/apple_popular_emojis.yaml @@ -76,7 +76,7 @@ deployment: - The choice of \\(\epsilon\\) was “based on the privacy characteristics of the underlying dataset” and is “consistent with the parameters proposed in the differential privacy research community”. - The CMS algorithm yields a high number of hash collisions (mapping 2,600 emojis to 1024 bits) providing further plausible deniability. - resources: + administrative: sources: | - Paper: https://docs-assets.developer.apple.com/ml-research/papers/learning-with-privacy-at-scale.pdf notes: | diff --git a/deployments/assistive_ai.yaml b/deployments/assistive_ai.yaml index b080d7a..195b7c7 100644 --- a/deployments/assistive_ai.yaml +++ b/deployments/assistive_ai.yaml @@ -30,7 +30,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Differentially Private Set Union justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - Blog post: https://www.microsoft.com/en-us/research/articles/assistive-ai-makes-replying-easier-2/ - Paper: https://arxiv.org/pdf/2002.09745 diff --git a/deployments/audience_engagement.yaml b/deployments/audience_engagement.yaml index 52a03c0..d4d6d55 100644 --- a/deployments/audience_engagement.yaml +++ b/deployments/audience_engagement.yaml @@ -34,7 +34,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Laplace, Gumbel justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - Paper: https://arxiv.org/pdf/2002.05839 notes: | diff --git a/deployments/autoplay_intent.yaml b/deployments/autoplay_intent.yaml index 8e5376b..5c1b981 100644 --- a/deployments/autoplay_intent.yaml +++ b/deployments/autoplay_intent.yaml @@ -32,7 +32,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Private Count Mean Sketch justification: '' # TODO: Fill in correct value - resources: + administrative: sources: https://docs-assets.developer.apple.com/ml-research/papers/learning-with-privacy-at-scale.pdf registry_authors: - Nicolas Berrios diff --git a/deployments/birth_dataset.yaml b/deployments/birth_dataset.yaml index 2ee9b6d..33fda20 100644 --- a/deployments/birth_dataset.yaml +++ b/deployments/birth_dataset.yaml @@ -34,7 +34,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: PrivBayes with Private Selection (Universal Microdata Scheme) justification: '' # TODO: Fill in correct value - resources: + administrative: sources: https://birth.dataset.pub registry_authors: - Shlomi Hod diff --git a/deployments/broadband_coverage.yaml b/deployments/broadband_coverage.yaml index 510bea0..467be19 100644 --- a/deployments/broadband_coverage.yaml +++ b/deployments/broadband_coverage.yaml @@ -77,7 +77,7 @@ deployment: - The number of devices connected to the internet at broadband speed per each zip code is counted based on the Federal Communications Commission (FCC)’s definition of broadband that is 25mbps per download. - "All differential privacy processing was done with the OpenDP SmartNoise library. The SmartNoise library includes a comprehensive set of differential privacy mechanisms, algorithms, and validator. The library is open source, and is maintained and vetted by OpenDP." - A Monte Carlo simulation process to estimate the error introduced by DP was chosen, the rationale being that this empirical method is better for estimating the combined error from several private sources and does not result in additional privacy losses. - resources: + administrative: sources: | - Paper: https://arxiv.org/abs/2103.14035 - Github: https://github.com/microsoft/USBroadbandUsagePercentages diff --git a/deployments/census_demographic_and_housing.yaml b/deployments/census_demographic_and_housing.yaml index 6a4e14b..9b380c7 100644 --- a/deployments/census_demographic_and_housing.yaml +++ b/deployments/census_demographic_and_housing.yaml @@ -101,7 +101,7 @@ deployment: - The Census Bureau solicited and incorporated detailed feedback from a wide array of stakeholders, including the National Academy of Sciences, federal and state partners, academic researchers, and tribal leaders. - The final production parameters were set by the Data Stewardship Executive Policy Committee (DSEP) after a final review of the privacy guarantees and data accuracy demonstrated in the tuning experiments. - resources: + administrative: sources: | - 2020 Census Demographic and Housing Characteristics File (DHC) Technical Documentation: https://www2.census.gov/programs-surveys/decennial/2020/technical-documentation/complete-tech-docs/demographic-and-housing-characteristics-file-and-demographic-profile/2020census-demographic-and-housing-characteristics-file-and-demographic-profile-techdoc.pdf - Census Bureau Releases New 2020 Census Data on Age, Sex, Race, Hispanic Origin, Households and Housing: https://www.census.gov/newsroom/press-releases/2023/2020-census-demographic-profile-and-dhc.html diff --git a/deployments/census_detailed_demographic_and_housing_a.yaml b/deployments/census_detailed_demographic_and_housing_a.yaml index 9f14d43..9f1b523 100644 --- a/deployments/census_detailed_demographic_and_housing_a.yaml +++ b/deployments/census_detailed_demographic_and_housing_a.yaml @@ -87,7 +87,7 @@ deployment: - The Census Bureau solicited and incorporated detailed feedback from a wide array of stakeholders, including the National Academy of Sciences, federal and state partners, academic researchers, and tribal leaders. - The final production parameters were set by the Data Stewardship Executive Policy Committee (DSEP) after a final review of the privacy guarantees and data accuracy demonstrated in the tuning experiments. - resources: + administrative: sources: | - 2020 Census Detailed Demographic and Housing Characteristics File A (Detailed DHC-A) Technical Documentation: https://www2.census.gov/programs-surveys/decennial/2020/technical-documentation/complete-tech-docs/detailed-demographic-and-housing-characteristics-file-a/2020census-detailed-dhc-a-techdoc.pdf - Census Bureau Announces Release Date for 2020 Census Data Product on Race and Ethnicity: https://www.census.gov/newsroom/press-releases/2023/2020-census-detailed-dhc-a.html diff --git a/deployments/census_detailed_demographic_and_housing_b.yaml b/deployments/census_detailed_demographic_and_housing_b.yaml index 735cc96..9c51605 100644 --- a/deployments/census_detailed_demographic_and_housing_b.yaml +++ b/deployments/census_detailed_demographic_and_housing_b.yaml @@ -87,7 +87,7 @@ deployment: - The Census Bureau solicited and incorporated detailed feedback from a wide array of stakeholders, including the National Academy of Sciences, federal and state partners, academic researchers, and tribal leaders. - The final production parameters were set by the Data Stewardship Executive Policy Committee (DSEP) after a final review of the privacy guarantees and data accuracy demonstrated in the tuning experiments. - resources: + administrative: sources: | - 2020 Census Detailed Demographic and Housing Characteristics File B (Detailed DHC-B) Technical Documentation: https://www2.census.gov/programs-surveys/decennial/2020/technical-documentation/complete-tech-docs/detailed-demographic-and-housing-characteristics-file-b/2020census-detailed-dhc-b-techdoc.pdf - Census Bureau to Hold Webinar on Release of 2020 Census Detailed Demographic and Housing Characteristics File B: https://www.census.gov/newsroom/press-releases/2024/webinar-2020-census-detailed-dhc-b.html diff --git a/deployments/census_redistricting_data.yaml b/deployments/census_redistricting_data.yaml index a8864cd..f88080f 100644 --- a/deployments/census_redistricting_data.yaml +++ b/deployments/census_redistricting_data.yaml @@ -101,7 +101,7 @@ deployment: - The Census Bureau solicited and incorporated detailed feedback from a wide array of stakeholders, including the National Academy of Sciences, federal and state partners, academic researchers, and tribal leaders. - The final production parameters were set by the Data Stewardship Executive Policy Committee (DSEP) after a final review of the privacy guarantees and data accuracy demonstrated in the tuning experiments. - resources: + administrative: sources: | - Disclosure avoidance handbook for the 2020 Census: https://www2.census.gov/library/publications/decennial/2020/2020-census-disclosure-avoidance-handbook.pdf - The 2020 Census Disclosure Avoidance System TopDown Algorithm: https://arxiv.org/abs/2204.08986 diff --git a/deployments/census_supplemental_demographic_and_housing.yaml b/deployments/census_supplemental_demographic_and_housing.yaml index ed4fcea..3ded7c3 100644 --- a/deployments/census_supplemental_demographic_and_housing.yaml +++ b/deployments/census_supplemental_demographic_and_housing.yaml @@ -76,7 +76,7 @@ deployment: - The Census Bureau solicited and incorporated detailed feedback from a wide array of stakeholders, including the National Academy of Sciences, federal and state partners, academic researchers, and tribal leaders. - The final production parameters were set by the Data Stewardship Executive Policy Committee (DSEP) after a final review of the privacy guarantees and data accuracy demonstrated in the tuning experiments. - resources: + administrative: sources: | - 2020 Supplemental Demographic and Housing Characteristics File (S-DHC) Technical Documentation: https://www2.census.gov/programs-surveys/decennial/2020/technical-documentation/complete-tech-docs/supplemental-demographic-and-housing-characteristics-file/2020census-supplemental-dhc-techdoc.pdf - Census Bureau Releases Final 2020 Census Data Product: https://www.census.gov/newsroom/press-releases/2024/final-2020-census-data-product-s-dhc.html diff --git a/deployments/county_business_patterns.yaml b/deployments/county_business_patterns.yaml index ff68345..fd57ad0 100644 --- a/deployments/county_business_patterns.yaml +++ b/deployments/county_business_patterns.yaml @@ -32,7 +32,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Gaussian justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - https://www.census.gov/topics/business-economy/disclosure/about.html - https://www.census.gov/data/academy/webinars/2023/differential-privacy-webinar.html diff --git a/deployments/covid19_notifications.yaml b/deployments/covid19_notifications.yaml index 463708c..c0334ee 100644 --- a/deployments/covid19_notifications.yaml +++ b/deployments/covid19_notifications.yaml @@ -34,7 +34,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Randomized Response justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - Paper: https://covid19-static.cdn-apple.com/applications/covid19/current/static/contact-tracing/pdf/ENPA_White_Paper.pdf registry_authors: diff --git a/deployments/covid19_search_trends_symptoms.yaml b/deployments/covid19_search_trends_symptoms.yaml index 6db333b..c58dc7c 100644 --- a/deployments/covid19_search_trends_symptoms.yaml +++ b/deployments/covid19_search_trends_symptoms.yaml @@ -88,7 +88,7 @@ deployment: - The decision to bound a user’s contribution to a maximum of three different symptom counts per day was based on the observation that approximately 75% of users search for three or fewer symptoms daily. - For each geographic region and symptom, data is released as either daily or weekly aggregates, depending on data quality. While the system provides daily aggregates whenever possible, it switches to weekly aggregates if privacy protections significantly affect the data's accuracy. Weekly aggregates are more robust because they are based on more data, which reduces the relative error from the added privacy noise. This choice was determined once in a differentially private manner using data from February to July 2020. After being set, the temporal granularity for a given region and symptom remains fixed for the entire duration of the dataset's release. - resources: + administrative: sources: | - Documentation: https://storage.googleapis.com/gcp-public-data-symptom-search/COVID-19%20Search%20Trends%20symptoms%20dataset%20documentation%20.pdf?utm_source=chatgpt.com - Paper: https://arxiv.org/pdf/2009.01265 diff --git a/deployments/employment_outcomes.yaml b/deployments/employment_outcomes.yaml index 00ac92d..778138d 100644 --- a/deployments/employment_outcomes.yaml +++ b/deployments/employment_outcomes.yaml @@ -47,7 +47,7 @@ deployment: # pre_processing_eda_hyperparameter_tuning: # mechanisms: # justification: - resources: + administrative: sources: | - Paper: https://journalprivacyconfidentiality.org/index.php/jpc/article/view/722/684 - Documentation: https://lehd.ces.census.gov/data/pseo_documentation.html diff --git a/deployments/google_mobility.yaml b/deployments/google_mobility.yaml index 285a912..4a8c8e9 100644 --- a/deployments/google_mobility.yaml +++ b/deployments/google_mobility.yaml @@ -95,7 +95,7 @@ deployment: - Statistically insignificant metrics are filtered out to ensure utility and reliability of the data. This is done through a strict confidence threshold, ensuring there is at most a 5% risk of any published value being wrong by more than 10 absolute percentage points. - The implementation decision to bound each user's contribution to a maximum of four (category, location) pairs per day is justified by an analysis of user behavior. This threshold was chosen because it "does not significantly affect data accuracy," given that 99% of users at the U.S. county level contribute to three or fewer pairs on average. - resources: + administrative: sources: https://arxiv.org/pdf/2004.04145 registry_authors: - Elena Ghazi diff --git a/deployments/healthkit.yaml b/deployments/healthkit.yaml index 2a121a5..b99a517 100644 --- a/deployments/healthkit.yaml +++ b/deployments/healthkit.yaml @@ -35,7 +35,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Private Count Mean Sketch justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - Paper: https://docs-assets.developer.apple.com/ml-research/papers/learning-with-privacy-at-scale.pdf - Additional version: https://machinelearning.apple.com/research/learning-with-privacy-at-scale diff --git a/deployments/korean_statistics_datahub.yaml b/deployments/korean_statistics_datahub.yaml index 83b43d1..04d9564 100644 --- a/deployments/korean_statistics_datahub.yaml +++ b/deployments/korean_statistics_datahub.yaml @@ -30,7 +30,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Unknown justification: '' # TODO: Fill in correct value - resources: + administrative: sources: https://unstats.un.org/wiki/display/UGTTOPPT/11.+Statistics+Korea%3A+Developing+a+privacy-preserving+Statistical+Data+Hub+Platform registry_authors: - Nicolas Berrios diff --git a/deployments/linkedin_hiring_reports.yaml b/deployments/linkedin_hiring_reports.yaml index 0d11da7..aedbbd6 100644 --- a/deployments/linkedin_hiring_reports.yaml +++ b/deployments/linkedin_hiring_reports.yaml @@ -38,7 +38,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Laplace, Gumbel justification: '' # TODO: Fill in correct value - resources: + administrative: sources: https://arxiv.org/abs/2010.13981 registry_authors: - Nicolas Berrios diff --git a/deployments/lookup_hints.yaml b/deployments/lookup_hints.yaml index 3cf739e..1e8f826 100644 --- a/deployments/lookup_hints.yaml +++ b/deployments/lookup_hints.yaml @@ -35,7 +35,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Count Mean Sketch or Hadamard Count Mean Sketch justification: '' # TODO: Fill in correct value - resources: + administrative: sources: https://www.apple.com/privacy/docs/Differential_Privacy_Overview.pdf registry_authors: - Nicolas Berrios diff --git a/deployments/microsoft_ctdc_victim_perpertartor_synthetic_dataset_2023.yaml b/deployments/microsoft_ctdc_victim_perpertartor_synthetic_dataset_2023.yaml index 23252ae..4b8a0cf 100644 --- a/deployments/microsoft_ctdc_victim_perpertartor_synthetic_dataset_2023.yaml +++ b/deployments/microsoft_ctdc_victim_perpertartor_synthetic_dataset_2023.yaml @@ -79,7 +79,7 @@ deployment: - In the description, the CTDC states: "The technology has enabled CTDC to share more data and conduct more robust research while protecting privacy and civil liberties". - Also, the CTDC states: "Both datasets preserve privacy by design". - resources: + administrative: sources: | - Microsoft Research Blog: https://www.microsoft.com/en-us/research/blog/iom-and-microsoft-release-first-ever-differentially-private-synthetic-dataset-to-counter-human-trafficking/ - Dataset: https://www.ctdatacollaborative.org/global-victim-perpetrator-synthetic-dataset#no-back diff --git a/deployments/mobility_trends_hurricane.yaml b/deployments/mobility_trends_hurricane.yaml index cfa4ba1..5956087 100644 --- a/deployments/mobility_trends_hurricane.yaml +++ b/deployments/mobility_trends_hurricane.yaml @@ -32,7 +32,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Laplace justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - https://spectus.ai/wp-content/uploads/2022/10/Spectus_DPWhitepaper_v01b.pdf - https://web.archive.org/web/20221209065442/https://spectus.ai/wp-content/uploads/2022/10/Spectus_DPWhitepaper_v01b.pdf diff --git a/deployments/movement_ranges_maps.yaml b/deployments/movement_ranges_maps.yaml index d7410ac..f35ee0d 100644 --- a/deployments/movement_ranges_maps.yaml +++ b/deployments/movement_ranges_maps.yaml @@ -31,7 +31,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Laplace justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - Blog post: https://research.facebook.com/blog/2020/06/protecting-privacy-in-facebook-mobility-data-during-the-covid-19-response/ - Downloadable data product: https://data.humdata.org/dataset/movement-range-maps diff --git a/deployments/on_device_browser_rec.yaml b/deployments/on_device_browser_rec.yaml index 651f24a..5d0384b 100644 --- a/deployments/on_device_browser_rec.yaml +++ b/deployments/on_device_browser_rec.yaml @@ -28,7 +28,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: DP SGD* justification: '' # TODO: Fill in correct value - resources: + administrative: sources: https://brave.com/blog/federated-learning/ registry_authors: - Nicolas Berrios diff --git a/deployments/private_third_party_audits.yaml b/deployments/private_third_party_audits.yaml index 57bbaf6..d610808 100644 --- a/deployments/private_third_party_audits.yaml +++ b/deployments/private_third_party_audits.yaml @@ -33,7 +33,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Not Specificed, Synthetic Data justification: '' # TODO: Fill in correct value - resources: + administrative: sources: https://unstats.un.org/wiki/display/UGTTOPPT/14.+Twitter+and+OpenMined%3A+Enabling+Third-party+Audits+and+Research+Reproducibility+over+Unreleased+Digital+Assets registry_authors: - Nicolas Berrios diff --git a/deployments/recurve_energy_dp.yaml b/deployments/recurve_energy_dp.yaml index fea9491..921c430 100644 --- a/deployments/recurve_energy_dp.yaml +++ b/deployments/recurve_energy_dp.yaml @@ -101,7 +101,7 @@ deployment: - Google's differential privacy Go library was used to mitigate floating point attacks. - An 'isolation attack' was simulated to demonstrate that the noise added by a Gaussian mechanism with \\(\epsilon=0.843\\) 'limits an attacker to a confidence interval of around ±1,000 kWh for hourly consumption.' - resources: + administrative: sources: | - Applying Energy Differential Privacy To Enable Measurement of the OhmConnect Virtual Power Plant (paper): https://assets.website-files.com/5cb0a177570549b5f11b9550/5ffddb83b5ea5d67f5c43661_Quantifying%20The%20OhmConnect%20Virtual%20Power%20Plant%20During%20the%20California%20Blackouts.pdf - Revenue-Grade Analysis of the OhmConnect Virtual Power Plant During the California Blackouts (blog): https://www.recurve.com/blog/revenue-grade-analysis-of-the-ohmconnect-virtual-power-plant-during-the-california-blackouts diff --git a/deployments/safari_energy_drain.yaml b/deployments/safari_energy_drain.yaml index d578d5d..9c74ea0 100644 --- a/deployments/safari_energy_drain.yaml +++ b/deployments/safari_energy_drain.yaml @@ -49,7 +49,7 @@ deployment: mechanisms: Private Hadamard Count Mean Sketch (PHCMS) justification: '"Formal proof, described in Theorem 4.1, Theorem 4.3"' - resources: + administrative: sources: https://docs-assets.developer.apple.com/ml-research/papers/learning-with-privacy-at-scale.pdf notes: | - There are multiple releases per device & per user -- statistics are reported automatically on IoS and MacOS devices, subject to daily aggregation. diff --git a/deployments/safety_classifier.yaml b/deployments/safety_classifier.yaml index 3764d68..e304d58 100644 --- a/deployments/safety_classifier.yaml +++ b/deployments/safety_classifier.yaml @@ -45,7 +45,7 @@ deployment: # pre_processing_eda_hyperparameter_tuning: # mechanisms: # justification: - resources: + administrative: sources: | - Blog post: https://research.google/blog/protecting-users-with-differentially-private-synthetic-training-data/ - Research paper: https://arxiv.org/pdf/2306.01684 diff --git a/deployments/sas_data_maker_vulnerable_persons.yaml b/deployments/sas_data_maker_vulnerable_persons.yaml index 809a8db..65de049 100644 --- a/deployments/sas_data_maker_vulnerable_persons.yaml +++ b/deployments/sas_data_maker_vulnerable_persons.yaml @@ -65,7 +65,7 @@ deployment: - PrivBayes network is chosen to model the static tables (Customers and Accounts) because it offers a strong balance between utility, privacy, computational efficiency, scalability, and interpretability. - Autoregressive model is chosen to model the Transactions table because it handles time-series data effectively. However, DP is not applied to this model since DP-SGD requires adding too much noise, reducing utility to an unacceptably low level. - resources: + administrative: sources: | - Usecase by ICO: https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/privacy-enhancing-technologies/case-studies/synthetic-data-to-test-the-effectiveness-of-a-vulnerable-persons-detection-system-in-financial-services/ - Usecase by Nationwide: https://medium.com/nationwide-technology/hazy-synthetic-data-to-fuel-rapid-innovation-fd24f2e21685 diff --git a/deployments/shared_mobility_dataset.yaml b/deployments/shared_mobility_dataset.yaml index a3d0aeb..dc7d4b7 100644 --- a/deployments/shared_mobility_dataset.yaml +++ b/deployments/shared_mobility_dataset.yaml @@ -37,7 +37,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Laplace Mechanism justification: '"All trips were anonymized and aggregated by jointly applying differential privacy via the Laplace mechanism in combination with k-anonymity."' - resources: + administrative: sources: https://www.nature.com/articles/s41467-019-12809-y registry_authors: - Nicolas Berrios diff --git a/deployments/spanish_language_next_word.yaml b/deployments/spanish_language_next_word.yaml index 5174d0b..6aea865 100644 --- a/deployments/spanish_language_next_word.yaml +++ b/deployments/spanish_language_next_word.yaml @@ -67,7 +67,7 @@ deployment: - DP-FedAvg algorithm's guarantee depended on amplification-via-sampling, which was deemed insufficient because ensuring that devices are subsampled precisely and uniformly at random from a large population would be complex and hard to verify in a real-world system where device availability fluctuates based on external factors (e.g., being idle, on Wi-Fi, and charging). - DP-FTRL was chosen to address this challenge, given the observation that training convergence depends on accuracy of cumulative sums of gradients rather than individual ones, and that it is possible to provide accurate estimates of cumulative sums with a strong DP guarantee by using negatively correlated noise, since some of the privacy noise cancels out from step to step, allowing the model's learning trajectory to stay closer to the true gradient descent steps and achieve better accuracy for a given level of privacy. - resources: + administrative: sources: | - Federated Learning with Formal Differential Privacy Guarantees (Article published on February 28, 2022): https://research.google/blog/federated-learning-with-formal-differential-privacy-guarantees/ - Paper that introduces the DP variant of Follow-The-Regularized-Leader (DP-FTRL) used in the deployment: Practical and Private (Deep) Learning Without Sampling or Shuffling: https://arxiv.org/pdf/2103.00039 diff --git a/deployments/synth_data_public_use.yaml b/deployments/synth_data_public_use.yaml index c99e356..415d3f9 100644 --- a/deployments/synth_data_public_use.yaml +++ b/deployments/synth_data_public_use.yaml @@ -29,7 +29,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: DP GAN* justification: '' # TODO: Fill in correct value - resources: + administrative: sources: https://datasciencecampus.ons.gov.uk/projects/generative-adversarial-networks-gans-for-synthetic-dataset-generation-with-binary-classes/ registry_authors: - Nicolas Berrios diff --git a/deployments/telemetry_collection_windows.yaml b/deployments/telemetry_collection_windows.yaml index c0db2b1..9678552 100644 --- a/deployments/telemetry_collection_windows.yaml +++ b/deployments/telemetry_collection_windows.yaml @@ -33,7 +33,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Laplace, Randomized Rounding justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - Paper: https://www.microsoft.com/en-us/research/publication/collecting-telemetry-data-privately/ registry_authors: diff --git a/deployments/uber.yaml b/deployments/uber.yaml index 90b441a..4ab4639 100644 --- a/deployments/uber.yaml +++ b/deployments/uber.yaml @@ -67,7 +67,7 @@ deployment: - Security measure to protect underlying data: none measured justification: | - Formal proof that elastic sensitivity is an upper bound of local sensistivity is presented. - resources: + administrative: sources: | - Paper: https://vldb.org/pvldb/vol11/p526-johnson.pdf # https://doi.org/10.1145/3177732.3177733 is not presently working - Blog post: https://medium.com/uber-security-privacy/uber-open-source-differential-privacy-57f31e85c57a diff --git a/deployments/user_url_privacy.yaml b/deployments/user_url_privacy.yaml index a662663..e22dba2 100644 --- a/deployments/user_url_privacy.yaml +++ b/deployments/user_url_privacy.yaml @@ -42,7 +42,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Laplace, Modified Report Noisy Max justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - Data: https://dataverse.harvard.edu/file.xhtml?persistentId=doi:10.7910/DVN/TDOAPG/DGSAMS&version=6.2 - FAQ: https://developers.facebook.com/docs/url-shares-dataset/support/faqs diff --git a/deployments/vaccine_search_insights.yaml b/deployments/vaccine_search_insights.yaml index 8a9e0e6..a95503d 100644 --- a/deployments/vaccine_search_insights.yaml +++ b/deployments/vaccine_search_insights.yaml @@ -34,7 +34,7 @@ deployment: pre_processing_eda_hyperparameter_tuning: '' # TODO: Fill in correct value mechanisms: Gaussian justification: '' # TODO: Fill in correct value - resources: + administrative: sources: | - Paper: https://arxiv.org/abs/2107.01179 registry_authors: diff --git a/deployments/wikimedia_current_usage_data.yaml b/deployments/wikimedia_current_usage_data.yaml index a1b1eab..5b487c0 100644 --- a/deployments/wikimedia_current_usage_data.yaml +++ b/deployments/wikimedia_current_usage_data.yaml @@ -69,7 +69,7 @@ deployment: - After comparison of available open-source tools, the Wikimedia Foundation decided to use Tumult Analytics, framework chosen for its robustness, production-readiness, compatibility with Wikimedia’s compute infrastructure, and support for advanced features like zCDP-based privacy accounting, and started a collaboration with Tumult Labs. - Choosing \\(\rho\\) =0.015 is considered to be conservative as it provides much stronger guarantee than \\((\epsilon,\delta)\\)-DP with \\(\epsilon=1\\) and \\(\delta = 10^{-7}\\) - resources: + administrative: sources: | - Paper: https://arxiv.org/pdf/2308.16298 - Several relevant links are provided in the paper’s references, among which: diff --git a/deployments/wikimedia_editor_activity_statistics_2023.yaml b/deployments/wikimedia_editor_activity_statistics_2023.yaml index a0bac81..ee064a9 100644 --- a/deployments/wikimedia_editor_activity_statistics_2023.yaml +++ b/deployments/wikimedia_editor_activity_statistics_2023.yaml @@ -63,7 +63,7 @@ deployment: - Pre-defined privacy/utility targets. - In their report, the Wikimedia Foundation states: "This dataset was optimized to perform well across several metrics: median relative error, relative error below 50% and relative error below 90%, spurious rate, and drop rate." - resources: + administrative: sources: | - Online description page: https://analytics.wikimedia.org/published/datasets/geoeditors_monthly/00_README.html - Code repository: https://gitlab.wikimedia.org/repos/security/differential-privacy/-/blob/main/differential_privacy/geoeditors_monthly.py diff --git a/deployments/wikimedia_historical_usage_data.yaml b/deployments/wikimedia_historical_usage_data.yaml index 1142678..17482e4 100644 --- a/deployments/wikimedia_historical_usage_data.yaml +++ b/deployments/wikimedia_historical_usage_data.yaml @@ -80,7 +80,7 @@ deployment: - The thresholds were chosen after extensive experimentation by measuring accuracy metrics (relative error distribution, drop rate, spurious rate) on the true data. Since these metrics could leak information, the team “kept fine-grained utility metrics confidential throughout the tuning process, minimizing data leakage”, and chose to “only publicly communicate approximate values of global utility metrics and the algorithmic parameters obtained from this tuning process”. - After an in-depth comparison of available open-source tools, the Wikimedia Foundation decided to use Tumult Analytics, framework chosen for its robustness, production-readiness, compatibility with Wikimedia’s compute infrastructure, and support for advanced features like zCDP-based privacy accounting, and started a collaboration with Tumult Labs. - resources: + administrative: sources: | - Paper: https://arxiv.org/pdf/2308.16298 - Several relevant links are provided in the paper’s references, among which: diff --git a/deployments/wikimedia_russian_editor_activity_statistics_2023.yaml b/deployments/wikimedia_russian_editor_activity_statistics_2023.yaml index c055289..7286fcc 100644 --- a/deployments/wikimedia_russian_editor_activity_statistics_2023.yaml +++ b/deployments/wikimedia_russian_editor_activity_statistics_2023.yaml @@ -62,7 +62,7 @@ deployment: - In their report, the Wikimedia Foundation states: "With a privacy loss of 0.1, they could at most be ~2.5% more certain of that account’s presence or absence." - In their report, they also state that "we have taken extra privacy precautions (specifically, using differential privacy) with this data release." - resources: + administrative: sources: | - Online description page: https://wikitech.wikimedia.org/wiki/Russian_editor_information_(2022-23) - Code or table building: https://gerrit.wikimedia.org/r/plugins/gitiles/analytics/refinery/%2B/refs/heads/master/hql/geoeditors/geoeditors_monthly.hql diff --git a/schemas/deployments-schema.yaml b/schemas/deployments-schema.yaml index 4a79365..42c7c05 100644 --- a/schemas/deployments-schema.yaml +++ b/schemas/deployments-schema.yaml @@ -369,7 +369,7 @@ properties: description_long: | Rationale for how decisions around implementation were made by the curator or others. In short, this information should answer the following questions laid out by Dwork Mulligan Kohli 2019: “What were the assumptions, modelling decisions, thresholds, and subjective decisions made in determining the implementation choices above? Why is the approach a thorough test of the stated assumptions? Was the process validated and verified? If so, how? type: string - resources: + administrative: type: object additionalProperties: False required: diff --git a/tests/bad_deployments/check_quoting_instead_of_use_to_preserve_whitespace_in_these_collapse_hurts_markdown_n.yaml b/tests/bad_deployments/check_quoting_instead_of_use_to_preserve_whitespace_in_these_collapse_hurts_markdown_n.yaml index a0d6d84..1611178 100644 --- a/tests/bad_deployments/check_quoting_instead_of_use_to_preserve_whitespace_in_these_collapse_hurts_markdown_n.yaml +++ b/tests/bad_deployments/check_quoting_instead_of_use_to_preserve_whitespace_in_these_collapse_hurts_markdown_n.yaml @@ -13,6 +13,6 @@ deployment: These collapse: hurts markdown! publication_date: "2026-01-01" - resources: + administrative: registry_authors: - ... diff --git a/tests/bad_deployments/check_schema_additional_properties_are_not_allowed_surprise_was_unexpected.yaml b/tests/bad_deployments/check_schema_additional_properties_are_not_allowed_surprise_was_unexpected.yaml index 52d3698..598df90 100644 --- a/tests/bad_deployments/check_schema_additional_properties_are_not_allowed_surprise_was_unexpected.yaml +++ b/tests/bad_deployments/check_schema_additional_properties_are_not_allowed_surprise_was_unexpected.yaml @@ -12,6 +12,6 @@ deployment: data_product_type: Summary statistics data_product_region: ... publication_date: "2026-01-01" - resources: + administrative: registry_authors: - ... diff --git a/tests/bad_deployments/check_schema_deployment_basic_publication_date_1_1_2026_is_not_a_date.yaml b/tests/bad_deployments/check_schema_deployment_basic_publication_date_1_1_2026_is_not_a_date.yaml index f2933c1..7cfdc79 100644 --- a/tests/bad_deployments/check_schema_deployment_basic_publication_date_1_1_2026_is_not_a_date.yaml +++ b/tests/bad_deployments/check_schema_deployment_basic_publication_date_1_1_2026_is_not_a_date.yaml @@ -11,6 +11,6 @@ deployment: data_product_type: Summary statistics data_product_region: ... publication_date: "1/1/2026" - resources: + administrative: registry_authors: - ... diff --git a/tests/bad_deployments/check_schema_deployment_basic_publication_date_datetime_date_2026_1_1_is_not_of_type_string.yaml b/tests/bad_deployments/check_schema_deployment_basic_publication_date_datetime_date_2026_1_1_is_not_of_type_string.yaml index 936ae9d..38f140f 100644 --- a/tests/bad_deployments/check_schema_deployment_basic_publication_date_datetime_date_2026_1_1_is_not_of_type_string.yaml +++ b/tests/bad_deployments/check_schema_deployment_basic_publication_date_datetime_date_2026_1_1_is_not_of_type_string.yaml @@ -11,6 +11,6 @@ deployment: data_product_type: Summary statistics data_product_region: ... publication_date: 2026-01-01 - resources: + administrative: registry_authors: - ... diff --git a/tests/bad_deployments/check_schema_status_is_a_required_property.yaml b/tests/bad_deployments/check_schema_status_is_a_required_property.yaml index 6147d0e..77871ce 100644 --- a/tests/bad_deployments/check_schema_status_is_a_required_property.yaml +++ b/tests/bad_deployments/check_schema_status_is_a_required_property.yaml @@ -10,6 +10,6 @@ deployment: data_product_type: Summary statistics data_product_region: ... publication_date: "2026-01-01" - resources: + administrative: registry_authors: - ... diff --git a/tests/bad_deployments/check_schema_tier_0_is_not_one_of_1_2_3.yaml b/tests/bad_deployments/check_schema_tier_0_is_not_one_of_1_2_3.yaml index 2be8ab7..77be098 100644 --- a/tests/bad_deployments/check_schema_tier_0_is_not_one_of_1_2_3.yaml +++ b/tests/bad_deployments/check_schema_tier_0_is_not_one_of_1_2_3.yaml @@ -11,6 +11,6 @@ deployment: data_product_type: Summary statistics data_product_region: ... publication_date: "2026-01-01" - resources: + administrative: registry_authors: - ... diff --git a/tests/good_deployments/google_tier_1.yaml b/tests/good_deployments/google_tier_1.yaml index 48db6ba..026f54c 100644 --- a/tests/good_deployments/google_tier_1.yaml +++ b/tests/good_deployments/google_tier_1.yaml @@ -11,6 +11,6 @@ deployment: data_product_type: Summary statistics data_product_region: Global publication_date: "2020-09-01" - resources: + administrative: registry_authors: - Elena Ghazi diff --git a/tests/good_deployments/google_tier_2.yaml b/tests/good_deployments/google_tier_2.yaml index a243ac5..e1c0b55 100644 --- a/tests/good_deployments/google_tier_2.yaml +++ b/tests/good_deployments/google_tier_2.yaml @@ -24,6 +24,6 @@ deployment: release_type: Many releases access_type: Non-interactive data_source_type: Static - resources: + administrative: registry_authors: - Elena Ghazi diff --git a/tests/good_deployments/google_tier_3.yaml b/tests/good_deployments/google_tier_3.yaml index fd8eb59..832ef13 100644 --- a/tests/good_deployments/google_tier_3.yaml +++ b/tests/good_deployments/google_tier_3.yaml @@ -44,6 +44,6 @@ deployment: pre_processing_eda_hyperparameter_tuning: Not available mechanisms: Laplace mechanism for both symptom search and normalization counts via Google’s open-source DP library justification: Not available - resources: + administrative: registry_authors: - Elena Ghazi diff --git a/tests/good_deployments/template.yaml b/tests/good_deployments/template.yaml index 70616fb..c4d3988 100644 --- a/tests/good_deployments/template.yaml +++ b/tests/good_deployments/template.yaml @@ -53,7 +53,7 @@ deployment: mechanisms: '' # Tier 3 justification: '' # Tier 3 - resources: + administrative: sources: '' # white papers notes: '' # other extra information registry_authors: diff --git a/tests/test_schema.py b/tests/test_schema.py index 80a7e33..903bd0c 100644 --- a/tests/test_schema.py +++ b/tests/test_schema.py @@ -38,9 +38,9 @@ def test_node_has_description(path, node): "/deployment/deployment_model/model_name_description", "/deployment/accounting", "/deployment/implementation", - "/deployment/resources", - "/deployment/resources/sources", - "/deployment/resources/notes", + "/deployment/administrative", + "/deployment/administrative/sources", + "/deployment/administrative/notes", ]: assert "description" not in node.keys() pytest.skip("TODO: More description would be nice to have") @@ -75,9 +75,9 @@ def test_node_has_description_long(path, node): "/deployment/deployment_model/release_type_description", "/deployment/deployment_model/data_source_type_description", "/deployment/deployment_model/access_type_description", - "/deployment/resources", - "/deployment/resources/notes", - "/deployment/resources/registry_authors", + "/deployment/administrative", + "/deployment/administrative/notes", + "/deployment/administrative/registry_authors", ] if path in skip_list: assert "description_long" not in node.keys() @@ -100,7 +100,7 @@ def test_node_has_tier(path, node): "/deployment/dp_flavor/bound_on_output_distance", "/deployment/deployment_model/model_name_description", "/deployment/deployment_model/release_type_description", - "/deployment/resources/registry_authors", + "/deployment/administrative/registry_authors", ]: assert "tier" not in node.keys() pytest.skip("TODO: More tiers would be nice to have")