1.9 KiB
1.9 KiB
Experiment: PREDICTIVE METRIC FOR OPTIMAL BUDGET ALLOCATION IN DIFFERENTIAL PRIVACY
Based on the original experiment, but using a real life dataset.
Key difference: sensitivity calculation to satisfy Differential Privacy.
The Dataset
PNAD Contínua (IBGE - Brasil) - 2026 (primeiro trimestre)
The Experiment
Objective: evaluate the metric on a t-test equation to compare Northeast and Southeast income (informal vs formal workers)
Steps:
- Clip and normalize dataset (C = 46_366, the constitutional salary cap, which is set at the salary of Supreme Court justices). Must be a value taken from the real world, not from the dataset.
- Calculate necessary statistics and their sensitivities
- For each possible budget allocation with total epsilon=12, granularity=0.5, epsilon > 0 (1,352,078 possibilities): evaluate the metric for the statistics and the budget allocation sequence
- Save all possible scores and budget allocation sequences to csv for further analysis, and print the best metric score.
Estimated Experiment Execution Time
To reproduce (in a Linux shell):
time python ./experiment.py real- Send an interrupt (Ctrl+C) when N exceeds 500.
- To find the number of hours, based on the last N and the elapsed time (t):
((1352078*t)/n)/3600
Tested Hardware:
- Laptop (Intel Core i5 1135G7 - 8GB RAM)
- Desktop (AMD Ryzen 7 8700G - 32GB RAM)
Partial Results:
- Laptop: 514 iterations in 12.13s
- Desktop: 493 iterations in 8.11s
Total Results:
- Estimated total time on the Laptop: 8h51m
- Estimated total time on the Desktop: 6h10m
Tasks
- Provide a way to run the experiment in parts, to allow parallel execution on multiple CPUs.
- Run the experiment to evaluate the metric.
- Combine the resulting CSVs into one.
- Analyze the csv to find notable metric scores for specific sequences.