time estimation and experiment details in README.md
This commit is contained in:
parent
000d1e1f2c
commit
02f69f6e54
8 changed files with 47 additions and 3 deletions
1
.gitignore
vendored
1
.gitignore
vendored
|
|
@ -3,3 +3,4 @@ venv/
|
||||||
*.parquet
|
*.parquet
|
||||||
*.zip
|
*.zip
|
||||||
*.csv
|
*.csv
|
||||||
|
__pycache__/
|
||||||
|
|
|
||||||
45
README.md
45
README.md
|
|
@ -1,3 +1,44 @@
|
||||||
# Budget Allocation in Differential Privacy
|
# Experiment: PREDICTIVE METRIC FOR OPTIMAL BUDGET ALLOCATION IN DIFFERENTIAL PRIVACY
|
||||||
|
|
||||||
O experimento está em experiment.py
|
Based on the [original experiment](https://github.com/conseg/TheImpactofDifferentialPrivacyondatautilityinfundamentalmathematicaloperations), but using a real life dataset.
|
||||||
|
|
||||||
|
Key difference: sensitivity calculation to satisfy Differential Privacy.
|
||||||
|
|
||||||
|
## The Dataset
|
||||||
|
|
||||||
|
PNAD Contínua (IBGE - Brasil) - 2026 (primeiro trimestre)
|
||||||
|
|
||||||
|
## The Experiment
|
||||||
|
|
||||||
|
Objective: evaluate the metric on a t-test equation to compare Northeast and Southeast income (informal vs formal workers)
|
||||||
|
|
||||||
|
Steps:
|
||||||
|
|
||||||
|
1. Clip and normalize dataset (C = 46_366, the constitutional salary cap, which is set at the salary of Supreme Court justices). Must be a value taken from the real world, not from the dataset.
|
||||||
|
2. Calculate necessary statistics and their sensitivities
|
||||||
|
3. For each possible budget allocation with total epsilon=12, granularity=0.5, epsilon > 0 (1,352,078 possibilities): evaluate the metric for the statistics and the budget allocation sequence
|
||||||
|
4. Save all possible scores and budget allocation sequences to csv for further analysis, and print the best metric score.
|
||||||
|
|
||||||
|
## Estimated Experiment Execution Time
|
||||||
|
|
||||||
|
To reproduce (in a Linux shell):
|
||||||
|
|
||||||
|
1. `time python ./experiment.py real`
|
||||||
|
2. Send an interrupt (Ctrl+C) when N exceeds 500.
|
||||||
|
3. To find the number of hours, based on the last N and the elapsed time (t): `((1352078*t)/n)/3600`
|
||||||
|
|
||||||
|
Laptop (Intel Core i5 1135G7 - 8GB RAM)
|
||||||
|
PC (AMD Ryzen 7 8700G - 32GB RAM)
|
||||||
|
|
||||||
|
Laptop: 514 iterations in 12.13s
|
||||||
|
PC: 493 iterations in 8.11s
|
||||||
|
|
||||||
|
Estimated total time on the laptop: 8h51m
|
||||||
|
Estimated total time on the PC: 6h10m
|
||||||
|
|
||||||
|
## Tasks
|
||||||
|
|
||||||
|
1. Provide a way to run the experiment in parts, to allow parallel execution on multiple CPUs.
|
||||||
|
2. Run the experiment to evaluate the metric.
|
||||||
|
3. Combine the resulting CSVs into one.
|
||||||
|
4. Analyze the csv to find notable metric scores for specific sequences.
|
||||||
|
|
|
||||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
|
|
@ -126,7 +126,9 @@ print(formal_sen)
|
||||||
results = []
|
results = []
|
||||||
best_metric = sys.float_info.max
|
best_metric = sys.float_info.max
|
||||||
|
|
||||||
|
n = 0
|
||||||
for bud in sequences:
|
for bud in sequences:
|
||||||
|
n += 1
|
||||||
metric = 0
|
metric = 0
|
||||||
sta = informal_sta + formal_sta
|
sta = informal_sta + formal_sta
|
||||||
sen = informal_sen + formal_sen
|
sen = informal_sen + formal_sen
|
||||||
|
|
@ -135,7 +137,7 @@ for bud in sequences:
|
||||||
metric += us_result
|
metric += us_result
|
||||||
metric += ue_result[0] + ue_result[1]
|
metric += ue_result[0] + ue_result[1]
|
||||||
metric = metric / 14.0
|
metric = metric / 14.0
|
||||||
print(metric)
|
print(f"n={n}; metric={metric}")
|
||||||
results.append([*bud, metric])
|
results.append([*bud, metric])
|
||||||
if metric < best_metric:
|
if metric < best_metric:
|
||||||
best_metric = metric
|
best_metric = metric
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue