time estimation and experiment details in README.md

This commit is contained in:
Gabriel Franco 2026-09-01 10:51:30 -03:00
parent 000d1e1f2c
commit 02f69f6e54
8 changed files with 47 additions and 3 deletions

1
.gitignore vendored
View file

@ -3,3 +3,4 @@ venv/
*.parquet *.parquet
*.zip *.zip
*.csv *.csv
__pycache__/

View file

@ -1,3 +1,44 @@
# Budget Allocation in Differential Privacy # Experiment: PREDICTIVE METRIC FOR OPTIMAL BUDGET ALLOCATION IN DIFFERENTIAL PRIVACY
O experimento está em experiment.py Based on the [original experiment](https://github.com/conseg/TheImpactofDifferentialPrivacyondatautilityinfundamentalmathematicaloperations), but using a real life dataset.
Key difference: sensitivity calculation to satisfy Differential Privacy.
## The Dataset
PNAD Contínua (IBGE - Brasil) - 2026 (primeiro trimestre)
## The Experiment
Objective: evaluate the metric on a t-test equation to compare Northeast and Southeast income (informal vs formal workers)
Steps:
1. Clip and normalize dataset (C = 46_366, the constitutional salary cap, which is set at the salary of Supreme Court justices). Must be a value taken from the real world, not from the dataset.
2. Calculate necessary statistics and their sensitivities
3. For each possible budget allocation with total epsilon=12, granularity=0.5, epsilon > 0 (1,352,078 possibilities): evaluate the metric for the statistics and the budget allocation sequence
4. Save all possible scores and budget allocation sequences to csv for further analysis, and print the best metric score.
## Estimated Experiment Execution Time
To reproduce (in a Linux shell):
1. `time python ./experiment.py real`
2. Send an interrupt (Ctrl+C) when N exceeds 500.
3. To find the number of hours, based on the last N and the elapsed time (t): `((1352078*t)/n)/3600`
Laptop (Intel Core i5 1135G7 - 8GB RAM)
PC (AMD Ryzen 7 8700G - 32GB RAM)
Laptop: 514 iterations in 12.13s
PC: 493 iterations in 8.11s
Estimated total time on the laptop: 8h51m
Estimated total time on the PC: 6h10m
## Tasks
1. Provide a way to run the experiment in parts, to allow parallel execution on multiple CPUs.
2. Run the experiment to evaluate the metric.
3. Combine the resulting CSVs into one.
4. Analyze the csv to find notable metric scores for specific sequences.

Binary file not shown.

Binary file not shown.

Binary file not shown.

View file

@ -126,7 +126,9 @@ print(formal_sen)
results = [] results = []
best_metric = sys.float_info.max best_metric = sys.float_info.max
n = 0
for bud in sequences: for bud in sequences:
n += 1
metric = 0 metric = 0
sta = informal_sta + formal_sta sta = informal_sta + formal_sta
sen = informal_sen + formal_sen sen = informal_sen + formal_sen
@ -135,7 +137,7 @@ for bud in sequences:
metric += us_result metric += us_result
metric += ue_result[0] + ue_result[1] metric += ue_result[0] + ue_result[1]
metric = metric / 14.0 metric = metric / 14.0
print(metric) print(f"n={n}; metric={metric}")
results.append([*bud, metric]) results.append([*bud, metric])
if metric < best_metric: if metric < best_metric:
best_metric = metric best_metric = metric