Introduction
DataDoom started as a question: what would a synthetic dataset look like if it were as version-controllable, shareable, and reproducible as source code? The product is a local-first, open-source Python engine that regenerates identical datasets byte-for-byte from a single seed, packaged with a CLI, ML framework adapters, and an interactive web Canvas.
The black-box tradeoff in synthetic data
Existing synthetic data tools force a trade-off: they are either realistic but act as a black box, preventing you from auditing causal relationships or quality failures, or they are controllable but throwaway, preventing you from recreating the exact same dataset tomorrow. For an ML engineer or educator, that makes benchmarking and debugging nearly impossible.
A dataset should be a recipe you version-control and share ? not a massive, static database of PII.
Architecture
I engineered a deterministic-by-construction pipeline powered by key derivation, hashing spec plus seed, and namespace-isolated PCG64 generators, ensuring identical dataset recreation across pinned runtimes. On top of this sits a Directed Acyclic Graph (DAG) for causal structural equations, supporting linear, logistic, and polynomial causal paths with per-node noise and do-interventions.
To make the engine accessible, I packaged it as a dual-surface developer tool: a CLI for automation and CI, and an interactive web Canvas served locally via FastAPI. The frontend leverages React Flow for causal diagramming, live WebSockets for generation progress tracking, and interactive controls for calibrating baseline AUROC difficulty and injecting data-quality failures such as MCAR, MAR, MNAR, drift, and leakage.
What I prioritized
A few of the package design and architectural decisions that mattered most:
- Bitwise Reproducibility. Seeded RNG factory and key-derivation pipeline guaranteeing identical dataset replication from a single spec and seed.
- Layered Architecture. Strict architectural boundaries, CLI to API to Jobs to Engine, enforced by import-linter to keep the core engine lightweight.
- Dual-Surface Package. A single pip-installable tool containing both a CLI for headless automation and a precompiled React web Canvas.
- Adaptive Difficulty. A bisection search loop that automatically tunes feature noise and label flips to hit a target baseline-model AUROC.


