[No. 006]ML Tooling & Open Source

DataDoom

Local-first synthetic data engine packaged as an extensible Python library, CLI, and interactive Web Canvas

SR
bySanthosh Reddy
RoleCreator & Core Architect
TimelineJune 2026 - Present
Read9 min
DataDoom
FIG. 01 - DataDoom overviewDataDoom.case

Introduction

DataDoom started as a question: what would a synthetic dataset look like if it were as version-controllable, shareable, and reproducible as source code? The product is a local-first, open-source Python engine that regenerates identical datasets byte-for-byte from a single seed, packaged with a CLI, ML framework adapters, and an interactive web Canvas.

The black-box tradeoff in synthetic data

Existing synthetic data tools force a trade-off: they are either realistic but act as a black box, preventing you from auditing causal relationships or quality failures, or they are controllable but throwaway, preventing you from recreating the exact same dataset tomorrow. For an ML engineer or educator, that makes benchmarking and debugging nearly impossible.

A dataset should be a recipe you version-control and share ? not a massive, static database of PII.

Built with
PythonNumPySciPyPandasNetworkXFastAPIReact Flow

Architecture

I engineered a deterministic-by-construction pipeline powered by key derivation, hashing spec plus seed, and namespace-isolated PCG64 generators, ensuring identical dataset recreation across pinned runtimes. On top of this sits a Directed Acyclic Graph (DAG) for causal structural equations, supporting linear, logistic, and polynomial causal paths with per-node noise and do-interventions.

To make the engine accessible, I packaged it as a dual-surface developer tool: a CLI for automation and CI, and an interactive web Canvas served locally via FastAPI. The frontend leverages React Flow for causal diagramming, live WebSockets for generation progress tracking, and interactive controls for calibrating baseline AUROC difficulty and injecting data-quality failures such as MCAR, MAR, MNAR, drift, and leakage.

What I prioritized

A few of the package design and architectural decisions that mattered most:

  • Bitwise Reproducibility. Seeded RNG factory and key-derivation pipeline guaranteeing identical dataset replication from a single spec and seed.
  • Layered Architecture. Strict architectural boundaries, CLI to API to Jobs to Engine, enforced by import-linter to keep the core engine lightweight.
  • Dual-Surface Package. A single pip-installable tool containing both a CLI for headless automation and a precompiled React web Canvas.
  • Adaptive Difficulty. A bisection search loop that automatically tunes feature noise and label flips to hit a target baseline-model AUROC.

Production hardening

Having completed Phases 0 through 5, the core package features time-series generation, 8 failure-injection mechanisms, and a plugin registry. The current roadmap focuses on 1.0 release hardening, expanding the cross-environment reproducibility matrix in CI, and deploying precompiled domain templates for hackathons and academic challenges.