Retrieved article excerpt
Open article Β· Retrieved 2026-09-16T05:21:31.745411+00:00
# DATAMIMIC β Governed Test Data for Regulated Enterprises
> **This repository contains the DATAMIMIC Community Edition (CE).** MIT-licensed, Python-native, MCP-ready.
>
> CE is fully usable standalone for deterministic synthetic data generation and PII-aware pseudonymization. The Enterprise Platform adds governed workflows, PII scanning, role-based access, audit logging, scheduling, multi-system execution, and the full operational layer that regulated enterprises require.
>
> π **Enterprise Platform:** [datamimic.io](https://datamimic.io) Β |Β π **Docs:** [docs.datamimic.io](https://docs.datamimic.io) Β |Β π
**Book a strategy call:** [datamimic.io/contact](https://datamimic.io/contact)
>
> π€ **AI agent?** Start at [`AGENTS.md`](https://github.com/rapiddweller/datamimic/blob/development/AGENTS.md) and use the project CLI: preserve new intent as `model.dm.json`, submit an early best attempt via `datamimic scaffold ... --format json`, repair from the structured issues, declare an expectation per stated requirement, and stop on `verified=true`. Existing raw XML uses lint plus bounded dry-run.
---
[CI](https://github.com/rapiddweller/datamimic/actions)
[Coverage](https://sonarcloud.io/summary/new_code?id=rapiddweller_datamimic)
[Maintainability](https://sonarcloud.io/summary/new_code?id=rapiddweller_datamimic)
[Python](https://www.python.org/downloads/)
[License: MIT](https://github.com/rapiddweller/datamimic/blob/development/LICENSE)
[MCP Ready](https://github.com/rapiddweller/datamimic/blob/development/docs/mcp_quickstart.md)
---
## What is DATAMIMIC?
**DATAMIMIC CE is the open-source deterministic data engine at the core of the DATAMIMIC Enterprise Platform.** It is usable standalone for synthetic data generation and PII-aware pseudonymization in any local, CI, or agent-driven workflow.
The Enterprise Platform adds the governed workflows, scanners, dashboards, and execution layer that regulated enterprises require for production-scale test-data operations.
**Available in CE (this repo):**
- **Generate** fully synthetic, deterministic datasets β model-driven, no source data required
- **Pseudonymize** staging/QA exports β deterministic (seeded) or privacy-maximized (non-seeded) field transformation; PII fields identified and modeled manually in the XML pipeline
- **Execute** single-system pipelines against PostgreSQL Β· MySQL Β· Oracle Β· MS SQL Β· SQLite Β· MongoDB Β· CSV Β· JSON Β· XML Β· XLSX Β· DbUnit Β· fixed-width (`.fcw`)
- **Model behavior** β weighted state machines, composite multi-field references, control flow (`<while>`, `<assert>`), and a scriptable memstore for staged aggregation
- **Emit provenance** β append-only execution logs and per-output content hash for audit re-execution
- **Guide agents** β machine-readable capabilities, progressive reference queries, and one canonical CLI scaffold transaction; an optional MCP adapter exposes the same authoring service
**The Enterprise Platform adds:**
- **PII scanner** β probability-scored field detection with configurable thresholds via DataWorkbench
- **Multi-system execution** β Oracle / MongoDB / Kafka in coordinated workflows with referential integrity
- **Industry message templates** β EDIFACT / SWIFT MT / HL7 v2.x / HL7 FHIR generated as deterministic test/training artefacts
- **Governance layer** β role-based dashboards, audit trails, approval flows, reusable enterprise templates, scheduler
- **Performance core** β Rust fastpath, ML/auto-regressive engine for complex distributions, keyset and manifest building, optimised distributed execution
- **On-premise / air-gapped deployment** β podman-compose or Helm, with consulting-led rollout
> Deployed in regulated EU banking environments for deterministic test data across Oracle, MongoDB, and Kafka pipelines. Reference customers available under NDA β see also [datamimic.io case studies](https://datamimic.io).
---
## AI agents: author, verify, and run data models
The CLI is the baseline agent contract. Install CE with `pip install datamimic-ce`;
inside this checkout, use `.venv/bin/datamimic` so a stale global installation cannot
change the available schema or commands.
| Need | CLI tool | Contract |
| --- | --- | --- |
| Discover the live structural surface | `datamimic capabilities` | Compact machine-readable JSON index by default; `--full` for the complete manifest, `--section <name>` for one section. |
| Learn the Intent Model progressively | `datamimic reference authoring`, then `datamimic reference authoring --category <category> --kind <kind>` | Start with the query catalogue, then load only the typed fragment needed. |
| Author a new model | Preserve `model.dm.json`; run `datamimic scaffold model.dm.json --format json` | One compile/lint/bounded-run/acceptance transaction per changed attempt. Stop on `verified=true`; generated XML is runtime output. |
| Work with existing raw XML | `datamimic lint model.xml --format json`, then `datamimic dry-run model.xml --format json` | Fix diagnostics, inspect bounded samples for intent, then use `datamimic run model.xml` only when real execution is requested. |
| Find a DSL detail | `datamimic reference overview`, then a narrow `reference` topic/name | Query the live model and rule registries instead of guessing elements, generators, scope, distributions, or rules. |
`capabilities`, authoring-reference projections, and the commands shown with
`--format json` return machine-readable JSON. On a failed scaffold attempt, change
`model.dm.json` using its structured validation issues, typed repair, or rule
diagnostics before retrying. A typed `max_count` remediation instead changes only
the bounded scaffold parameter to at least its reported minimum. Never repeat an
identical failed call. A successful scaffold result is terminal for authoring, so
do not lint or dry-run its generated XML again. Exact source fragments are
discoverable through queries such as `--category source --kind memstore`.
### Optional MCP adapter
When the calling environment already exposes DATAMIMIC MCP tools, they map to the same
canonical contracts and implementations: `reference` β `datamimic_reference`,
`scaffold` β `datamimic_scaffold`, `lint` β `datamimic_check`, and `dry-run` β
`datamimic_run`. Install the adapter with `pip install "datamimic-ce[mcp]"`;
registration details belong in the [`MCP quickstart`](https://github.com/rapiddweller/datamimic/blob/development/docs/mcp_quickstart.md),
not in the authoring workflow. The adapter intentionally exposes only the four
canonical reference, scaffold, check, and bounded-run operations; domain generation
remains a Python/CLI capability rather than a parallel MCP authoring path.
### Prompts to paste into your agent
**Author and verify a new model**
```
Create the dataset I describe with DATAMIMIC.
Read AGENTS.md first. In a repository checkout use `.venv/bin/datamimic`;
otherwise use the current `datamimic` CLI. Preserve my intent as
`model.dm.json`; do not hand-write XML.
Start from the minimal valid document shape in AGENTS.md ("Authoring a new
model"). Two rules prevent most rejections: the top level allows ONLY
version, seed, products, expectations; product-level "kind"
(generated/source/time_series) is a different vocabulary from field-level
"kind" (increment, values, weighted, int_range, decimal_range, pattern,
constant, script). Range fields take minimum/maximum, never min/max.
Submit EARLY: run `datamimic scaffold model.dm.json --format json` with your
best attempt after at most one discovery call. Repair from the structured
issues (path/code/message/allowed_fields) and diagnostics (fix_hint) β they
teach the schema faster than more discovery. Never resubmit an unchanged
document. If a remediation requests a larger max_count, retry scaffold with
at least that value without changing the intent.
Declare an expectation for every requirement I state (counts as exact_count
with a "count" field, uniqueness, allowed values, ranges, foreign keys) β
verified=true certifies only what you declared. Stop on verified=true; do
not lint or dry-run the generated XML. If I request real execution, save the
returned XML as a generated artifact and run that descriptor. Return the
model.dm.json path and concise verification evidence.
```
**Relational hierarchy with referential integrity (fully supported β no XML
needed)**
```
Seed a relational dataset with referential integrity: 4 customers, each with
exactly 2 orders.
Customers get an incrementing unique id and a region from
{north, south, east, west}. Each order carries the REAL parent customer id
as a foreign key and an amount between 10.0 and 500.0.
Follow AGENTS.md's "Authoring a new model" and its structural recipes:
orders nest inside the customer product's "children" array; the FK field is
{"kind": "script", "script": "parent.id"} with a foreign_key role β a
randomly generated FK passes schema validation but fails per-parent-count
acceptance. Declare expectations for the customer count, customer id
uniqueness, exactly 2 orders per customer (per_parent_count), the
orders->customers foreign key, and the amount range. Stop on verified=true
and show the acceptance evidence.
```
Raw XML remains supported for existing descriptors (lint β dry-run β run;
see AGENTS.md). For new models it is a last resort: only when a scaffold
issue explicitly classifies the requirement as `unsupported_intent` should
an agent hand-author XML, preserving that evidence.
---
## CE vs Enterprise Platform
CE and EE are **not the same engine with a feature flag**. They share the DSL and determinism contract, but EE is an independently optimised execution engine built for enterprise-scale throughput and operational control.
### Engine comparison
| Capability | Community Edition (CE) | Enterprise Platform (EE) |
| --- | --- | --- |
| Deterministic data generation | β
| β
|
| Deterministic seeding in the DSL | β
entities + standalone literal `<key generator>` (4.0.0) | β
same, plus sandboxed script expressions and stdlib `random` calls |
| **Pseudonymization β seeded** *(GDPR Art. 4(5); supports Art. 25 / Art. 32)* | β
manual model | β
automated via DataWorkbench |
| **Pseudonymization β non-seeded (privacy-maximized)** | β
manual model | β
automated via DataWorkbench |
| Python API + XML pipelines | β
| β
|
| Domain models: Finance, Healthcare, Demographics | β
| β
|
| Time-series generation (`<generate start/end/interval>`, ISO 8601, prefix-stable) | β
| β
|
| MCP server for AI agent integration | β
| β
|
| CLI + local execution | β
| β
|
| **Scale** | millions of records via Python multiprocessing (and optional Ray) | **designed for billion-record workloads** β Rust fastpath, optimised multi-process execution, and keyset/manifest building on top of the shared Ray distribution layer |
| **PII scanner** | β | β
probability-scored field detection, configurable threshold, DataWorkbench integration |
| **Runtime configuration profiles** | β | β
Performance Β· Balanced Β· Flexibility |
| **Memory management** | standard | optimised for high-volume batch and streaming |
| **Logging granularity** | flat execution log | configurable: minimal Β· standard Β· deep nested tracing |
| **Nested structure evaluation** | basic | deep nested generation with extended condition + ruleset evaluation |
| **Importer / exporter logging** | β | per-stage logging for importers and exporters |
| **Error handling** | standard exceptions | structured error catalog with recovery strategies |
| **Rust fastpath** | β | performance-critical paths in Rust |
| **Keyset and manifest building** | β | reads live DB schemas to build coordinated multi-table generation plans |
| **ML / auto-regressive engine** | β | combine statistical models with conditions, rulesets, validators for complex distributions |
### Platform capabilities (EE only)
| Capability | EE |
| --- | --- |
| Multi-user collaboration | β
|
| Role-based access control (RBAC) | β
|
| Audit logs + provenance dashboards | β
|
| **PII scanne