2026-10-11 17:10 UTC

Rapiddweller claims DATAMIMIC CE's released CLI and MCP adapter let coding agents generate deterministic test datasets and verify declared requirements, potentially replacing ad hoc fixtures with reproducible, constraint-checked test data.

state: seedheat: mediumuncertainty: mediumknownscott: lowcoding-agents software-testing mcpRapiddweller

What is this?

DATAMIMIC CE is rapiddweller’s model-driven synthetic test-data tool, whose repository documentation presents a CLI and Python API as interfaces for AI agents to generate seeded, reproducible datasets. The documentation describes an optional MCP adapter exposing reference, scaffold, check, and bounded-run operations, plus output content hashes for verifying repeatability. These are maintainer claims: the supplied snippets do not establish a dated release of these capabilities, the scope of declared-requirement verification, or independent evidence that they replace ad hoc fixtures effectively.

Why it matters to Scott

The underlying position is already held in Scott’s Generative Test Scaffolding and Deterministic test seams via injected interfaces pages; DATAMIMIC is another implementation candidate, with no supplied evidence that it changes his existing harnesses or establishes a consequential new endorsement. Seeded datasets and repeatability hashes do not establish the independent requirement verification demanded by his Mechanically Different Verifiers concept, so the stronger replacement claim remains unproven.
ip:concept.generative-test-scaffoldingdev:concept.deterministic-test-seamsip:concept.mechanically-different-verifiersradar:concept.agent-verificationradar:concept.reproducibilityradar:five-bugs-test-spec-blindspot
queries asked of Scott's wikis
  • coding agent self-authored tests independent verification
  • deterministic fixtures seeded datasets regression harnesses
  • executable requirements constraint validation test oracles
  • CLI-first agent tools optional MCP adapters
  • reproducible execution content hashes provenance

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 611h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-16 05:21 (minted)⭐ origin echo-reconstructedDATAMIMIC CE provides deterministic synthetic data generation with a CLI scaffold transaction and optional MCP adapter; its documentation ex
rapiddweller on github (echo) Β· attributed from hn.story.49722276 Β· published time unknown
β€”
09-16 04:58first on hacker news Β· published Β· lag ?Datamimic – don't let your coding agent invent its own test world
ake2l
β€”
09-16 04:58amplified on hacker news πŸ‘‘hn.story.49722276
ake2l
peak 59 Β· 8 comments Β· 100% of case engagement
09-16 05:20our radar first saw it Β· lag ?discovery anchor: hn.story.49722276β€”
pace: p63 vs 1032 stories at the 336h mark (now 611h old) β€” ahead of cache-tax-idle-session-warming (1.0x), behind hunterbench-live-pentesting-benchmark (1.0x)

Evidence (2) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnDatamimic – don't let your coding agent invent its own test world
Retrieved article excerpt

Open article Β· Retrieved 2026-09-16T05:21:31.745411+00:00

# DATAMIMIC β€” Governed Test Data for Regulated Enterprises

> **This repository contains the DATAMIMIC Community Edition (CE).** MIT-licensed, Python-native, MCP-ready.
>
> CE is fully usable standalone for deterministic synthetic data generation and PII-aware pseudonymization. The Enterprise Platform adds governed workflows, PII scanning, role-based access, audit logging, scheduling, multi-system execution, and the full operational layer that regulated enterprises require.
>
> πŸ‘‰ **Enterprise Platform:** [datamimic.io](https://datamimic.io) Β |Β  πŸ“˜ **Docs:** [docs.datamimic.io](https://docs.datamimic.io) Β |Β  πŸ“… **Book a strategy call:** [datamimic.io/contact](https://datamimic.io/contact)
>
> πŸ€– **AI agent?** Start at [`AGENTS.md`](https://github.com/rapiddweller/datamimic/blob/development/AGENTS.md) and use the project CLI: preserve new intent as `model.dm.json`, submit an early best attempt via `datamimic scaffold ... --format json`, repair from the structured issues, declare an expectation per stated requirement, and stop on `verified=true`. Existing raw XML uses lint plus bounded dry-run.

---

[CI](https://github.com/rapiddweller/datamimic/actions)
[Coverage](https://sonarcloud.io/summary/new_code?id=rapiddweller_datamimic)
[Maintainability](https://sonarcloud.io/summary/new_code?id=rapiddweller_datamimic)
[Python](https://www.python.org/downloads/)
[License: MIT](https://github.com/rapiddweller/datamimic/blob/development/LICENSE)
[MCP Ready](https://github.com/rapiddweller/datamimic/blob/development/docs/mcp_quickstart.md)

---

## What is DATAMIMIC?

**DATAMIMIC CE is the open-source deterministic data engine at the core of the DATAMIMIC Enterprise Platform.** It is usable standalone for synthetic data generation and PII-aware pseudonymization in any local, CI, or agent-driven workflow.

The Enterprise Platform adds the governed workflows, scanners, dashboards, and execution layer that regulated enterprises require for production-scale test-data operations.

**Available in CE (this repo):**

- **Generate** fully synthetic, deterministic datasets β€” model-driven, no source data required
- **Pseudonymize** staging/QA exports β€” deterministic (seeded) or privacy-maximized (non-seeded) field transformation; PII fields identified and modeled manually in the XML pipeline
- **Execute** single-system pipelines against PostgreSQL Β· MySQL Β· Oracle Β· MS SQL Β· SQLite Β· MongoDB Β· CSV Β· JSON Β· XML Β· XLSX Β· DbUnit Β· fixed-width (`.fcw`)
- **Model behavior** β€” weighted state machines, composite multi-field references, control flow (`<while>`, `<assert>`), and a scriptable memstore for staged aggregation
- **Emit provenance** β€” append-only execution logs and per-output content hash for audit re-execution
- **Guide agents** β€” machine-readable capabilities, progressive reference queries, and one canonical CLI scaffold transaction; an optional MCP adapter exposes the same authoring service

**The Enterprise Platform adds:**

- **PII scanner** β€” probability-scored field detection with configurable thresholds via DataWorkbench
- **Multi-system execution** β€” Oracle / MongoDB / Kafka in coordinated workflows with referential integrity
- **Industry message templates** β€” EDIFACT / SWIFT MT / HL7 v2.x / HL7 FHIR generated as deterministic test/training artefacts
- **Governance layer** β€” role-based dashboards, audit trails, approval flows, reusable enterprise templates, scheduler
- **Performance core** β€” Rust fastpath, ML/auto-regressive engine for complex distributions, keyset and manifest building, optimised distributed execution
- **On-premise / air-gapped deployment** β€” podman-compose or Helm, with consulting-led rollout

> Deployed in regulated EU banking environments for deterministic test data across Oracle, MongoDB, and Kafka pipelines. Reference customers available under NDA β€” see also [datamimic.io case studies](https://datamimic.io).

---

## AI agents: author, verify, and run data models

The CLI is the baseline agent contract. Install CE with `pip install datamimic-ce`;
inside this checkout, use `.venv/bin/datamimic` so a stale global installation cannot
change the available schema or commands.

| Need | CLI tool | Contract |
| --- | --- | --- |
| Discover the live structural surface | `datamimic capabilities` | Compact machine-readable JSON index by default; `--full` for the complete manifest, `--section <name>` for one section. |
| Learn the Intent Model progressively | `datamimic reference authoring`, then `datamimic reference authoring --category <category> --kind <kind>` | Start with the query catalogue, then load only the typed fragment needed. |
| Author a new model | Preserve `model.dm.json`; run `datamimic scaffold model.dm.json --format json` | One compile/lint/bounded-run/acceptance transaction per changed attempt. Stop on `verified=true`; generated XML is runtime output. |
| Work with existing raw XML | `datamimic lint model.xml --format json`, then `datamimic dry-run model.xml --format json` | Fix diagnostics, inspect bounded samples for intent, then use `datamimic run model.xml` only when real execution is requested. |
| Find a DSL detail | `datamimic reference overview`, then a narrow `reference` topic/name | Query the live model and rule registries instead of guessing elements, generators, scope, distributions, or rules. |

`capabilities`, authoring-reference projections, and the commands shown with
`--format json` return machine-readable JSON. On a failed scaffold attempt, change
`model.dm.json` using its structured validation issues, typed repair, or rule
diagnostics before retrying. A typed `max_count` remediation instead changes only
the bounded scaffold parameter to at least its reported minimum. Never repeat an
identical failed call. A successful scaffold result is terminal for authoring, so
do not lint or dry-run its generated XML again. Exact source fragments are
discoverable through queries such as `--category source --kind memstore`.

### Optional MCP adapter

When the calling environment already exposes DATAMIMIC MCP tools, they map to the same
canonical contracts and implementations: `reference` β†’ `datamimic_reference`,
`scaffold` β†’ `datamimic_scaffold`, `lint` β†’ `datamimic_check`, and `dry-run` β†’
`datamimic_run`. Install the adapter with `pip install "datamimic-ce[mcp]"`;
registration details belong in the [`MCP quickstart`](https://github.com/rapiddweller/datamimic/blob/development/docs/mcp_quickstart.md),
not in the authoring workflow. The adapter intentionally exposes only the four
canonical reference, scaffold, check, and bounded-run operations; domain generation
remains a Python/CLI capability rather than a parallel MCP authoring path.

### Prompts to paste into your agent

**Author and verify a new model**

```
Create the dataset I describe with DATAMIMIC.

Read AGENTS.md first. In a repository checkout use `.venv/bin/datamimic`;
otherwise use the current `datamimic` CLI. Preserve my intent as
`model.dm.json`; do not hand-write XML.

Start from the minimal valid document shape in AGENTS.md ("Authoring a new
model"). Two rules prevent most rejections: the top level allows ONLY
version, seed, products, expectations; product-level "kind"
(generated/source/time_series) is a different vocabulary from field-level
"kind" (increment, values, weighted, int_range, decimal_range, pattern,
constant, script). Range fields take minimum/maximum, never min/max.

Submit EARLY: run `datamimic scaffold model.dm.json --format json` with your
best attempt after at most one discovery call. Repair from the structured
issues (path/code/message/allowed_fields) and diagnostics (fix_hint) β€” they
teach the schema faster than more discovery. Never resubmit an unchanged
document. If a remediation requests a larger max_count, retry scaffold with
at least that value without changing the intent.

Declare an expectation for every requirement I state (counts as exact_count
with a "count" field, uniqueness, allowed values, ranges, foreign keys) β€”
verified=true certifies only what you declared. Stop on verified=true; do
not lint or dry-run the generated XML. If I request real execution, save the
returned XML as a generated artifact and run that descriptor. Return the
model.dm.json path and concise verification evidence.
```

**Relational hierarchy with referential integrity (fully supported β€” no XML
needed)**

```
Seed a relational dataset with referential integrity: 4 customers, each with
exactly 2 orders.

Customers get an incrementing unique id and a region from
{north, south, east, west}. Each order carries the REAL parent customer id
as a foreign key and an amount between 10.0 and 500.0.

Follow AGENTS.md's "Authoring a new model" and its structural recipes:
orders nest inside the customer product's "children" array; the FK field is
{"kind": "script", "script": "parent.id"} with a foreign_key role β€” a
randomly generated FK passes schema validation but fails per-parent-count
acceptance. Declare expectations for the customer count, customer id
uniqueness, exactly 2 orders per customer (per_parent_count), the
orders->customers foreign key, and the amount range. Stop on verified=true
and show the acceptance evidence.
```

Raw XML remains supported for existing descriptors (lint β†’ dry-run β†’ run;
see AGENTS.md). For new models it is a last resort: only when a scaffold
issue explicitly classifies the requirement as `unsupported_intent` should
an agent hand-author XML, preserving that evidence.

---

## CE vs Enterprise Platform

CE and EE are **not the same engine with a feature flag**. They share the DSL and determinism contract, but EE is an independently optimised execution engine built for enterprise-scale throughput and operational control.

### Engine comparison

| Capability | Community Edition (CE) | Enterprise Platform (EE) |
| --- | --- | --- |
| Deterministic data generation | βœ… | βœ… |
| Deterministic seeding in the DSL | βœ… entities + standalone literal `<key generator>` (4.0.0) | βœ… same, plus sandboxed script expressions and stdlib `random` calls |
| **Pseudonymization β€” seeded** *(GDPR Art. 4(5); supports Art. 25 / Art. 32)* | βœ… manual model | βœ… automated via DataWorkbench |
| **Pseudonymization β€” non-seeded (privacy-maximized)** | βœ… manual model | βœ… automated via DataWorkbench |
| Python API + XML pipelines | βœ… | βœ… |
| Domain models: Finance, Healthcare, Demographics | βœ… | βœ… |
| Time-series generation (`<generate start/end/interval>`, ISO 8601, prefix-stable) | βœ… | βœ… |
| MCP server for AI agent integration | βœ… | βœ… |
| CLI + local execution | βœ… | βœ… |
| **Scale** | millions of records via Python multiprocessing (and optional Ray) | **designed for billion-record workloads** β€” Rust fastpath, optimised multi-process execution, and keyset/manifest building on top of the shared Ray distribution layer |
| **PII scanner** | ❌ | βœ… probability-scored field detection, configurable threshold, DataWorkbench integration |
| **Runtime configuration profiles** | ❌ | βœ… Performance Β· Balanced Β· Flexibility |
| **Memory management** | standard | optimised for high-volume batch and streaming |
| **Logging granularity** | flat execution log | configurable: minimal Β· standard Β· deep nested tracing |
| **Nested structure evaluation** | basic | deep nested generation with extended condition + ruleset evaluation |
| **Importer / exporter logging** | ❌ | per-stage logging for importers and exporters |
| **Error handling** | standard exceptions | structured error catalog with recovery strategies |
| **Rust fastpath** | ❌ | performance-critical paths in Rust |
| **Keyset and manifest building** | ❌ | reads live DB schemas to build coordinated multi-table generation plans |
| **ML / auto-regressive engine** | ❌ | combine statistical models with conditions, rulesets, validators for complex distributions |

### Platform capabilities (EE only)

| Capability | EE |
| --- | --- |
| Multi-user collaboration | βœ… |
| Role-based access control (RBAC) | βœ… |
| Audit logs + provenance dashboards | βœ… |
| **PII scanne
ake2l598
🟧 echo.github ⭐DATAMIMIC CE provides deterministic synthetic data generation with a CLI scaffold transaction and optional MCP adapter; its documentation exrapiddwellerβ€”β€”

Interpretation history

Decision trace