Syntho
Syntho generates privacy-safe, production-like test datasets when real customer or patient records are too sensitive or slow to share. The Amsterdam company packs data masking, rule-based generation, and statistical synthesis into one engine, so you pick or mix methods per table instead of juggling separate tools.
Where many vendors sell one synthesis technique, Syntho lets you run masking, formula-driven rules, and statistical twins in a single job. That matters when one column needs deterministic mock values and another needs statistical realism for analytics. Licensing is feature-tiered with unlimited workspaces and no per-generation consumption charges.
The Syntho Engine deploys via Docker inside your own infrastructure, and Syntho never receives your source records during synthesis. Data science, QA, and data engineering teams in finance, healthcare, software vendors, and public organizations use it for compliant test data, analytics sandboxes, tailored demos, and secure sharing. Customers named on the site include Philips, KPN, and Nederlandse Spoorwegen.
You connect source and target databases through built-in connectors, run jobs from the web UI or REST API, and automate recurring generation workflows. A quality assurance report validates privacy and fidelity with three industry-standard metrics.
Three synthesis modes in one platform: masking, rule-based, and statistical generation
Self-hosted Docker deployment with 20+ database and 20+ filesystem connectors
200+ mock data generators plus PII column and open-text scanners
Preserves referential integrity across multi-table databases with consistent mapping
QA report with three industry-standard privacy metrics for synthetic output validation
Typical starting hardware: 12 to 20 vCPUs, 32GB RAM, and 128GB storage
Tables under 1 million rows typically synthesize in under 5 minutes
Combines masking, rule-based, and statistical synthesis in one platform instead of separate tools.
On-premise Docker deployment keeps sensitive source data inside your own environment.
Feature-based licensing with no per-generation consumption fees and unlimited workspaces.
QA reports include three industry-standard privacy metrics for validating synthetic output.
Preserves referential integrity across complex multi-table relational databases automatically.
No public pricing; all plans require a sales quote and annual license agreement.
Enterprise deployment expects Docker or Kubernetes infrastructure on your side.
Typical starting hardware is 12 to 20 vCPUs and 32GB RAM, which adds infra cost beyond the license.
Does Syntho see or process my data?
No. Syntho deploys the Syntho Engine inside your trusted environment via Docker, so your source data stays on your infrastructure. Syntho does not receive, view, or process your records during synthesis.
What deployment options does Syntho support?
Syntho ships as a Docker container deployable on-premise or in your private cloud via Docker Compose or Kubernetes. The platform offers a web UI and a REST API for pipeline integration.
What databases does Syntho connect to?
Syntho supports PostgreSQL, SQL Server, Oracle, MySQL, Databricks, IBM DB2, MariaDB, Azure Data Lake, and Amazon S3, with 20+ database and filesystem connectors documented on the pricing page.
How does Syntho pricing work?
Syntho uses feature-based annual licensing with Basic, Standard, and Ultimate tiers quoted on request. Plans include unlimited workspaces, no per-generation consumption fees, and support bundled in the license.
What data types can Syntho synthesize?
Syntho works best on structured tabular data including categorical and numerical columns, PII, geographic data, time series, multi-table databases with referential integrity, and open text fields.
How long does Syntho take to generate synthetic data?
Generation time depends on database size. Syntho states that a table with fewer than 1 million records is typically synthesized in under 5 minutes.

