Database subsetting

Test with smaller, masked schemas

Challenges

You need realistic test data files and the creation of test reports:

  • Develop and maintain JCL and COBOL programs on the mainframe
  • create and simulate new BI and analytics applications
  • Evaluation of CRM, ERP, EAI, EII, and Other Systems
  • Prototype ETL operation and bulk load scenarios
  • Benchmarking new hardware or software platforms

The data you have in production files or reports may be inaccessible, restricted by company policies or privacy laws, or simply not exist. And if you have access to test data files or test reports, they may not contain the scope or volume of data that will be needed in the future.

Solutions

In addition to the powerful database parsing, generation, and population capabilities, which IRI RowGen To provide structurally and referentially correct test data, you can now also generate (and mask) referentially intact database subsets from standard relational sources as well as from complex and/or very large databases (VLDBs).

A proven, high-performance Database Subsetting Assistant for relational databases is in the IRI Workbench, the Eclipse IDE for IRI Voracity Test Data Management Platform or the test data security tools of the IRI Data Protector Suite (IRI FieldShield (for data masking and/or RowGen for test data generation). This ergonomic test data sub-setting utility allows you to quickly create custom subsets of manageable, referentially correct data determined by your master table, while simultaneously applying consistent data masking and/or mapping rules across all sub-tables.

It is also possible to retrieve data from individual To selectively partition and mask tables on an ad-hoc basis for testing purposes, using the same IRI Metadata Framework used. To do this, you simply write a FieldShield job script with the table details or use a wizard to automatically create such a script, and then either insert SQL SELECT syntax directly into the input area or create and use custom /INCOLLECT row filters and/or qualitative /INCLUDE or /OMIT statements to define the size and content of each subset.

You can use DB subsets in various ways for test data users provide, which work on-premises or in the cloud. These options for managing database subsets include: new persistent or virtual (federated) test schemas, flat-file targets, and DevOps pipelinesSee this example).

In addition to database subsets, you can also Test file subsets create. Use RowGen to generate synthetic, highly realistic test files in any format and create for any size. Use FieldShield to extract (and mask) test data subsets from structured (flat) files in fixed position or delimited formats using the built-in Selection and Filter Functions utility. FieldShield is based on the same big data engine as Voracity (IRI CoSort SortCL program) and can therefore also handle structured files with any volume process.

Subsetting strategies like these not only minimize the risk of PII disclosure and violations of data privacy laws, but also drastically reduce the costs of database and application testing infrastructures—some talk of up to $50,000 per database. Here Learn how to automatically set up and create test data subsetting jobs in Workbench.