Virtualization of Test Data

Create, deploy, and manage virtual test data

What is test data virtualization?

Test data virtualization combines test data generation—whether through masking, subsetting, or synthesis—with efficient test data provisioning. The key word is efficiency, as most test data provisioning today still involves creating many physical copies of test tables or files, rather than fewer, more consistent, up-to-date golden copies of test data that can be easily accessed and reset.

Manual, database-centric, and off-the-shelf approaches to test data virtualization have proven to be time-consuming and costly, leading to insufficient testing or database-centric approaches that delay SLA/delivery dates. And testing with the latest, still unmasked production data as a fallback is simply unsafe.

The ultimate goal is to provide dynamic, more or less on-demand (self-service) test data for database and software application development, system testing, and outsourcing. The creation and management of virtual test environments with secure, intelligent test data remains a tedious part of quality assurance and development cycles.

Solutions

By leveraging the long-standing capabilities of IRI RowGen-Use tools for generating synthetic test data and for subsetting – or the FieldShield and DarkShield data masking tools, which are also in the IRI Voracity-data management platform are included – you can meet multiple test data management requirements. You can also meet many of your test data provisioning requirements through virtual test environments without the cost or complexity associated with commercial test data virtualization solutions.

One of the inherent advantages of the IRI Voracity data management platform is the combination of robust data integration, test data generation, and Data replication features. Together, they enable you to quickly and easily create and deploy customized, virtual test data solutions for DevOps.

Voracity can combine static and streaming ETL or real-time incremental database replication with data masking, subsetting, synthesis, data transformation, and custom formatting. Without impacting live systems or being limited to a specific database or cloud platform, Voracity users can leverage and automate the capture, manipulation, and deployment of both ad-hoc (virtual) and persistent test data sets that: reflect production data characteristics, preserve data and referential integrity enterprise-wide (not just database-wide), anonymize PII, and don't become stale.

Suggestions

First, consider under which business rules you need an ad hoc solution. IRI will give you in this article series for test data management advice on how to consider these rules and provides you with various functions available, that will help you with the data you need to work with, found in such sources, i.e., in files, databases, and dark data documents.

Next, you should consider what kind of test data you need, depending on who needs it and how and where it will be used. You may need to be creative; some testing goals benefit from a combination of data masking and synthesis. like this. Or you want to mask data to generate realistic test data, while you

  • a subset from a database environment how here, or a replication how here
  • The integration of SQL and file sources how here, or the preview of ETL jobs how here
  • Integrating into a DevOps pipeline (CI/CD) for test automation how here
  • Real-time refresh of a virtual test database how here
  • Streaming from an IoT data broker, how here

Also note that each Voracity data generation process allows you to define multiple, differently formatted persistent and virtual targets simultaneously. This efficiency and flexibility is particularly valuable for DevOps- Teams that need to work in parallel.

Once techniques and goals are established, you can also choose how to design, modify and/or release the job(s), and how and where you want to run them. Voracity supports multiple job design and runtime methods; see the IRI Workbench section on this page.

Additional benefits

Unlike other virtual TDM solutions, IRI doesn't require you to clone databases, set up a virtual TDM appliance, or do anything else that's complex (or expensive). Test data engineers can deploy as many persistent or virtual copies as they need and instantly populate their testers' repositories as test data is generated. However, if you want a fully masked or synthetic database clone, IRI FieldShield and RowGen jobs can be executed as scripts that simultaneously ActifioCommvault– and Windocks-Operations (virtualized container images) can be called!

IRI-Subsetting, Masking, and Synthesis jobs for structured data are also supported in Cigniti and Value Lab TDM Portals, enabling you to create and manage on-demand test datasets for file, DB, and API targets. For TDM with semi-structured (e.g., HL7, JSON, XML) and unstructured text or file sources (e.g., PDF, MS Office, image data), you can IRI DarkShield use to mask them or replace real values in them with test data generated by IRI RowGen; See this article.

Finally, controlling test data can be just as important as controlling your production data. In addition to the inherent data security governance in the many static data masking functions from Voracity, allow you multiple Data Quality Functions the validation and stabilization of your test data sets, whether virtual or not. Workflow diagrams and the automatic generation of batch files support the graphical design of independent and dependent work chains. Furthermore, several options for the Data and metadata sequence supported, so you can track changes to the source data and your test data projects.