Test data for benchmarking

Large, intelligent datasets for system testing

Challenges

Evaluating the performance of hardware platforms and software applications requires the use of realistic production data. Files and tables must have the correct size and contain the right data types, file formats, record layouts and counts, as well as value distributions (data frequencies).

Standard benchmarks used by organizations such as the Transaction Processing Performance Council (TPC) may require a wide range of predetermined volumes and layouts of test data.

Creating and loading large files and tables can take a very long time without the right tools and techniques. Extracting sample data from production can be time-consuming and may violate data privacy regulations.

Test data tools like TDG or Snowfakery for Salesforce can also be difficult to use or require special programming skills (like Java or YAML). More complex test data management centers that generate synthetic test data are too expensive and not designed for the customization and speed of data volume that many system benchmarks require.

Solutions

The IRI RowGen Test Data Tool - or the IRI Voracity Data management platform that includes RowGen – can synthesize secure, large test data files – in CSV, JSON, XML, LDIF, ASN.1, COBOL, and many other structured formats (even reports) – and insert or bulk load intelligent data into relational and NoSQL database platforms.

With RowGen, you can generate a complete and consistent battery of files and tables for stress testing various software and hardware platforms. Its unique embedded data transformation functionality can also help you perform a data quality assessment or evaluate the best processing paradigms for your environment.

RowGen can create any number (and size) of files or relational tables with any number of columns in a fixed or delimited position, with over 100 different data types available. It can also automatically generate and load test data for multiple targets in various formats simultaneously.

With RowGen, you can filter or select synthetic dataset (row) and field (column) data, and even transform it to emulate production data and simulate how downstream transformation logic will affect that data. You can also decide whether to maintain or change the generated values across successive runs through random seed management.

When benchmarking database prototypes, Data Vault architectures, or Data Warehouse ETL processes is needed, RowGen considers the layout and relationships of production tables from existing DDL. It creates a batch script that you can run to quickly build and populate test DB targets that are structurally and referentially correct.

Each value in your test datasets can contain either randomly generated data or data that is randomly selected from specific files or numerical ranges to be as realistic as necessary.