Accelerate or abandon DataStage

Improve your ETL job performance

Challenges

Even after consultation and tuning, large amounts of data (i.e., more than one million rows) can only be transformed slowly, especially without an expensive hardware or version upgrade of DataStage.

ETL performance bottlenecks include large sorts, joins, aggregations, loads, and sometimes unloads. Parallelization or optimization in other layers or tools can be cumbersome if not expensive and can impact performance for other users.

From a security perspective, IBM's data masking solutions can be expensive or cumbersome for some, or not provide all the functionality for PII detection or privacy protection for others.

Solutions

Accelerate sorts, joins, and aggregations in DataStage with a one-pass operation by calling CoSort Sort Control LanguageSortCL) in a sequential file stage or a Subroutine before job routine. Perform large data transformations without burdening other jobs in DataStage, your database, or your BI tool. Also specify file format and data type conversions, field-level masking functions, custom reports, and pre-sorted load files.

Improve DataStage performance by adding a sequential file stage before the aggregation, running a SortCL job to externally pre-sort the file by the partitioning keys, and then defining the sorted fields in the aggregation stage.

Data stored in tables and flat files within DataStage can be sensitive and may contain personally identifiable information subject to confidentiality restrictions and data privacy laws. SortCL, which is CoSort licensed – or via compatible IRI FieldShield-Data masking products or IRI VoracityData Management (and ETL) Platform Operations – can column/field values in any ODBC-connected database or standalone flat fileWhat protect.

Your business rules determine the function you choose for each column, i.e., format-preserving AES-256, FIPS-compliant OpenSSL, 3DES and/orGPG encryption, lookup value substitution (pseudonymization), character masking, hashing, redaction, custom expression logic, substring, or user field function.

IRI Voracity generated by its constitutive (or intrinsic) IRI RowGen-Software product secure, realistic test data using COBOL or CoSort-Metadata, .dsx-defined files, and all JDBC-connected RDB data models. Use RowGen to create compliant, realistic test data from random generation and/or set file selection, customizing it further with built-in data manipulation and formatting functions. Voracity also includes database subsetting and masking for testing in lower environments.

Streamline the migration of DataStage to a faster, more cost-effective ETL operation in IRI Voracity with Erwin Mapping Manager (Analytix DS) or Code Automation Frameworks (CATfx). This proven technology, along with ADS Lite Speed Conversion Services, finally gives ETL architects and the CIO/CFO suite the ability to save hundreds of thousands of euros immediately and transition to cost-effective operating costs in the future.

Other Resources