Sort / Merge Operations

Fast sorting and merging of large, structured data

Challenges

Sorting remains a critical component of data processing. Data sorting is part of:

Database load, index, and query/search operations

· Data Warehouse Sorting, Linking, and Aggregation of Transformations (ETL jobs)

Reporting, analysis, and testing environments

However, with increasing data source sizes, from hundreds of megabytes to terabytes and beyond, sorting can cause an exponential demand for computing resources. Most standalone data sorting tools and techniques, as well as many data merging methods, are simply not scalable to meet the demands of Big Data.

The sorting functions in databases, ETL and BI reporting tools, operating systems, and compilers are also not designed for Big Data. The old sort/merge programs are expensive to operate, use cryptic JCL syntax, and are limited in their functionality.

This diagram summarizes some of the problems that occur in the market for large-scale data sorting tools:

Robustness issues

Management Concerns

Sorting speed and scalability in volume

Sorting and associated functionality

Support for data and file types

Simplicity of GUI and/or Parm syntax

Event Monitoring and Debugging

Logging and Metadata Frameworks

Performance Optimization and Logging

Pricing and Licensing Models

Plugin compatibility or accuracy of Parm conversion

Technical Support Speed

Interoperability of third-party hardware and software

Supplier Capabilities and Reputation

Implementation Paradigm

Skills gap (e.g., Hadoop), maintenance costs

Solutions

With the increasing volume of data, the value of IRI CoSort. CoSort is the world's first commercial sort and merge package for use on „open systems“ and has been a leader in commercial sorting since 1978, and is a proven product:

Unix file sorting program

Windows Sorting Program

· ETL, BI and Database Sorting Alternative

Mainframe JCL Sorting

Consolidation Replacement

with state-of-the-art performance, industry-leading functionality, and the most familiar, intuitive user interfaces…. and without additional hardware, Hadoop, in-memory DBs, or appliances.

CoSort sorts any number, size, and type of structured fields, keys, records, and files – including mainframe binary files, IP addresses, Asian multibyte characters, Unicode,… The CoSort Engine scales linearly with volume and allows for granular tuning of CPU, memory, disk, and related resources. Multi-gigabyte sorting in seconds on multi-CPU servers.

___________________________________________________________________________________

127,268,900 lines * 405 bytes/line = 51.5GB input file

CoSort Job Time w/ 20-byte sort key @ 131 seconds = 2 minutes and 11 seconds

Platform: x86 Linux Development Server with 32 of 64 cores in use

___________________________________________________________________________________

CoSort can also replace or convert third-party sort functionalities with proven libraries, tools, or services – saving time and money on batch operations and integrated applications. Ask about special incentives for migrating from an older sort product and for discounts on embedded sales.

Sorting is just the beginning

CoSort also offers the unique capability to simultaneously protect data transformto migrateto report and to protect. The CoSort Sort Sort Control Language (SortCLcombines these functions in the same job script and I/O pass. Assign multiple sources to multiple destinations and formats, while You are sorting.

SortCL is just one of several interfaces in the CoSort package available for standalone or integrated sort/merge operations. All sort and transform jobs can be in the IRI Workbench GUI, which is based on Eclipse™, are planned, monitored, logged, audited, and otherwise managed.

Beyond the CoSort package, the same SortCL-driven operations are also an integral part of CoSort – including the IRI Voracity-Data management platform that performs and combines large-scale data discovery, integration, migration, management, and analysis. In Voracity, the CoSort sort engine (and SortCL scripts) are automatically used in (and for): ETL, change data capture, DB subsetting, pseudonymization, synthetic test data, data wrangling, and bulk DB loading.