Sort / Merge Operations
Fast sorting and merging of large, structured data
Challenges
Sorting remains a critical component of data processing. Data sorting is part of:
Database load, index, and query/search operations
· Data Warehouse Sorting, Linking, and Aggregation of Transformations (ETL jobs)
Reporting, analysis, and testing environments
However, with increasing data source sizes, from hundreds of megabytes to terabytes and beyond, sorting can cause an exponential demand for computing resources. Most standalone data sorting tools and techniques, as well as many data merging methods, are simply not scalable to meet the demands of Big Data.
The sorting functions in databases, ETL and BI reporting tools, operating systems, and compilers are also not designed for Big Data. The old sort/merge programs are expensive to operate, use cryptic JCL syntax, and are limited in their functionality.
This diagram summarizes some of the problems that occur in the market for large-scale data sorting tools:
Robustness issues | Management Concerns |
Sorting speed and scalability in volume | Sorting and associated functionality |
Support for data and file types | Simplicity of GUI and/or Parm syntax |
Event Monitoring and Debugging | Logging and Metadata Frameworks |
Performance Optimization and Logging | Pricing and Licensing Models |
Plugin compatibility or accuracy of Parm conversion | Technical Support Speed |
Interoperability of third-party hardware and software | Supplier Capabilities and Reputation |
Implementation Paradigm | Skills gap (e.g., Hadoop), maintenance costs |
Solutions
With the increasing volume of data, the value of IRI CoSort. CoSort is the world's first commercial sort and merge package for use on „open systems“ and has been a leader in commercial sorting since 1978, and is a proven product:
Unix file sorting program
Windows Sorting Program
· ETL, BI and Database Sorting Alternative
Mainframe JCL Sorting
Consolidation Replacement
with state-of-the-art performance, industry-leading functionality, and the most familiar, intuitive user interfaces…. and without additional hardware, Hadoop, in-memory DBs, or appliances.
CoSort sorts any number, size, and type of structured fields, keys, records, and files – including mainframe binary files, IP addresses, Asian multibyte characters, Unicode,… The CoSort Engine scales linearly with volume and allows for granular tuning of CPU, memory, disk, and related resources. Multi-gigabyte sorting in seconds on multi-CPU servers.
___________________________________________________________________________________
127,268,900 lines * 405 bytes/line = 51.5GB input file
CoSort Job Time w/ 20-byte sort key @ 131 seconds = 2 minutes and 11 seconds
Platform: x86 Linux Development Server with 32 of 64 cores in use
___________________________________________________________________________________
CoSort can also replace or convert third-party sort functionalities with proven libraries, tools, or services – saving time and money on batch operations and integrated applications. Ask about special incentives for migrating from an older sort product and for discounts on embedded sales.
Sorting is just the beginning
CoSort also offers the unique capability to simultaneously protect data transform, to migrate, to report and to protect. The CoSort Sort Sort Control Language (SortCLcombines these functions in the same job script and I/O pass. Assign multiple sources to multiple destinations and formats, while You are sorting.
SortCL is just one of several interfaces in the CoSort package available for standalone or integrated sort/merge operations. All sort and transform jobs can be in the IRI Workbench GUI, which is based on Eclipse™, are planned, monitored, logged, audited, and otherwise managed.
Beyond the CoSort package, the same SortCL-driven operations are also an integral part of CoSort – including the IRI Voracity-Data management platform that performs and combines large-scale data discovery, integration, migration, management, and analysis. In Voracity, the CoSort sort engine (and SortCL scripts) are automatically used in (and for): ETL, change data capture, DB subsetting, pseudonymization, synthetic test data, data wrangling, and bulk DB loading.