Big data wrangling
Fast Data Preparation for BI & Analytics
In a world of big data, time to insight, data quality, and data security are paramount.
The larger the source data, the greater the challenges in preparing it for today's Business Intelligence (BI) and analytics tools:
Volume
Regarding performance, BI and analytics tools cannot handle large datasets. If they work at all, it takes a long time for the results to be displayed.
Diversity
Speed
BI tools typically cannot handle real-time or mediated data without pre-processing. And if possible, feeding individual displays rather than a regulated environment creates redundancy and uncertainty.
Truthfulness
Data protection
Complexity
Designing and modifying report and dashboard layouts are already complicated. Integrating data from different sources for each new report is another challenge for data and ETL architects, DBAs, and business users. Data silos also pose challenges for data quality, security, storage, and synchronization.
Is there a way to solve all these problems at once… integrate, clean, and mask all this data so your analysis or data visualization tool can use it?
Yes, this provision of reusable data for the use and display of BI and analysis tools is called Data Blending, Data preparation, Data Franchising, referred to as Data Munging or Data Wrangling. This is such an important process that several VC-backed tools have recently been launched just to tackle this challenge.
Data integration, cleansing, and masking all run simultaneously in a consolidated SortCL-Program (of the standard CoSort and Voracity engine). Use it to quickly and reliably differentiate Data sources to manage for use and reuse by your BI or analytics platform. Choose from multiple data preparation options in the free Eclipse™ GUI for Voracity and process this data in Windows, Unix, or Hadoop file systems without purchasing additional hardware or impacting a database.
Data Wrangling
During data preparation or franchising, different data sources are captured, filtered, de-normalized, sorted, aggregated, protected, and reformatted. This approach allows your BI tool to import only the data it needs, in the table or flat file format (e.g., CSV, XML) it requires.
Data visualizations – and therefore answers to your business questions – are generated faster when you use Voracity or CoSort:
Filter, scrub, sort, merge, aggregate, and otherwise transform large data in a single job script and I/O pass.
Create subsets that can require and process dashboards, scatter plots, scorecards, or other analytical tools.
Centralized data preparation also avoids the reproduction or synchronization of data when another report is needed.
Data protection (masking)
Deidentify BI and analytics applications fed with PII with integrated field-level anonymization features such as:
Encoding
Encryption (format-preserving or not)
Expressions
Hashing
Masking (Obfuscation)
Randomization
editorial revision
Substring Manipulation
Apply the desired function – using data classes and rules – based on appearance, reversibility, and authorization.
Did you know that?
The free graphical IDE for job design across all IRI software products is called IRI Workbench. The IRI Workbench, based on Eclipse™, supports:- automatic data profiling, classification, ERD, and metadata creation
- Generation of job scripts (or flow) with multiple modification methods
- Batch, Remote, and HDFS Executable Shortcuts
- Data, metadata, and job version control
- Master data management