Improving data quality
Solutions for data enrichment and data cleansing
Ensuring and improving data quality are essential components of data integration, data management and analytics. Why? Simply put, „garbage in, garbage out.“
In addition to normal errors and duplication of work, company data is subject to constant change. For example, in the US alone, 240 companies change addresses and 5,769 people change jobs every hour. Your employees need to be able to trust the data sources for their reports and applications.
According to Gartner, in 2017, „33% of Fortune 100 organizations will face an information crisis due to their inability to effectively use, control and trust their enterprise information.“ 36% of the participants in the Gartner study estimated that they lose more than 1 million dollars annually due to data quality issues, while 35% were unable to estimate the cost impact.
Learn more in this article by Bloor Research on improving data quality with the IRI Voracity platform software.
Capability | Options |
Profile & Classify | Discover and analyze You sources in data viewer tools and the metadata discovery wizard. These functions, together with the wizards for profiling flat files, databases and dark data in the IRI Workbench (Eclipse GUI) allow you to find data values that exactly match (literal, pattern or lookup) these values, or fuzzy match (up to a probability threshold). Output reports are provided in CSV format and extracted dark data values are packed into flat files. New classification functions allow you to apply transformation rules (and masking rules) to data categories. |
Bulk filter | Remove unwanted rows, columns and duplicate data records with the same sort keys in the CoSort / Voracity program SortCL. Identify, remove or isolate bad values with a special selection logic. To this page you will find further information. |
Validate | Use the SortCL at field level „if-then-else logic“ and „iscompare“, to isolate null values and incorrect data formats. Use „Outer Joins“ to obtain silo source values, that do not match master (reference) data records. Use data formatting templates and their date validation options to e.g. the correctness of input days and dates. |
Standardize | Use the wizard to Data standardization in consolidation style (MDM) in IRI Voracity to find and evaluate data similarities and eliminate redundancies. Sort the remaining master data values into files or tables. Another wizard can transfer master values back to your original sources, and a pending registry hub supports a reporting front end for finding data that has been searched in different silos. |
Replace | Specify one-to-one replacement via pattern matching functions or create multiple values in sets that are used for many-to-one mappings. |
De-duplicate | Eliminate duplicate rows with the same keys in SortCL orders. |
Clean up | Specify custom, complex include/omit conditions in SortCL based on data values. To this page you will find further information. |
Enrich | Combine, sort, join, aggregate, search and segment data from multiple sources to improve row and column detail in SortCL. Create new data forms and layouts through conversions, calculations and expressions. Improve layouts through remapping and templating (composite formats), see IRI NextForm. Create additional or new test data for extrapolation with IRI RowGen. |
Extended DQ | Field-level integration with SortCL for Trillium and Mellissa data standardization APIs, etc. |
Generate | Use RowGen to create good and bad data, including more realistic Values and formats, valid days and dates, national ID numbers, master data formats, etc. |