Improving data quality

Solutions for data enrichment and data cleansing

Ensuring and improving data quality are essential components of data integration, data management and analytics. Why? Simply put, „garbage in, garbage out.“

In addition to normal errors and duplication of work, company data is subject to constant change. For example, in the US alone, 240 companies change addresses and 5,769 people change jobs every hour. Your employees need to be able to trust the data sources for their reports and applications.

According to Gartner, in 2017, „33% of Fortune 100 organizations will face an information crisis due to their inability to effectively use, control and trust their enterprise information.“ 36% of the participants in the Gartner study estimated that they lose more than 1 million dollars annually due to data quality issues, while 35% were unable to estimate the cost impact.

 

Users of the overall data management platform IRI Voracity or its component products such as IRI CoSort can control data quality and analyze data in many different sources in different ways clean up:

Learn more in this article by Bloor Research on improving data quality with the IRI Voracity platform software.

Capability

Options

Profile & Classify

Discover and analyze You sources in data viewer tools and the metadata discovery wizard. These functions, together with the wizards for profiling flat files, databases and dark data in the IRI Workbench (Eclipse GUI) allow you to find data values that exactly match (literal, pattern or lookup) these values, or fuzzy match (up to a probability threshold). Output reports are provided in CSV format and extracted dark data values are packed into flat files. New classification functions allow you to apply transformation rules (and masking rules) to data categories.

Bulk filter

Remove unwanted rows, columns and duplicate data records with the same sort keys in the CoSort / Voracity program SortCL. Identify, remove or isolate bad values with a special selection logic. To this page you will find further information.

Validate

Use the SortCL at field level „if-then-else logic“ and „iscompare“, to isolate null values and incorrect data formats. Use „Outer Joins“ to obtain silo source values, that do not match master (reference) data records. Use data formatting templates and their date validation options to e.g. the correctness of input days and dates.

Standardize

Use the wizard to Data standardization in consolidation style (MDM) in IRI Voracity to find and evaluate data similarities and eliminate redundancies. Sort the remaining master data values into files or tables. Another wizard can transfer master values back to your original sources, and a pending registry hub supports a reporting front end for finding data that has been searched in different silos.

Replace

Specify one-to-one replacement via pattern matching functions or create multiple values in sets that are used for many-to-one mappings.

De-duplicate

Eliminate duplicate rows with the same keys in SortCL orders.

Clean up

Specify custom, complex include/omit conditions in SortCL based on data values. To this page you will find further information.

Enrich

Combine, sort, join, aggregate, search and segment data from multiple sources to improve row and column detail in SortCL. Create new data forms and layouts through conversions, calculations and expressions. Improve layouts through remapping and templating (composite formats), see IRI NextForm. Create additional or new test data for extrapolation with IRI RowGen.

Extended DQ

Field-level integration with SortCL for Trillium and Mellissa data standardization APIs, etc.

Generate

Use RowGen to create good and bad data, including more realistic Values and formats, valid days and dates, national ID numbers, master data formats, etc.