Semi-structured and unstructured data

Profiling, Processing, Protection, and Presentation

Semi-structured data

Data comes in various forms from different sources, including an explosion of machine-generated data. Semi-structured data formats with flexible schemas, such as JSON and XML, are now standard formats for exchanging and storing data. However, traditional databases and data warehouses, which are based on a fixed schema, cannot easily store or process them. Therefore, they must be stored and carried in their raw form (which affects performance) or transformed before loading (loss of information while increasing complexity).

The IRI software is designed to process semi-structured data without these compromises. The CoSort data manipulation engine in the big data management platform IRI Voracity now natively processes semi-structured formats, allowing you to process this data without converting its format or creating a new external schema.

In this direct mode, Voracity users can leverage the „unparalleled parallel performance“ of CoSortMachine processing use without sacrificing functionality or flexibility. For this reason, IRI software can, in addition to large structured data also process certain classes of static and streaming semi-structured data, for example:

  • ASN.1 call detail record (CDR) files
  • IDMS, IMS and other legacy sources
  • MF-ISAM and Vision index files
  • NoSQL databases, including MongoDB (BSON), Cassandra, Elasticsearch
  • Excel, JSON, and XML files
  • NoSQL, Hive and cloud / SaaS Sources (e.g., AWS S3)
  • Internet of Things and message queues via MQTT, Kafka, MQseries, etc.
Unstructured data

You can now also in the IRI Workbench-GUI Data from unstructured Text file sources search, extract and structure – and then everything with ...the flat-file results in that environment. This means that with Voracity, you also get something like a text-based ETL tool. Additionally, it's possible to find and mask PII in unstructured data files and mask it in place or into new destinations with the same file names.

With the Dark Data Discovery Assistant in IRI Workbench, the IRI Voracity data management platform or the IRI DarkShield Data Masking Product Can users simultaneously find, mask/replace/delete, and extract (and then further process) strings based on patterns, explicit or lookup table values, machine-learned NLP models, path filters, or defined bounding box regions, across: email repositories; NoSQL DBs like Cassandra, Elasticsearch, and MongoDB; .pdf, .rtf, and MS Office (.doc/x, .ppt/x, .xls/x) documents; .txt, .xml, .html, .hl7/x12, JSON, XML, and other unstructured text and log files—as well as image files and faces—all at once.

And from the same Eclipse GUI, IRI software users can work with the flat-file extracts and their metadata:

The quintessence

The entire stock of Big Data – whether analyzed in batch processing or real-time feeds – is of great interest to companies and government service providers. IRI software – and in particular Voracity Total Data Management Platform – the fastest, easiest, and most cost-effective way to integrate and prepare structured, semi-structured, and unstructured data sources…. into your existing IT infrastructureProcessingProtection and Provision).