Data forensics
Discover, Filter, Track, Audit
Challenges
Profiling data, finding matching patterns or related values, and quickly identifying the positions and ancestry attributes of different data sources are all ways to reveal the content of data and the manner in which it was created, deleted, or modified. Most tools available for this purpose are expensive and designed for a specific data source (e.g., a database).
After the data has been discovered and transformed, application audit trails require comprehensive information about target layouts and job flows. The details must be immediately available and secure. Logs should also track sensitive data backups and enable: user accountability, job replication, parameter modification, and issue analysis.
For example, computer forensics can expose whether a record count or a value range changes beyond a set threshold; this could indicate a data loss or fraud issue. One of the challenges with student and health data is mitigating re-identification risk and the need to measure that risk. Few data management software platforms or purpose-built applications support all of this.
Finally, in the context of a database firewall, you need the ability to log your protection policy settings, as well as all traffic and activity in a custom, queryable audit trail that is secure and subject to recovery from deletion.
Solutions
Searching & Profiling Data Sources Using state-of-the-art database, file, and dark Data Discovery Tools in the free IRI Workbench GUI For all IRI software products, you can find the location of exact (and fuzzy-matching) data patterns and automatically discover source-specific metadata that reveals file authorship and other attributes. For example, when you find PII values in databases, flat files, spreadsheets, text documents, images, and other repositories, you can also automatically display the location, owner, security, and other properties of those files.
The automated Data classification for databases and files go a step further. This wizard allows you to define data classes and groups to which you can apply global data transformation and masking rules. IRI is also in the process of adding policy control at the task, source, and field level for IAM and lineage reports.
Auditing of data masking jobs and re-identification risk: The job scripts, statistical reports, and audit logs in the IRI Voracity Data Management–Platform and its constituent IRI CoSort (SortCL)Data transformation and the IRI FieldShield data masking programs Your data layout specifications, query syntax, and manipulation details are included.
An integrated RE-ID Risk Assessment Assistant measures the statistical probability that a masked dataset can still be traced back to an individual based on the remaining quasi-identifiers in the dataset.
Read more about IRI Job Audit Logs
The XML audit log of most IRI jobs (how CoSort Data Transformation and FieldShield Data Masking) provides details for each input, inrec (virtual), and output definition – including the specified field attributes and modification functions.
The entire job script as well as user, runtime, and environment variable information are also logged in the audit log. Queries and reports on the logs are easily possible with your preferred XML parsing tool or with SortCL itself (via the included data definition files for the logs). For example, you can query file and field names, run data, and job duration. You can quickly investigate specific jobs without having to manually sift through a huge audit trail.
Phase-specific dataset counts—including the number of accepted, rejected, and processed records—along with job reconciliation details are available in optional statistical log files that are created with each SortCL script execution. An upcoming SortCL release (for CoSort, FieldShield, RowGen, NextForm, and Voracity users) will include an even more robust, granular JSON audit log with authorized access. An integrated utility assists in querying specific details from the logs, and provides extracts for visual analysis in tools like Excel and Splunk.
IRI DarkShield generates multiple audit log formats for search and masking operations (or both simultaneously). Text files, Eclipse-modeled tree editors, a dynamic dashboard, and feeds to the Splunk Enterprise Security (ES) SIEM environment in the cloud.
IRI CellShield EE creates a audit trail directly in the Excel Interchange Format (.EIF) report file, which is generated in a sheet after data discovery, and is used with the Spreadsheet Selector dialog in Excel to correct vulnerable columns.
Data and metadata lineage
Free data and Metadata Matching Functions are also available in the IRI Workbench, through the use of search tools and hubs such as EGit for sharing and securing master data and metadata in the cloud. Graphical Data Lineage and Metadata Impact Analysis for users of the IRI Voracity ETL Platform are over Quest (formerly AnalytiX DS and erwin) Mapping Manager or the Data Advantage Group MetaCenter platform is available.
A new logging system for SortCL-compatible operations (which support CoSort, Voracity, FieldShield, NextForm, and RowGen) will generate machine-readable data for lineage analysis in purpose-built platforms.
Database Activity Logs
IRI Ripcurrent recognizes changes in data and schema structure in multiple relational databases for which notifications can be configured. See this article for more information.
Related solutions
Product links