Back
Can enterprises protect data without making AI less reliable?
SiTech AI Team2 წთ. საკითხავი

Can enterprises protect data without making AI less reliable?

A new Perforce Delphix report finds that privacy controls often make production-grade data harder to use: 26% of organizations struggle to obtain it, and 25% cannot preserve the relationships between records. More than half cite data quality challenges.

Enterprises pouring money into AI are running into an unexpected bottleneck: obtaining data that is both protected and useful. In a sponsored post on The New Stack, Mayank Ahluwalia, a senior product manager at Perforce Delphix, describes how privacy initiatives can make the datasets that AI and engineering teams depend on harder to reach, less representative of real-world conditions — or unable to preserve the relationships between records.

The numbers behind the bottleneck

The argument draws on Perforce Delphix's "2026 State of AI and Data Privacy Report". Among surveyed organizations, 26% said privacy controls make production-quality data harder to obtain, and 25% reported struggling to preserve relationships across data entities. More than half — 51% — cited data quality challenges. The most striking figure: 84% of respondents maintain a data privacy exception in their non-production environments, a sign that companies still trade compliance against innovation instead of reconciling the two.

"Protecting data isn't enough if it can no longer support the systems that depend on it," Ahluwalia writes. Low-quality datasets, he notes, lead to inaccurate analytics, poorly trained models, incomplete test coverage, more rework and delayed releases.

Why referential integrity matters

Masking sensitive fields is only part of the problem. If the relationships connecting records break during protection, the data no longer resembles reality: a customer, order or payment record may still exist, but billing validation spanning several database tables can no longer produce correct results. Analytics pipelines depend on consistent identifiers to join information across sources, machine learning workflows need complete business context, and software testing needs realistic relationships between records.

These failures are hard to detect, because pipelines keep running while quietly producing degraded outcomes. That is why a quarter of surveyed organizations flagged preserving relationships across data entities as a significant challenge.

From compliance metrics to a data strategy

The post argues privacy initiatives should be measured on data quality, realism, referential integrity, accessibility and provisioning speed — not compliance metrics alone. It recommends designing for governance by default, mapping the data lifecycle from request and discovery to reuse and retirement, and adopting a portfolio approach: data virtualization for speed, masking for security and synthetic data for coverage.

"In the AI era, trusted data, not fast model output, may be the ultimate competitive advantage," Ahluwalia concludes.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.