Building AI You Can Trust Starts With Data You Can Defend 

Why responsible data infrastructure is the missing layer in AI security  By the CyberSG R&D Programme Office, with BetterData and Data Illusion April 2026 

 

The AI Security Conversation Is Starting In The Wrong Place 

Most conversations about AI security centre on the model — its tendency to hallucinate, its embedded biases, the guardrails needed to keep it within acceptable boundaries. These are real and important concerns. But they share a common blind spot: they start too late. 

By the time an AI model is being evaluated for security, the most consequential decisions have already been made. They were made in the data. What was collected, what was excluded, how it was stored, who could access it, and whether it could be shared at all — these choices shape everything a model can and cannot do, long before a single parameter is trained. 

The Data Dilemma 

The relationship between AI and data is straightforward in principle: better data produces better models. In practice, it creates a structural problem that most organisations have not adequately resolved. 

The data that would most improve AI models in high-value sectors — patient records in healthcare, transaction histories in finance, case files in government — is precisely the data that is most restricted. It carries legal obligations, privacy risks, and reputational exposure that make sharing it, in its raw form, effectively impossible. 

Organisations respond to this in predictable ways. They work with whatever data they can access, which is often incomplete, outdated, or unrepresentative. They anonymise datasets and treat them as safe, without examining whether anonymisation actually holds. They build data silos and accept that collaboration is too risky, forgoing insights that could only emerge across organisational boundaries. 

Each of these responses is understandable. None of them is adequate. And each reflects the same underlying confusion: treating data security and AI capability as opposing forces, when they are in fact two sides of the same problem. 

 

Why “Safe Enough” Is No Longer Safe Enough 

The most widespread assumption in data governance is that removing identifying information from a dataset makes it safe to use and share. Anonymisation has become a default step in governance frameworks — treated as a reliable control rather than a partial measure. 

The deeper problem runs further than anonymisation. In the race to develop robust AI models, organisations often misunderstand the relationship between data security and data utility. A prevailing misconception is that to keep sensitive data — such as healthcare records, financial models, or government datasets — secure, it must remain absolutely isolated within organisational boundaries. While well-intentioned, this lock-and-key approach creates rigid data silos, starving AI algorithms of the high-value, real-world data they need to evolve and deliver meaningful insights. 

The fundamental problem is not the desire for security, but the traditional belief that data ownership cannot be decoupled from data usage. In the AI era, forcing a binary choice between absolute isolation and risky data sharing is a false dichotomy — and one that is quietly costing organisations their competitive edge and stalling cross-industry innovation. 

“The question is not how tightly data can be locked away. It is how securely it can be put to work.” 

The solution lies in shifting the paradigm: from moving data, to moving algorithms. By establishing a Trusted Data Space — such as a Data Clean Room built on zero-trust principles and kernel-level sandboxing — organisations can create secure environments where AI models compute and extract necessary insights without ever exposing or transferring the underlying raw data. The governing principle is straightforward: data stays, programs move. 

This approach enables enterprises and government bodies to collaborate on critical AI initiatives — from medical research to financial modelling — with full confidence in data sovereignty. Federated data infrastructure and clean rooms do not weaken security. They redefine what security makes possible. 

 

Building AI Before the Data Exists 

Not every organisation has the luxury of waiting for the right data to exist. In sectors where production data is restricted, legally sensitive, or insufficient in volume, the prevailing assumption is that AI development must wait too. It does not have to. 

Synthetic data has crossed a meaningful threshold. Modern generative approaches can produce datasets that preserve the statistical properties, distributions, and relationships present in real data, without any generated record corresponding to an actual individual, customer, or transaction. Organisations no longer need to choose between model quality and data safety. 

For AI development teams, the question is less whether a customer record is “real” and more whether it teaches the model the right behaviour. Good synthetic data offers the same statistics that real data has, but in a safer form. It keeps the patterns that matter: how customers behave, what type of customer profiles are valid, and where rare cases sit. Privacy controls such as differential privacy can tighten that separation further, while programmable generation rules can rebalance underrepresented groups or introduce scenarios an organisation needs to test. 

This matters because AI systems are often needed before the ideal dataset exists. A bank launching a new card product, for example, cannot wait a year for enough fraud cases before building alerting dashboards or validation processes. It can start with related products, generate plausible transaction patterns, then adjust those patterns toward the expected risk profile of the new product. Where even adjacent data is thin, tabular foundation models can now help produce a first statistically coherent dataset because they have learned common structures across many other datasets in the industry. 

The technology has crossed a threshold because modern models can preserve relationships in real data that older masking and sampling methods cannot. Synthetic data still needs validation against real outcomes when they arrive. But when a model trained or tested on synthetic data behaves close to the real data benchmark, with far less privacy exposure, it becomes a serious and legitimate way to build trustworthy AI earlier. 

 When Security Enables Development 

Perhaps the most damaging assumption is that security teams and AI development teams are fundamentally at odds. One protects data. The other demands access to it. The negotiation that follows typically produces bad outcomes for both — concessions that compromise security, and constraints that compromise capability. Data security and AI development only feel opposed when raw data movement is the default answer. A more mature approach starts with the purpose of the work. 

If an AI model needs to learn from sensitive records that should not leave their source, trusted compute environments can bring algorithms to the data. If development teams need realistic data for building, testing, or collaboration, synthetic data can carry the statistical behaviour of production data without carrying the same exposure. These are not alternative approaches — they are complementary, each operating at a different stage of the same pipeline. 

Together, they give organisations more than protection. They give them options. Security teams can maintain control of sensitive data while AI teams get credible evidence that their models and pipelines work. That is the practical foundation for trustworthy AI: not maximum access, and not total isolation, but the right form of access for the risk at hand. 

Putting It Together 

A responsible data infrastructure for AI does not begin at model training. It begins when an organisation first decides what data it holds, how it governs access, and under what conditions it can be used. 

Federated data infrastructure and clean rooms address the foundational layer — creating environments in which sensitive data can be computed on without being exposed, enabling collaboration between organisations that could never share raw data. Synthetic data addresses what comes next: once governance over real data is established, synthetic generation creates privacy-preserving proxies that allow development, testing, and validation to proceed without production data present. 

The answer to where synthetic data belongs in the development lifecycle is often earlier than teams expect. Production-grade AI moves through development, UAT, staging, and production — but only production typically sees production data. That leaves teams testing pipelines on stale extracts or small sanitised samples. The tests pass, yet the first real run exposes missing values, skewed categories, and edge cases that were never properly tested. 

Synthetic data closes that gap without requiring every environment to be a copy of production. Teams can generate privacy-preserving datasets for migrations, feature pipelines, monitoring dashboards, and model validation. Enterprise data is not one homogeneous thing — it is customers linked to accounts, devices, claims, events, and transactions. Synthetic generation that preserves those relational links does not just produce safer training data. It produces a dependable foundation for testing every system that carries AI into production. 

The sequence matters: clean room infrastructure determines what can be done with sensitive data in place. Synthetic data determines what can be done without it. Together, they represent a complete and defensible approach to AI data security. 

 The Question Worth Asking 

The organisations building the most trustworthy AI are not just asking whether their model is safe. They are asking whether their data infrastructure was ready before the model was built. 

That is a harder question, and it does not have a single answer. But it is the right question — and it is one that security teams, AI development teams, and the executives who oversee both should be asking together, before the next model reaches production. 

The capabilities to answer it exist. The challenge now is ensuring that the organisations who need them most understand what is possible, and are equipped to build it. 

 About the Contributors 

BetterData develops synthetic data generation technology for AI training and testing, enabling organisations to build and validate AI systems without direct exposure to sensitive production data. 

Data Illusion builds federated data infrastructure and data clean room environments, allowing organisations to collaborate on AI initiatives while maintaining full sovereignty over their underlying data.