Harvest Now, Decrypt Later and AI Infrastructure: Why Training Data Is Newly Exposed

Illuminated server racks in a modern data center powering AI infrastructure

Harvest now, decrypt later isn’t a new risk, but in May 2026 the Cloud Security Alliance’s AI Safety Initiative gave it a sharper target. Its paper, “Harvest Now, Decrypt Later: Quantum Risk to AI Infrastructure,” argues that AI systems aren’t just another category of data worth stealing – they’re a concentrated one, and that concentration is exactly what makes them worth an adversary’s patience.

Why AI infrastructure concentrates the risk

CSA’s framing is direct: AI systems are built on what amounts to encrypted gold mines – proprietary model weights, massive training datasets, and the ongoing communication streams that connect increasingly complex, multi-agent systems. Three categories in particular carry HNDL exposure that a conventional IT risk model doesn’t naturally capture.

  • Model weights represent months or years of compute investment and proprietary research. If captured today and decrypted once a cryptographically relevant quantum computer exists, an adversary doesn’t get a copy of raw data – they get the finished product, without doing any of the work that produced it.
  • Training datasets frequently contain sensitive personal, financial, or medical data carrying its own long-retention compliance requirements, layered on top of whatever confidentiality obligation the data already carried before it became training input.
  • Inter-agent and inference traffic is a newer, expanding surface. Multi-agent systems generate continuous streams of encrypted communication between components – a moving target harder to reason about than a single application’s data-at-rest.

What makes this qualitatively different

CSA’s own description of the risk is worth quoting directly: for organisations building and operating AI infrastructure, the breach may already have occurred, the data may already be in adversarial hands, and the organisation may not know until years from now. That’s the general character of harvest-now-decrypt-later risk, but sharpened. Where a traditional HNDL scenario spreads exposure across thousands of scattered records, AI infrastructure concentrates enormous value into a comparatively small number of artefacts – a handful of model checkpoints, a defined set of training corpora. Fewer things to protect sounds like an advantage, until the same fact is read the other way: fewer things for an adversary to bother capturing in the first place.

This isn’t a fringe concern

Thales’ 2026 Quantum & AI Threat Report found that 61% of respondents named harvest-now-decrypt-later their top quantum-related concern – a majority position, not an edge case, among organisations already engaging with the question at all.

What this changes for discovery

The practical implication is scope. Cryptographic discovery across an AI-heavy estate can’t stop at the systems a traditional inventory would naturally cover – customer-facing applications, PKI, network infrastructure. Training pipelines, model registries, vector databases, and the encryption protecting inter-agent and inter-service communication all belong in an Enterprise CBOM, not as an afterthought but as core scope from the start.

Model weights specifically deserve HNDL-style exposure scoring on their own terms. A model representing months of compute and genuine intellectual property has an effectively indefinite confidentiality requirement – the “how long must this stay secret” variable in any harvest-now-decrypt-later exposure calculation doesn’t have a natural expiry the way a time-limited session key does. That alone tends to put model weights near the front of any exposure-based prioritisation, once they’re actually in scope to be prioritised.

Further reading: the general harvest-now-decrypt-later risk, cryptographic discovery and CBOM, and cryptographic exposure assessment.

Scroll to Top