Courses Practitioner
Data Protection for AI Systems
Proves the ability to govern personal and confidential data across training, retrieval and inference without relying on privacy theatre.
This practitioner exam assesses working knowledge of how data protection actually behaves in AI systems: provenance and purpose limitation for training corpora, memorisation and inference attacks, what differential privacy does and does not promise, the gap between de-identification and anonymisation, and the privacy properties of embeddings and retrieval indexes. It is aimed at engineers, privacy engineers and architects who build or review data pipelines feeding models. Candidates are expected to reason about residual risk and enforceable controls rather than recite regulation.
What it covers
- Training data governance: provenance, licensing, lineage records and purpose limitation on reuse
- Privacy attacks: memorisation and extraction, membership inference, attribute inference and model inversion
- Formal and informal protections: differential privacy guarantees and limits, de-identification, k-anonymity and re-identification
- Retrieval surfaces: embedding leakage, permission-aware indexes and tenant isolation in RAG
- Lifecycle obligations: minimisation in prompts, retention and deletion including deletion from trained weights
- Jurisdiction and assessment: cross-border transfer, DPIA triggers and residual-risk documentation
How it is marked
- Questions and answer options are shuffled for every sitting.
- Multi-answer questions are marked as a set: you need all of the correct options and none of the wrong ones. There is no partial credit.
- You need 75% to pass.
- You can revisit and change any answer until you submit.
- Afterwards you see every question, the answer you gave, whether it was right, and the reasoning behind it. The answer key itself is never printed, so the paper stays worth sitting.
- You can re-sit the paper, but not immediately: there is a short wait between attempts.