Practitioner

Data Protection for AI Systems

Proves the ability to govern personal and confidential data across training, retrieval and inference without relying on privacy theatre.

7 modules 49 min of reading 18 questions 75% to pass Free
About this course

This practitioner exam assesses working knowledge of how data protection actually behaves in AI systems: provenance and purpose limitation for training corpora, memorisation and inference attacks, what differential privacy does and does not promise, the gap between de-identification and anonymisation, and the privacy properties of embeddings and retrieval indexes. It is aimed at engineers, privacy engineers and architects who build or review data pipelines feeding models. Candidates are expected to reason about residual risk and enforceable controls rather than recite regulation.

What it covers
  • Training data governance: provenance, licensing, lineage records and purpose limitation on reuse
  • Privacy attacks: memorisation and extraction, membership inference, attribute inference and model inversion
  • Formal and informal protections: differential privacy guarantees and limits, de-identification, k-anonymity and re-identification
  • Retrieval surfaces: embedding leakage, permission-aware indexes and tenant isolation in RAG
  • Lifecycle obligations: minimisation in prompts, retention and deletion including deletion from trained weights
  • Jurisdiction and assessment: cross-border transfer, DPIA triggers and residual-risk documentation
Part 1

Learn the material

7 modules, about 49 minutes of reading. Work through them in order, or jump to whatever you need. The assessment is drawn from exactly this material.

  1. 01 Knowing what you trained on: provenance, lineage and lawful reuse Why dataset-level paperwork cannot answer the questions legal will ask, and what record-level lineage has to carry to make reuse defensible. 7 min
  2. 02 What the model remembers: memorisation, extraction and inference attacks How trained models leak individual records, why memorisation is not hallucination, and why membership inference is fundamentally a generalisation problem. 6 min
  3. 03 Differential privacy: reading the guarantee literally What an epsilon actually bounds, what composition does to a privacy budget, and the specific claims differential privacy does not support. 8 min
  4. 04 De-identification, anonymisation and the re-identification gap Why stripping identifiers does not produce anonymous data, what k-anonymity guarantees and what it misses, and how to reason about re-identification risk. 7 min
  5. 05 Minimisation at inference, retention, and the deletion problem How to cut exposure in prompts without losing answer quality, how to design retention that survives review, and what erasure honestly means once data is in the weights. 7 min
  6. 06 Embeddings, vector stores and retrieval isolation Why an index of vectors derived from personal data is itself personal data, and how to make retrieval enforce permissions and tenant boundaries instead of hoping. 7 min
  7. 07 Jurisdiction, impact assessment and residual risk When an inference call is a cross-border transfer, which AI deployments trigger a formal impact assessment, and how to write down what your controls do not cover. 7 min

Start the course

Part 2

Take the assessment

18 questions drawn from the material above. Pass and you can put your name to a certificate with a serial anyone can verify.

How it is marked

  • Questions and answer options are shuffled for every sitting.
  • Multi-answer questions are marked as a set: you need all of the correct options and none of the wrong ones. There is no partial credit.
  • You need 75% to pass.
  • You can revisit and change any answer until you submit.
  • Afterwards you see every question, the answer you gave, and whether it was right. The answer key is never printed.
  • The reasoning behind each answer is released once you pass. Held back on a fail, it would hand over most of the key to anyone willing to sit the paper once and read it, which is why the taught material above is the intended route back.
  • You can re-sit the paper, but not immediately: there is a ten minute wait between attempts on the same course.