Adversarial ML and Red Teaming
The hardest exam in the ladder: proves expert judgement on adversarial attacks, defence evaluation and red team practice.
The capstone examination for practitioners who attack AI systems or defend against those who do. It assesses adversarial example theory, poisoning and extraction attacks, the evaluation discipline that separates real robustness from gradient masking, jailbreak methodology, and the scoping, measurement and reporting standards of professional AI red teaming. Expect questions that turn on subtle distinctions; partial knowledge will not pass.
- Attack theory: adversarial examples, transferability, and white, grey and black box threat models
- Training-time attacks: data poisoning, backdoor triggers, and clean-label techniques
- Model confidentiality: extraction, membership inference, and the trade-offs of every defence
- Robustness evaluation: gradient masking, adaptive attacks, and certified versus empirical guarantees
- LLM offence: jailbreak taxonomy, automated attack generation, and multi-turn escalation
- Professional practice: scoping, rules of engagement, attack success measurement, and reproducible reporting
Learn the material
8 modules, about 54 minutes of reading. Work through them in order, or jump to whatever you need. The assessment is drawn from exactly this material.
- 01 Adversarial Examples, Transferability and Threat Models What an adversarial example actually is, why perturbations transfer between independently trained models, and how to classify an adversary honestly. 7 min
- 02 The Attack Taxonomy and Training-Time Attacks Evasion, poisoning and extraction as distinct classes, and how backdoors and clean-label poisoning defeat the obvious controls. 7 min
- 03 Model Extraction and the Honest Cost of Every Defence Why query access leaks the model, and what rate limiting, watermarking and output perturbation actually buy you. 7 min
- 04 Membership Inference, Memorisation and Their Mitigations Why a small generalisation gap does not imply privacy, how per-example leakage must be measured, and what differential privacy does and does not promise. 6 min
- 05 Gradient Masking and the Adaptive Attack Principle How defences manufacture false confidence by obstructing the attacker optimiser, and why a defence is only tested by an adversary who knows it is there. 7 min
- 06 Certified Robustness, Empirical Robustness and the Trade-off What a certificate actually bounds, why certified numbers sit below measured ones, and how to choose a perturbation budget as a risk decision. 6 min
- 07 Jailbreaks, Automated Red Teaming and Multi-Turn Attacks Classifying policy attacks correctly, the technique families that compose into working jailbreaks, and why per-turn evaluation misses escalation. 7 min
- 08 Professional Practice: Scoping, Measurement and Reporting Rules of engagement for testing a live system, how to make an attack success rate interpretable, and the artefact set that makes a finding actionable. 7 min
Take the assessment
20 questions drawn from the material above. Pass and you can put your name to a certificate with a serial anyone can verify.
How it is marked
- Questions and answer options are shuffled for every sitting.
- Multi-answer questions are marked as a set: you need all of the correct options and none of the wrong ones. There is no partial credit.
- You need 80% to pass.
- You can revisit and change any answer until you submit.
- Afterwards you see every question, the answer you gave, and whether it was right. The answer key is never printed.
- The reasoning behind each answer is released once you pass. Held back on a fail, it would hand over most of the key to anyone willing to sit the paper once and read it, which is why the taught material above is the intended route back.
- You can re-sit the paper, but not immediately: there is a ten minute wait between attempts on the same course.