Nov 2026
9 Mon
10 Tue
11 Wed
12 Thu
13 Fri 09:00 AM – 06:00 PM IST
14 Sat 09:00 AM – 06:00 PM IST
15 Sun
simran
Submitted Oct 6, 2026
A high-accuracy ML classifier can still become a security weakness when an attacker deliberately manipulates its input. This session explores how adversarial evasion attacks work and gives participants a hands-on environment to attack a working classifier, measure its robustness, and think about ML models as part of the application attack surface.
Engineering teams usually evaluate classifiers using metrics such as accuracy, precision, recall, and F1 score. These measurements tell us how a model performs on expected data, but they do not necessarily tell us how it behaves when someone is deliberately trying to make it fail.
That distinction becomes important when classifiers influence security-sensitive decisions such as malware detection, phishing or spam filtering, fraud detection, anomaly detection, or other automated allow/block decisions.
The engineering question behind this session is:
If an attacker controls or can modify the input to a classifier, how do we test what they can make the model believe?
The session explores the difference between normal model accuracy and adversarial robustness. Rather than treating adversarial ML as only a research topic, I want to approach it from the perspective of a security engineer: define the attacker, attack the decision boundary, measure what happens, and determine what engineering controls are needed around the model.
Other — Security Engineers, Developers, ML/AI Engineers, MLOps Engineers, Application Security Engineers, Penetration Testers and Red Teamers
The session is particularly useful for developers and security practitioners who are starting to encounter ML-backed functionality but have not previously performed adversarial ML testing.
Intermediate
No deep background in machine-learning mathematics is required. I will explain the necessary concepts visually and through the attack itself, but basic familiarity with software engineering, security, or machine learning will help.
Participants will learn a repeatable approach for evaluating a classifier from an attacker’s perspective: establish normal behaviour, define attacker capabilities, generate adversarial inputs, measure attack success, and evaluate robustness.
Participants will understand why high model accuracy does not automatically imply adversarial robustness, and what additional testing should be considered when an ML model influences a security-sensitive decision.
The session combines technical explanation with a hands-on adversarial ML lab.
I will first show a small classifier operating normally: the input, expected class, model prediction, and confidence.
We will then attack it.
Participants will work with a controlled model where they can generate adversarial inputs and modify parameters such as the perturbation budget. We will examine how small changes to an input affect the model’s loss, confidence, and final prediction.
The technical walkthrough will include a first-order gradient-based evasion attack such as FGSM (Fast Gradient Sign Method) and unpack what is happening behind:
x_adv = x + ε · sign(∇x J(θ, x, y))
Rather than presenting the equation in isolation, we will connect each part to the attacker’s objective: which direction to modify the input, how much modification is permitted, and whether that modification successfully changes the model’s decision.
Participants will compare clean and adversarial inputs and measure metrics such as:
We will then connect the controlled experiment to security-sensitive classifiers and discuss how the same adversarial mindset applies when a model participates in decisions such as identifying malicious or benign inputs.
The goal is not simply to demonstrate one successful adversarial example. It is to show how an engineer can systematically test an ML decision mechanism as an attack surface.
Experiment or prototype
Research / investigation
Internal engineering project / hands-on security experimentation
My background is in cybersecurity and security testing, and I became interested in adversarial ML from the perspective of how ML changes the attack surface of modern applications.
I have been experimenting with adversarial evasion techniques against classifiers and studying how model behaviour changes under deliberately manipulated inputs. The work is currently experimental rather than a claim about a production compromise.
The session is built around what I learned while trying to move from understanding adversarial attacks theoretically to actually implementing, observing, measuring, and explaining them as a security testing problem.
One of the biggest problems I encountered was treating a successful adversarial example as proof that a system is insecure.
Making a model misclassify one modified input is visually impressive, but from an engineering perspective it raises much harder questions:
Was the modification realistic?
What capabilities did the attacker require?
How much could the input change before the attack stopped being meaningful?
Does the attack work consistently or only on carefully selected examples?
Would the modification survive the transformations that occur in a real system?
This changed the direction of the work from simply “Can I fool the model?” toward “Under a clearly defined threat model, how robust is the classifier?”
Another challenge is that simple demonstrations can easily overstate what they prove. An attack demonstrated with white-box access does not automatically imply that an attacker with only API access could reproduce the same result.
These limitations will be part of the session rather than hidden from it.
I would start with the threat model before the attack algorithm.
Before choosing FGSM or another technique, I would first define:
Only then would I select an attack technique and evaluate the model.
I would also avoid using a single adversarial example as evidence of robustness or vulnerability. Instead, I would evaluate attacks across multiple inputs and perturbation budgets and report measurable results such as attack success rate and robust accuracy.
The biggest trade-off is realism versus explainability.
A complex production security classifier would make the scenario more realistic, but it would also hide much of the attack behind feature engineering, preprocessing, and domain-specific behaviour.
For the hands-on component, I therefore use a deliberately understandable classifier where participants can observe the complete attack path:
input → prediction → gradient → perturbation → adversarial input → new prediction
This sacrifices some production realism in exchange for making the mechanics observable and reproducible.
The session then maps those mechanics onto more realistic security scenarios while explicitly discussing where the assumptions change.
Another important trade-off is white-box versus black-box testing. White-box attacks provide an excellent way to understand the mechanics of adversarial optimization because gradients are available directly. Real attackers may have much less information. I therefore use white-box testing as the starting point while discussing how attacker knowledge fundamentally changes the threat model.
The main takeaway is a different way of thinking about ML-backed applications:
The model is not only an implementation component. If its output influences a consequential decision, its behaviour can become part of the attack surface.
Practitioners can take the evaluation process from this session and apply it to their own systems:
1. Identify the ML-driven decision
2. Determine what inputs an attacker can influence
3. Define realistic attacker capabilities and constraints
4. Establish baseline model performance
5. Generate adversarial inputs
6. Measure attack success and robust accuracy
7. Test mitigations and surrounding controls
This provides security teams with a starting methodology for adding adversarial robustness testing to existing application-security and ML evaluation processes.
It also highlights a mistake to avoid: treating conventional model accuracy as sufficient evidence that a classifier will behave reliably when its inputs are actively hostile.
Experimental/prototype
The attack environment and demonstrations are experimental and intended to explore adversarial robustness in a controlled setting.
The work does not claim that every production classifier can be bypassed using the demonstrated technique. Instead, the prototype provides a reproducible environment for understanding the attack mechanics, evaluating assumptions, and developing a methodology that practitioners can adapt to their own ML systems.
#security #adversarialML #AIsecurity #machinelearning #MLsecurity #evasion #redteam #inference #mlops #demo #workshop #handson #workinprogress
{{ gettext('Login to leave a comment') }}
{{ gettext('Post a comment…') }}{{ errorMsg }}
{{ gettext('No comments posted yet') }}