Misleading Performance in Sysmon-Based Machine Learning
When 100% accuracy measures the wrong thing.
Remove identifier leakage, change the unit of analysis from events to behavior, and the perfect detector becomes an honest one.
Kian Esmaeili · Research Archive
An actively maintained record of publications, experiments, platforms, and applied work in AI security, benchmark integrity, and reproducible research.
Selected work
Flagship work on evaluation, autonomous defense, and reproducible evidence.
Misleading Performance in Sysmon-Based Machine Learning
Remove identifier leakage, change the unit of analysis from events to behavior, and the perfect detector becomes an honest one.
An AI-driven validation framework for the class of update-integrity failures exemplified by the 2024 CrowdStrike Falcon content-update outage.
Read on IEEE XploreProject index
Search by topic or filter by the kind of work. Every record points to its strongest available evidence: a publication, interactive article, live system, source repository, or research record.
6 records shown.
Near-perfect ransomware-detection accuracy can be an artifact of event-level data representation. This work tests five models on 6,258 labeled Sysmon events and shows how temporal aggregation exposes the gap.
An AI-driven validation framework for the class of update-integrity failures exemplified by the 2024 CrowdStrike Falcon content-update outage.
A publication format that releases a manuscript, dataset, and evaluation code as one interactive object, allowing readers to audit the pipeline and test its claims.
A multi-variant dynamic flag platform that salts and hashes flags per student, making cybersecurity coursework resistant to answer-sharing while keeping verification automatic.
Side-channel and fault-injection work on password-verification firmware using ChipWhisperer-Nano, including timing, power, electromagnetic, and acoustic analysis.
Power Apps automation for data analysis and reporting in cybersecurity examination workflows, alongside enterprise AI use cases for cybersecurity and operational work.
Lab notes
Short technical essays about the questions, design decisions, and lessons behind selected research and systems.
Benchmark integrity
Reproducible research
Cybersecurity education
Endpoint security
Research themes
Does an evaluation measure the threat—or an artifact of its own design?
What must validation look like before autonomous change deserves trust?
Manuscript, dataset, and evaluation code should belong together.
Research direction
The next direction combines behavioral telemetry with AI reasoning: detecting threats earlier, evaluating the detector with the same scrutiny this work argues for, and requiring the system to justify a decision when asked.