ORCID
Mehdi E. Manaa: https://orcid.org/0000-0001-6498-8562
Article Type
Original Study
Abstract
Breach rates and unparalleled vulnerabilities are a constant feature of the cyber landscape these days, and the increasing complexity of the proliferation of Internet of Things (IoT) nodes is to be expected. With these challenges, the conventional intrusion detection systems (IDS) are proven to be unable to deal with the extensive and varied data streams. Such systems can be fundamentally attributed to the classical nature of these systems, which are lacking in flexibility to analyze traffic in real-time and thus have no proactive capabilities of identifying patterns of unknown attacks. Considering these technical barriers, in this paper, an offensive-defensive system based on Deep Reinforcement Learning (DRL) algorithms is proposed. The novelty of the proposed method is the unique synergic combination of mathematical feature engineering and a dynamically changing Deep Q-Network (DQN) agent, dynamically adapting to changing traffic patterns. A comprehensive four-stage data preparation pipeline was designed and tested on the benchmark CICIoT2023 dataset, prior to the training phase. To reduce the dimensionality, non-influential features of the network were eliminated using entropy and the Synthetic Minority Over-Sampling Technique (SMOTE) was applied to address the statistical imbalance between the classes. It was tested under strict experimental conditions, with an 80:20 train/test split, on a set of 231,250 samples, before an isolated test set of 46,250 samples. This organization's initiation led to a stable mathematical context of the DQN agent, which can better formulate inferential policies to enable it to make real-time directional choices, including blocking or passing packets, by optimization of the reward function, which is ideally consistent with the concepts of zero-trust architecture. At the experimental level, the suggested framework proved to be highly efficient in its operation, with a total accuracy of 98.74%, a precision rate of 99.80%, a recall rate of 98.93%, and an F1-score of 99.37%. These numerical metrics outperform several state-of-the-art machine learning approaches in the literature, revealing that the systematic incorporation of accurate data engineering and reinforcement learning frameworks generates a field-tested security barrier that offers an expedient reaction to counteract multifaceted threats to IoT networks.
Keywords
Internet of Things (IoT), Intrusion Detection Systems (IDS), Deep Reinforcement Learning (DRL), Deep Q-Network (DQN), Anomaly detection
How to Cite This Article
Habeeb, Hawraa A. and Manaa, Mehdi E.
(2026)
"Proactive Deep Q-Learning Approach for Anomaly Detection in IoT IDSs,"
Journal of Intelligent Informatics, Networking, and Cybersecurity: Vol. 2
:
Iss.
2
, Article 8.
Available at:
https://doi.org/10.65445/3106-1192.1019
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.