Cognitive Detection at Big-Data Scale: A CNN-LSTM-DQN Framework with Prioritized Experience Replay for Cross-Attack-Family Generalization and Multi-Seed Initialization Sensitivity Analysis
Abstract
1. Introduction
- Cross-attack-family generalization experiment. We re-instantiate the CNN–LSTM–DQN+PER methodology [19]—architecture template, reward design, and λ calibration inherited from the CSE-CIC-IDS2018 (E7) study, backbone retrained on TON_IoT—on the TON_IoT Processed_Network dataset (5,000,000 flows, temporal split, 94.53% attack ratio, exclusively XSS test condition [5,6]), achieving 98.02% attack recall and 98.73% F1-score on the best-performing seed (42; five-seed mean recall 0.833 ± 0.306)—compared against the previously published CICIDS2018 benchmark recall of 91.40% [4].
- Five-seed initialization sensitivity protocol. We evaluate both X1 (CNN–LSTM supervised baseline) and X2 (CNN–LSTM–DQN+PER proposed) under five independent random seeds {7, 13, 21, 42, 99} with paired Wilcoxon signed-rank and t-tests, quantifying the initialization sensitivity of cognitive DQN agents [3,18] under extreme class imbalance and deriving an empirical minimum-seed recommendation for statistical adequacy.
- Window-based temporal stability analysis. We partition the 1,000,000-flow test set into four sequential windows of 250,000 flows and evaluate per-window consistency, demonstrating a recall variance of 4.63 × 10−7 and F1 variance of 1.38 × 10−7—confirming that the CNN–LSTM encoder’s latent representations [23,24] generalize temporally across an unseen attack family.
- Scale-aware ARMF decomposition. We formally separate ARMF into a structural component (driven by attack prevalence, 99.46% of observed ARMF on TON_IoT) and a model-induced component (driven by false positives, 0.54%), establishing that cross-dataset ARMF comparisons require prevalence normalization and providing a correction formula for environment-invariant reward calibration [19].
- Degenerate initialization analysis. We identify, reproduce, and mechanistically explain a previously undocumented failure mode of value-based cognitive DQN agents: under a pathological 68:1 attack-to-normal test ratio, a single initialization (seed 21) fails to escape a poor basin of attraction and collapses to chronic under-alerting. This is a degenerate initialization-and-optimization failure rather than reward exploitation—the asymmetric reward (αFN = −5.0 vs. αFP = −0.5) penalizes under-alerting an order of magnitude more heavily than over-alerting, so the collapse runs counter to the reward gradient [3,18]. Because such under-alerting would be further amplified if an explicit alert-volume penalty (Equation (7), λ > 0) were activated and transferred across environments of differing attack prevalence, we derive a prospective prevalence-adaptive λ normalization (Equation (13))—a theoretically motivated correction that would remove the prevalence dependence of an explicit alert-volume penalty under cross-environment transfer; its empirical validation is left to future work.
- Systematic comparative positioning. We compare the proposed framework against 16 representative recent works in a structured literature table (Table 1); unlike the surveyed literature, the present evaluation combines all four criteria: naturalistic big-data TON_IoT distribution [5,6], cognitive DQN+PER adaptive detection [18,20], ARMF operational metric reporting [19], and multi-seed Wilcoxon statistical validation [15].
2. Related Work
2.1. Deep Learning-Based Hybrid Architectures for NIDS
2.2. Reinforcement Learning and Cognitive Agents for NIDS
2.3. Cross-Dataset Generalization and Big-Data Evaluation
2.4. Comparative Literature Summary
3. Dataset and Big-Data Processing Pipeline
3.1. TON_IoT Processed_Network Dataset
3.2. Systematic Sampling and Scale Characterization
3.3. Attack Family Composition and Unseen-Attack-Family Condition
3.4. Feature Engineering for Big-Data Scale
3.5. Experimental Design and Configuration
4. CNN-LSTM-DQN+PER Cognitive Detection Framework
4.1. Architecture Overview
4.2. CNN-LSTM Backbone
4.2.1. Convolutional Feature Extractor
4.2.2. LSTM Temporal Encoder
4.2.3. Supervised Training—Phase 1 (X1)
4.3. Cognitive DQN Policy with ARMF-Aware Reward
4.3.1. Markov Decision Process Formulation
4.3.2. Reward Function Design
4.3.3. Double Q-Learning Update
4.4. Prioritised Experience Replay at Big-Data Scale
4.5. Training Protocol and Computational Requirements
5. Experimental Results
5.1. Multi-Seed Benchmark Results and Latent Space Analysis
5.2. Cross-Environment Methodology-Transfer Analysis
5.3. Window-Based Temporal Stability Analysis
5.4. DQN Cognitive Training Dynamics
5.5. Statistical Significance Testing and Power Assessment
6. Discussion
6.1. ARMF Decomposition: Structural vs. Model-Induced Alert Volume
6.2. Degenerate Seed Analysis: Initialisation-Induced Optimisation Failure Under High Prevalence
6.3. Implications for Big-Data NIDS Benchmarking
6.4. Limitations
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Chen, M.; Herrera, F.; Hwang, K. Cognitive Computing: Architecture, Technologies and Intelligent Applications. IEEE Access 2018, 6, 19774–19783. [Google Scholar] [CrossRef]
- Khraisat, A.; Gondal, I.; Vamplew, P.; Kamruzzaman, J. Survey of intrusion detection systems: Techniques, datasets and challenges. Cybersecurity 2019, 2, 20. [Google Scholar] [CrossRef]
- Kheddar, H.; Dawoud, D.W.; Awad, A.I.; Himeur, Y.; Khan, M.K. Reinforcement-Learning-Based Intrusion Detection in Communication Networks: A Review. IEEE Commun. Surv. Tutor. 2025, 27, 2420–2469. [Google Scholar] [CrossRef]
- Sharafaldin, I.; Lashkari, A.H.; Ghorbani, A.A. Toward generating a new intrusion detection dataset and intrusion traffic characterization. ICISSp 2018, 1, 108–116. [Google Scholar] [CrossRef]
- Moustafa, N. A new distributed architecture for evaluating AI-based security systems at the edge: Network TON_IoT datasets. Sustain. Cities Soc. 2021, 72, 102994. [Google Scholar] [CrossRef]
- Alsaedi, A.; Moustafa, N.; Tari, Z.; Mahmood, A.; Anwar, A. TON_IoT Telemetry Dataset: A New Generation Dataset of IoT and IIoT for Data-Driven Intrusion Detection Systems. IEEE Access 2020, 8, 165130–165150. [Google Scholar] [CrossRef]
- Ahmad, Z.; Khan, A.S.; Shiang, C.W.; Abdullah, J.; Ahmad, F. Network intrusion detection system: A systematic study of machine learning and deep learning approaches. Trans. Emerg. Telecommun. Technol. 2021, 32, e4150. [Google Scholar] [CrossRef]
- Ferrag, M.A.; Maglaras, L.; Moschoyiannis, S.; Janicke, H. Deep learning for cyber security intrusion detection: Approaches, datasets, and comparative study. J. Inf. Secur. Appl. 2020, 50, 102419. [Google Scholar] [CrossRef]
- Halbouni, A.; Gunawan, T.S.; Habaebi, M.H.; Halbouni, M.; Kartiwi, M.; Ahmad, R. CNN-LSTM: Hybrid Deep Neural Network for Network Intrusion Detection System. IEEE Access 2022, 10, 99837–99849. [Google Scholar] [CrossRef]
- Abdallah, M.; Le-Khac, N.; Jahromi, H.Z.; Jurcut, A.D. A Hybrid CNN-LSTM Based Approach for Anomaly Detection Systems in SDNs. In Proceedings of the 16th International Conference on Availability, Reliability and Security, Vienna, Austria, 17–20 August 2021; pp. 1–7. [Google Scholar] [CrossRef]
- Altunay, H.C.; Albayrak, Z. A hybrid CNN+LSTM-based intrusion detection system for industrial IoT networks. Eng. Sci. Technol. Int. J. 2023, 38, 101322. [Google Scholar] [CrossRef]
- Bamber, S.S.; Katkuri, A.V.R.; Sharma, S.; Angurala, M. A hybrid CNN-LSTM approach for intelligent cyber intrusion detection system. Comput. Secur. 2025, 148, 104146. [Google Scholar] [CrossRef]
- Karthik, M.G.; Keerthika, V.; Mantena, S.V.; Siri, D.; Yeluri, L.P.; Lella, K.K.; Ganesh, B.R. Energy-efficient intrusion detection with a protocol-aware transformer–spiking hybrid model. Sci. Rep. 2026, 16, 7095. [Google Scholar] [CrossRef] [PubMed]
- Alkasassbeh, M.; Omoush, E.H.; Almseidin, M.; Aldweesh, A. A Self-Adaptive Intrusion Detection System for Zero-Day Attacks Using Deep Q-Networks. IEEE Access 2025, 13, 174280–174296. [Google Scholar] [CrossRef]
- Sarhan, M.; Layeghy, S.; Portmann, M. Towards a standard feature set for network intrusion detection system datasets. Mob. Netw. Appl. 2022, 27, 357–370. [Google Scholar] [CrossRef]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction; MIT Press: Cambridge, UK, 1998; Volume 1. [Google Scholar]
- Nguyen, T.T.; Reddi, V.J. Deep reinforcement learning for cyber security. IEEE Trans. Neural Netw. Learn. Syst. 2021, 34, 3779–3795. [Google Scholar] [CrossRef] [PubMed]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.A.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [PubMed]
- Rushendra; Ramli, K.; Purnamasari, P.D. Stability-Aware Evaluation of a CNN–LSTM–DQN Intrusion Detection System for Zero-Day and Drifted Network Traffic. IIUM Eng. J. 2026, 27, 227–256. [Google Scholar] [CrossRef]
- Schaul, T.; Quan, J.; Antonoglou, I.; Silver, D.; Deepmind, G. Prioritized Experience Replay. In Proceedings of the 4th International Conference on Learning Representations (ICLR), San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar] [CrossRef]
- Ren, K.; Zeng, Y.; Cao, Z.; Zhang, Y. ID-RDRL: A deep reinforcement learning-based feature selection intrusion detection model. Sci. Rep. 2022, 12, 15370. [Google Scholar] [CrossRef] [PubMed]
- Tan, H.; Wang, L.; Zhu, D.; Deng, J. Intrusion detection based on adaptive sample distribution dual-experience replay reinforcement learning. Mathematics 2024, 12, 948. [Google Scholar] [CrossRef]
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [PubMed]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [PubMed]
- Sinha, P.; Sahu, D.; Prakash, S.; Yang, T.; Rathore, R.S.; Pandey, V.K. A high performance hybrid LSTM CNN secure architecture for IoT environments using deep learning. Sci. Rep. 2025, 15, 9684. [Google Scholar] [CrossRef] [PubMed]
- Alavizadeh, H.; Alavizadeh, H.; Jang-Jaccard, J. Deep Q-Learning Based Reinforcement Learning Approach for Network Intrusion Detection. Computers 2022, 11, 41. [Google Scholar] [CrossRef]
- Alam, K.; Monir, M.F.; Hossain, M.J.; Uddin, M.S.; Habib, M.T. Adaptive Defense: Zero-Day Attack Detection in NIDS with Deep Reinforcement Learning. IEEE Access 2025, 13, 116345–116361. [Google Scholar] [CrossRef]
- Wu, Y.; Hu, Y.; Wang, J.; Feng, M.; Dong, A.; Yang, Y. An active learning framework using deep Q-network for zero-day attack detection. Comput. Secur. 2024, 139, 103713. [Google Scholar] [CrossRef]
- Hossain, M.A. Deep Q-learning intrusion detection system (DQ-IDS): A novel reinforcement learning approach for adaptive and self-learning cybersecurity. ICT Express 2025, 11, 875–880. [Google Scholar] [CrossRef]
- Shaikh, J.A.; Wang, C.; Sima, M.W.U.; Arshad, M.; Owais, M.; Hassan, D.S.M.; Alkanhel, R.; Muthanna, M.S.A. A deep Reinforcement learning-based robust Intrusion Detection System for securing IoMT Healthcare Networks. Front. Med. 2025, 12, 1524286. [Google Scholar] [CrossRef] [PubMed]
- Lin, Y.D.; Huang, H.X.; Sudyana, D.; Lai, Y.C. AI for AI-based intrusion detection as a service: Reinforcement learning to configure models, tasks, and capacities. J. Netw. Comput. Appl. 2024, 229, 103936. [Google Scholar] [CrossRef]
- Susilo, B.; Muis, A.; Sari, R.F. Intelligent Intrusion Detection System Against Various Attacks Based on a Hybrid Deep Learning Algorithm. Sensors 2025, 25, 580. [Google Scholar] [CrossRef] [PubMed]
- Sajid, M.; Malik, K.R.; Almogren, A.; Malik, T.S.; Khan, A.H.; Tanveer, J.; Rehman, A.U. Enhancing intrusion detection: A hybrid machine and deep learning approach. J. Cloud Comput. 2024, 13, 123. [Google Scholar] [CrossRef]
- Zhang, J.Z.; Srivastava, P.R.; Sharma, D.; Eachempati, P. Big data analytics and machine learning: A retrospective overview and bibliometric analysis. Expert Syst. Appl. 2021, 184, 115561. [Google Scholar] [CrossRef]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep Reinforcement Learning with Double Q-Learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016. [Google Scholar] [CrossRef]








| Reference | Year | Dataset (Scale) | Approach | Cognitive/Big-Data Angle | Performance | Key Gap |
|---|---|---|---|---|---|---|
| Halbouni et al. [9] | 2022 | CICIDS2017 | CNN-LSTM | Hybrid spatio-temporal | Acc 99.3% | Single dataset; no RL; no ARMF |
| Abdallah et al. [10] | 2021 | NSL-KDD/SDN | CNN-LSTM | Anomaly in SDN | F1 98.9% | No cross-dataset; no statistical test |
| Altunay & Albayrak [11] | 2023 | CIC-IDS/UNSW | CNN+LSTM | Industrial IoT hybrid | F1 98.2% | Single dataset; no RL; no ARMF |
| Bamber et al. [12] | 2025 | Multiple | CNN-LSTM | Intelligent cyber IDS | Acc 99.1% | No RL adaptive policy; no multi-seed |
| Sinha et al. [25] | 2025 | IoT datasets | LSTM-CNN | Secure IoT architecture | F1 98.7% | No cognitive RL; no cross-dataset |
| Alavizadeh et al. [26] | 2022 | KDDCup99/NSL | DQN | Cognitive RL agent | Recall 92.3% | Outdated dataset; no PER; no ARMF |
| Alam et al. [27] | 2025 | CICIDS2017 | DRL-based NIDS | Zero-day via DRL | Acc 98.5% | Single dataset; no cross-dataset; no ARMF |
| Wu et al. [28] | 2024 | CICIDS2018 | DQN active learning | Zero-day detection | Recall 94.7% | No PER; no multi-seed; no big-data scale |
| Hossain [29] | 2025 | NSL-KDD | DQ-IDS | Adaptive self-learning | Acc 99.2% | Single dataset; no ARMF; no cross-dataset |
| Alkasassbeh et al. [14] | 2025 | TON_IoT (balanced) | DQN self-adaptive | Zero-day via DQN | Acc 98.9% | Pre-balanced partition; no multi-seed; no ARMF |
| Shaikh et al. [30] | 2025 | IoMT dataset | DRL | Healthcare IoMT IDS | Recall 96.3% | Domain-specific; no network traffic; no ARMF |
| Ren et al. [21] | 2022 | CICIDS2017 | DRL feature select | RL feature importance | F1 97.8% | No PER; no cross-dataset; no ARMF |
| Tan et al. [22] | 2024 | NSL-KDD | Dual-ER RL | Adaptive replay | Acc 98.6% | No CNN-LSTM; no cross-dataset; no ARMF |
| Lin et al. [31] | 2024 | CICIDS datasets | RL config | AI-for-IDS-as-service | Acc 97.4% | No PER; no multi-seed; no TON_IoT |
| Susilo et al. [32] | 2025 | Multiple IoT | Hybrid DL | Intelligent detection | Acc 99.0% | No RL policy; no cross-dataset validation |
| Sajid et al. [33] | 2024 | Multiple | Hybrid ML+DL | Ensemble approach | F1 98.5% | No RL adaptive; no ARMF; no multi-seed |
| This Work | 2026 | TON_IoT Proc. Network (5 M flows, temporal split) | CNN-LSTM-DQN+PER (5 seeds) | Adaptive cognitive agent; ARMF-aware reward; big-data naturalistic | Recall 98.02%, F1 98.73% (seed 42) | Combination not reported in the 16 surveyed works: PER-RL + ARMF-aware evaluation + 5-seed Wilcoxon + prior-benchmark comparison + unseen-attack-family (XSS) test on 5 M naturalistic TON_IoT |
| Partition | Total Flows | Normal (0) | Attack (1) | Attack Ratio | Attack Families |
|---|---|---|---|---|---|
| Full Sample | 5,000,000 | 273,618 (5.47%) | 4,726,382 (94.53%) | 17.27:1 | 5 families |
| Training | 4,000,000 | 259,211 (6.48%) | 3,740,789 (93.52%) | 14.43:1 | 4 (scan, ddos, dos, inj) |
| Test | 1,000,000 | 14,407 (1.44%) | 985,593 (98.56%) | 68.41:1 | XSS only (unseen family) |
| Attack Family | Train Count | Train % | Test Count | Test % | Status |
|---|---|---|---|---|---|
| Scanning | 1,777,848 | 47.53% | 0 | 0% | Training only |
| DDoS | 998,109 | 26.68% | 0 | 0% | Training only |
| DoS | 839,637 | 22.44% | 0 | 0% | Training only |
| Injection | 125,195 | 3.35% | 0 | 0% | Training only |
| XSS | 0 | 0% | 985,593 | 98.56% | Unseen Family |
| Normal | 259,211 | 6.48% | 14,407 | 1.44% | Both partitions |
| Category | Count | Features (Index: Name) | NIDS Relevance |
|---|---|---|---|
| Traffic Volume | 8 | 00:duration, 01:src_bytes, 02:dst_bytes, 03:missed_bytes, 04:src_pkts, 05:src_ip_bytes, 06:dst_pkts, 07:dst_ip_bytes | Byte/packet volume discriminates DoS, scanning from benign; missed_bytes flags dropped segments |
| Protocol-Level | 6 | 08:dns_qclass, 09:dns_qtype, 10:dns_rcode, 11:http_request_body_len, 12:http_response_body_len, 13:http_status_code | Application-layer behaviors distinguish XSS/SQLi (anomalous HTTP bodies) from benign browsing |
| Flag Encoding | 8 | 14:proto_enc, 15:dns_AA_enc, 16:dns_RD_enc, 17:dns_RA_enc, 18:dns_rejected_enc, 19:ssl_resumed_enc, 20:ssl_established_enc, 21:weird_notice_enc, 22:http_trans_depth_enc | Protocol state abnormalities indicate malformed packets, injection, or scanning attempts |
| Connection State (One-Hot) | 13 | 23:cs_S0, 24:cs_S1, 25:cs_S2, 26:cs_S3, 27:cs_SF, 28:cs_SH, 29:cs_SHR, 30:cs_OTH, 31:cs_REJ, 32:cs_RSTO, 33:cs_RSTOS0, 34:cs_RSTR, 35:cs_RSTRH | Connection lifecycle patterns: S0 flood→DoS; REJ→port scan; SF→normal session |
| Service Type (One-Hot) | 9 | 36:svc_-, 37:svc_dns, 38:svc_http, 39:svc_ftp, 40:svc_ssl, 41:svc_gssapi, 42:svc_dce_rpc, 43:svc_smb, 44:svc_smb;gssapi | Service distribution differs across attack families; XSS concentrates in svc_http/svc_ssl |
| Total | 45 | — | Covers volumetric, behavioral, structural, and protocol-state attack indicators |
| Parameter | X1 (CNN-LSTM Baseline) | X2 (CNN-LSTM-DQN+PER Proposed) | Rationale |
|---|---|---|---|
| CNN filter counts | [32, 64, 128] | [32, 64, 128]—frozen in Phase 2 | Progressive feature abstraction |
| LSTM hidden units | 128 | 128—frozen in Phase 2 | Latent dim = DQN state space |
| Optimizer | Adam | Adam (DQN head only) | Adaptive learning rate |
| Learning rate | 1 × 10−3 | 1 × 10−4 (Phase 2) | Fine-tuning regime for DQN |
| Batch size | 256 | 256 | GPU memory-bounded |
| Epochs | 10 (Phase 1) | 10 (Phase 2) | Equal training budget |
| Class weights | Inverse frequency | Inverse frequency (Phase 1) | 14.43:1 imbalance correction |
| Loss function | Class-weighted cross-entropy | IS-weighted TD MSE (Phase 2) | Phase-appropriate objective |
| Discount factor γ | N/A | 0.8 | Moderate future reward horizon |
| ε (exploration) | N/A | 0.20 → 0.01 (linear decay) | Decaying greedy exploration |
| αper (PER priority) | N/A | 0.6 | Moderate prioritization sharpness |
| β (IS correction) | N/A | 0.40 → 1.00 (annealed) | Gradually removes sampling bias |
| λ (ARMF penalty) | N/A | Inherited from E7 calibration | Cross-env transfer test condition |
| Random seeds | {7, 13, 21, 42, 99} | {7, 13, 21, 42, 99} | Multi-seed robustness protocol |
| Epoch | Epsilon (ε) | Beta (β) | Updates | TD Loss | Avg Reward | Δ Loss | Notes |
|---|---|---|---|---|---|---|---|
| 1 | 0.181 | 0.46 | 960 | 0.3799 | 1.1089 | — | High IS bias; initial exploration |
| 2 | 0.162 | 0.52 | 978 | 0.2430 | 1.1410 | −0.1369 | Rapid early policy improvement |
| 3 | 0.143 | 0.58 | 978 | 0.3916 | 1.1886 | +0.1486 | Exploration transient (ε-greedy active) |
| 4 | 0.124 | 0.64 | 978 | 0.3450 | 1.2277 | −0.0466 | Post-transient stabilisation |
| 5 | 0.105 | 0.70 | 978 | 0.2963 | 1.2642 | −0.0487 | Continued convergence |
| 6 | 0.086 | 0.76 | 978 | 0.2469 | 1.3044 | −0.0494 | IS correction strengthening |
| 7 | 0.067 | 0.82 | 978 | 0.2111 | 1.3459 | −0.0358 | Near-exploitation regime |
| 8 | 0.048 | 0.88 | 978 | 0.1754 | 1.3867 | −0.0357 | Low exploration; stable policy |
| 9 | 0.029 | 0.94 | 978 | 0.1382 | 1.4260 | −0.0372 | Near full IS correction |
| 10 | 0.010 | 1.00 | 978 | 0.1291 | 1.4693 | −0.0091 | Full IS correction; final policy |
| Model/Configuration | Dataset | Accuracy | Recall | Precision | F1-Score | ARMF | Seeds |
|---|---|---|---|---|---|---|---|
| X1: CNN-LSTM Baseline (seed 7) | TON_IoT Proc. | 55.78% | 55.57% | 99.22% | 71.24% | 552,055 | 1 |
| X1: CNN-LSTM Baseline (seed 13) | TON_IoT Proc. | 97.58% | 98.12% | 99.42% | 98.76% | 972,708 | 1 |
| X1: CNN-LSTM Baseline (seed 21) | TON_IoT Proc. | 83.99% | 84.24% | 99.43% | 91.21% | 835,080 | 1 |
| X1: CNN-LSTM Baseline (seed 42) | TON_IoT Proc. | 97.08% | 97.58% | 99.45% | 98.51% | 967,002 | 1 |
| X1: CNN-LSTM Baseline (seed 99) | TON_IoT Proc. | 93.07% | 93.48% | 99.45% | 96.38% | 926,314 | 1 |
| X1: Mean ± Std (5 seeds) | TON_IoT Proc. | 85.50 ± 17.5% | 85.80 ± 17.8% | 99.39 ± 0.10% | 91.22 ± 11.6% | 850,632 ± 175,757 | 5 |
| X2: CNN-LSTM-DQN+PER (seed 7) | TON_IoT Proc. | 96.77% | 97.26% | 99.45% | 98.34% | 963,816 | 1 |
| X2: CNN-LSTM-DQN+PER (seed 13) | TON_IoT Proc. | 96.25% | 96.77% | 99.42% | 98.07% | 959,329 | 1 |
| X2: CNN-LSTM-DQN+PER (seed 21) | TON_IoT Proc. | 29.34% | 28.59% | 99.02% | 44.37% | 284,528 | 1 |
| X2: CNN-LSTM-DQN+PER (seed 42) | TON_IoT Proc. | 97.52% | 98.02% | 99.46% | 98.73% | 971,319 | 1 |
| X2: CNN-LSTM-DQN+PER (seed 99) | TON_IoT Proc. | 95.24% | 95.70% | 99.45% | 97.54% | 948,402 | 1 |
| X2: Mean ± Std (5 seeds) | TON_IoT Proc. | 83.02 ± 30.0% | 83.26 ± 30.6% | 99.36 ± 0.19% | 87.41 ± 24.1% | 825,479 ± 302,515 | 5 |
| E7: DQN+PER—CICIDS2018 [19] | CICIDS2018 | 99.90% | 91.40% | 6.11% | 11.46% | 1031 | 1 |
| Seed | X1 Recall | X2 Recall | Δ Recall (pp) | X1 ARMF | X2 ARMF | Δ ARMF | Status |
|---|---|---|---|---|---|---|---|
| 7 | 55.57% | 97.26% | +41.69 | 552,055 | 963,816 | +411,761 | STABLE (X2 ≫ X1) |
| 13 | 98.12% | 96.77% | −1.35 | 972,708 | 959,329 | −13,379 | STABLE (X2 ≈ X1) |
| 21 | 84.24% | 28.59% | −55.65 | 835,080 | 284,528 | −550,552 | DEGENERATE |
| 42 | 97.58% | 98.02% | +0.44 | 967,002 | 971,319 | +4317 | STABLE (X2 > X1) ★ |
| 99 | 93.48% | 95.70% | +2.22 | 926,314 | 948,402 | +22,088 | STABLE (X2 > X1) |
| Mean | 85.80% | 83.26% | −2.54 | 850,632 | 825,479 | −25,153 | — |
| Std | 17.79% | 27.35% | — | 175,757 | 302,515 | — | — |
| Window | Samples | Attacks | Alerts | ARMF | Accuracy | Recall | F1-Score |
|---|---|---|---|---|---|---|---|
| W1 | 250,000 | 246,333 | 242,547 | 970,188 | 97.41% | 97.92% | 98.68% |
| W2 | 250,000 | 246,370 | 242,784 | 971,136 | 97.54% | 98.02% | 98.74% |
| W3 | 250,000 | 246,339 | 242,916 | 971,664 | 97.55% | 98.06% | 98.75% |
| W4 | 250,000 | 246,551 | 243,072 | 972,288 | 97.56% | 98.06% | 98.75% |
| Mean | -- | -- | -- | 971,319 | 97.52% | 98.02% | 98.73% |
| Variance | -- | -- | -- | 790,212 | -- | 4.63 × 10−7 | 1.38 × 10−7 |
| (a) | |||||||
| Metric | X1 Mean | X1 Std | X2 Mean | X2 Std | Wilcoxon p | t-Test p | Sig. (p < 0.05) |
| ARMF | 850,632 | 175,757 | 825,479 | 302,515 | 0.594 | 0.439 | No |
| Recall | 0.858 | 0.178 | 0.833 | 0.306 | 1.000 | 0.878 | No |
| F1 | 0.912 | 0.116 | 0.874 | 0.241 | 0.500 | 1.000 | No |
| (b) | |||||||
| Metric | X1 Mean | X2 Mean | Δ Mean | Wilcoxon p | t-Test p | Sig. (p < 0.05) | Interpretation |
| Recall | 0.862 | 0.969 | +0.107 | 0.375 | 0.375 | No | Three of four stable seeds improve; seed 13 has a small reversal; +10.7 pp mean |
| ARMF | 854,520 | 960,717 | +106,197 | 0.375 | 0.375 | No | Higher X2 ARMF reflects more TP detected (higher recall) |
| F1 | 0.912 | 0.982 | +0.069 | 0.375 | 0.376 | No | Three of four stable seeds improve; seed 13 has a small reversal; +6.9 pp mean |
| Component | CICIDS2018 E7 [19] | TON_IoT X1 (s42) | TON_IoT X2 (s42) | Interpretation |
|---|---|---|---|---|
| ) | 2,697,128 | 1,000,000 | 1,000,000 | Per-million normalisation denominator |
| ) | 186 | 985,593 | 985,593 | Raw attack count; not a summand |
| ) | 69 | 985,593 | 985,593 | Reference only; not a summand |
| Attack prevalence | 0.007% | 98.56% | 98.56% | 14,080× higher in TON_IoT |
| False positives (FP) | 2612 | 5284 | 5284 | Identical—DQN adds no FP |
(TP-driven) | 63 | 961,718 | 966,035 | TP-driven alert mass at achieved recall |
| (FP-driven) | 968 | 5284 | 5284 | Model-induced alert burden |
| Total ARMF (observed) | 1031 | 967,002 | 971,319 | Structural + model-induced |
| Model-induced share | 93.89% | 0.55% | 0.54% | Reversal of dominant component |
| Structural share | 6.11% | 99.45% | 99.46% | ARMF is prevalence-dominated in TON_IoT |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Rushendra; Ramli, K.; Purnamasari, P.D.; Gunawan, T.S.; Salman, M. Cognitive Detection at Big-Data Scale: A CNN-LSTM-DQN Framework with Prioritized Experience Replay for Cross-Attack-Family Generalization and Multi-Seed Initialization Sensitivity Analysis. Big Data Cogn. Comput. 2026, 10, 239. https://doi.org/10.3390/bdcc10070239
Rushendra, Ramli K, Purnamasari PD, Gunawan TS, Salman M. Cognitive Detection at Big-Data Scale: A CNN-LSTM-DQN Framework with Prioritized Experience Replay for Cross-Attack-Family Generalization and Multi-Seed Initialization Sensitivity Analysis. Big Data and Cognitive Computing. 2026; 10(7):239. https://doi.org/10.3390/bdcc10070239
Chicago/Turabian StyleRushendra, Kalamullah Ramli, Prima Dewi Purnamasari, Teddy Surya Gunawan, and Muhammad Salman. 2026. "Cognitive Detection at Big-Data Scale: A CNN-LSTM-DQN Framework with Prioritized Experience Replay for Cross-Attack-Family Generalization and Multi-Seed Initialization Sensitivity Analysis" Big Data and Cognitive Computing 10, no. 7: 239. https://doi.org/10.3390/bdcc10070239
APA StyleRushendra, Ramli, K., Purnamasari, P. D., Gunawan, T. S., & Salman, M. (2026). Cognitive Detection at Big-Data Scale: A CNN-LSTM-DQN Framework with Prioritized Experience Replay for Cross-Attack-Family Generalization and Multi-Seed Initialization Sensitivity Analysis. Big Data and Cognitive Computing, 10(7), 239. https://doi.org/10.3390/bdcc10070239

