RISK-SENSITIVE Q-LEARNING IN A TWO-AGENT MARKOV ATTACK–DEFENSE MODEL

Authors

DOI:

https://doi.org/10.31891/csit-2026-3-4

Keywords:

reinforcement learning, machine learning, Markov model, risk-sensitive Q-learning, time to compromise, security testing, cybersecurity, cyberattack, cyberdefense, artificial intelligence

Abstract

A risk-sensitive method for forming a web application security testing policy based on entropic-risk Q-learning in a two-agent Markov attack–defense model is proposed. The study addresses the need to detect short critical compromise trajectories under partial observability, noisy estimation of defensive attributes, and stochastic changes in the defensive profile of the environment. Unlike risk-neutral approaches, the proposed method employs a dense reward oriented towards the time to compromise (TTC) and a post-episode evaluation of tail losses in terms of the conditional value at risk at the 0.95 confidence level (CVaR0.95), where the risk-sensitivity parameter β of the entropic-risk operator controls the degree of risk aversion. Two modes of the simulation environment are introduced, namely the controlled environment Env1 and the stochastic environment Env2, which are used to analyze policy stability under different levels of information uncertainty. The ToN-IoT and CIC-IDS2017 datasets were used for the statistical parameterization of the environments. Aggregation of semantically related MAL attack steps into web components reduces the dimensionality of the state representation by a factor of 8–12, which makes the training of the risk-sensitive policy computationally manageable. The obtained results show that the proposed method provides the lowest average time to compromise among all the compared approaches. In Env2 the TTC value decreased from 36.9±1.5 to 27.9±1.1 steps for ToN-IoT and from 39.7±1.7 to 31.4±1.3 steps for CIC-IDS2017, the CVaR0.95 of normalized tail losses decreased by approximately 31 %, and the number of episodes required for policy convergence was reduced by approximately 35.4%. The most effective mode is achieved at β ≈ -0.6. The practical significance lies in the possibility of applying the method for automated penetration testing and for evaluating the robustness of web application protection mechanisms.

Downloads

Published

2026-09-30

How to Cite

PRYTULA, A., & KUPERSHTEIN, L. (2026). RISK-SENSITIVE Q-LEARNING IN A TWO-AGENT MARKOV ATTACK–DEFENSE MODEL. Computer Systems and Information Technologies, (3), 39–48. https://doi.org/10.31891/csit-2026-3-4