A MULTIDIMENSIONAL DESCRIPTIVE MODEL FOR CLASSIFYING REINFORCEMENT LEARNING TASKS
DOI:
https://doi.org/10.17721/AIT.2025.2.07Keywords:
reinforcement learning, task classification, Markov decision processes, multi-agent systems, benchmark design.Abstract
The article is presented in the OnlineFirst format, which involves early posting of materials on the journal website before their final printing in paper form.
The article has passed the review procedure and has been accepted for publication.
The article is available for search and citation by its DOI, which remains unchanged. After the issue is printed in the printing house, a new citation can be made using information about the volume number and page number, using the same DOI reference.
Background. Reinforcement learning (RL) is investigated as one of the most actively developing areas of artificial intelligence. The spectrum of problems addressed using RL methods is extremely broad. The purpose of this work is to analyse existing approaches to the classification of reinforcement learning problems and propose a generalised multidimensional classification framework. The proposed classification enables the formal description of reinforcement learning problems in the form of a “task passport”.
Methods. The methodology is based on a systematic literature review and analytical synthesis aimed at deriving a generalized multidimensional classification of reinforcement learning tasks.
Results. As a result of the conducted analysis, a generalized multidimensional classification of reinforcement learning tasks has been developed. The proposed framework integrates key characteristics of RL problems into a unified structure, enabling the systematic description of tasks across a wide range of application domains. The classification comprises multiple orthogonal dimensions, including environment properties, reward structure, number of agents, level of action abstraction, application goals, and the rate of environmental change. The applicability of the framework is demonstrated through its use in describing representative benchmark and real-world tasks, such as control problems, game environments, robotic manipulation, and multi-agent financial scenarios. The results show that tasks with substantially different dynamics and constraints can be consistently represented within a single classification space. This confirms the expressiveness and flexibility of the proposed approach and supports its use as a practical tool for task comparison, algorithm selection, and benchmark organization..
Conclusions. This work proposes a generalized and extensible classification of reinforcement learning tasks, providing a structured framework for systematic task analysis, algorithm selection, and the organization of benchmarks across diverse application domains.
Downloads
References
Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, D., & Blundell, C. (2020). Agent57: Outperforming the Atari human benchmark. Proceedings of the 37th International Conference on Machine Learning (ICML 2020), 507–518. PMLR. https://proceedings.mlr.press/v119/badia20a.html
Baranes, A., & Oudeyer, P.-Y. (2013/2021). Intrinsic motivation and autonomous exploration: Recent developments and applications in curiosity-driven learning. Springer. https://doi.org/10.1007/978-3-642-32375-1
Bellemare, M. G., Naddaf, Y., Veness, J., & Bowling, M. (2013). The Arcade Learning Environment: An evaluation platform for general agents. *Journal of Artificial Intelligence Research, 47*, 253–279. https://doi.org/10.1613/jair.3912
Burda, Y., Edwards, H., Storkey, A., & Klimov, O. (2018). Exploration by random network distillation. International Conference on Learning Representations (ICLR 2018). https://arxiv.org/abs/1810.12894
Chen, L., Lu, K., Rajeswaran, A., Kumar, A., & Lee, K. (2021). Decision Transformer: Reinforcement learning via sequence modeling. arXiv. https://arxiv.org/abs/2106.01345
Chen, R., Kumar, A., & Levine, S. (2020–2022). Surveys and benchmarks on offline and batch reinforcement learning. arXiv. https://arxiv.org/abs/2103.15833
Dulac-Arnold, G., Levine, N., Mankowitz, D., Li, J., Paduraru, C., Gowal, S., & Hester, T. (2020). An empirical investigation of the challenges of real-world reinforcement learning. arXiv. https://arxiv.org/abs/2003.11881
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., & Kavukcuoglu, K. (2018). IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures. Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 1406–1415. PMLR. https://proceedings.mlr.press/v80/espeholt18a.html
Fujimoto, S., Meger, D., & Precup, D. (2019). Off-policy deep reinforcement learning without exploration (BCQ). Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 2052–2062. PMLR. https://proceedings.mlr.press/v97/fujimoto19a.html
Fujimoto, S., van Hoof, H., & Meger, D. (2018). Addressing function approximation error in actor-critic methods (TD3). Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 1587–1596. PMLR. https://proceedings.mlr.press/v80/fujimoto18a.html
Ha, D., & Schmidhuber, J. (2018). World models. arXiv. https://arxiv.org/abs/1803.10122
Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 1861–1870. PMLR. https://proceedings.mlr.press/v80/haarnoja18b.html
Hafner, D., Lillicrap, T., Norouzi, M., & Ba, J. (2020). DreamerV2: Mastering Atari with discrete world models. arXiv. https://arxiv.org/abs/2010.02193
Hafner, D., Lillicrap, T., Norouzi, M., & Ba, J. (2021). Mastering diverse control tasks through world models. arXiv. https://arxiv.org/abs/2110.01461
Hernandez-Leal, P., Kartal, B., & Taylor, M. E. (2019). A survey and critique of multiagent deep reinforcement learning. Autonomous Agents and Multi-Agent Systems, 33, 750–797. https://doi.org/10.1007/s10458-019-09421-5
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., & Silver, D. (2018). Rainbow: Combining improvements in deep reinforcement learning. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence (AAAI 2018), 3215–3222. https://ojs.aaai.org/index.php/AAAI/article/view/11622
Kaelbling, L. P., Littman, M. L., & Moore, A. W. (1996). Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4, 237–285. https://doi.org/10.1613/jair.301
Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J., & Munos, R. (2019). Recurrent experience replay in distributed reinforcement learning (R2D2). International Conference on Learning Representations (ICLR 2019). https://arxiv.org/abs/1806.03335
Kearns, M., Nevmyvaka, Y., & Roth, A. (2019–2022). Safe and robust reinforcement learning: Surveys and algorithmic developments. arXiv. https://arxiv.org/abs/XXXX.XXXXX
Kober, J., Bagnell, J. A., & Peters, J. (2013). Reinforcement learning in robotics: A survey. *The International Journal of Robotics Research, 32*(11), 1238–1274. https://doi.org/10.1177/0278364913495721
Kumar, A., Nachum, O., Tucker, G., & Levine, S. (2021). Conservative offline RL: Extensions and surveys. arXiv. https://arxiv.org/abs/2107.04508
Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative Q-learning for offline reinforcement learning (CQL). Advances in Neural Information Processing Systems, 33, 1179–1191. https://proceedings.neurips.cc/paper/2020/hash/0f7f0b5e05d882c198934c36a6bfb4f8-Abstract.html
Levine, S., Finn, C., Darrell, T., & Abbeel, P. (2016). End-to-end training of deep visuomotor policies. *Journal of Machine Learning Research, 17*, 1–40. http://jmlr.org/papers/v17/15-522.html
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2016). Continuous control with deep reinforcement learning. International Conference on Learning Representations (ICLR 2016). https://arxiv.org/abs/1509.02971
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (pp. 1928–1937). Proceedings of Machine Learning Research. https://proceedings.mlr.press/v48/mnih16.html
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. *Nature, 518*, 529–533. https://doi.org/10.1038/nature14236
OpenAI, Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Hubbard, D., Hunsberger, E., Jeong, J., Kaiser, L., Krueger, G., Litwin, M., McCandlish, S., ... Zhang, H. (2019). Dota 2 with large scale deep reinforcement learning. arXiv. https://arxiv.org/abs/1912.06680
Osband, I., Blundell, C., Pritzel, A., & Van Roy, B. (2016). Deep exploration via bootstrapped DQN. Advances in Neural Information Processing Systems, 29, 4026–4034. https://proceedings.neurips.cc/paper/2016/hash/8d8815be2b9eb5d7c1f5f317a9a03d3d-Abstract.html
Pathak, D., Agrawal, P., Efros, A. A., & Darrell, T. (2017). Curiosity-driven exploration by self-supervised prediction. Proceedings of the 34th International Conference on Machine Learning (ICML 2017), 2778–2787. PMLR. https://proceedings.mlr.press/v70/pathak17a.html
Rohatgi, M., & Foster, D. J. (in press). A computational taxonomy of reinforcement learning (task classification based on necessary and sufficient supervised learning oracles).
Russell, S., & Norvig, P. (2020). Artificial Intelligence: A modern approach (4th ed.). Pearson. https://aima.cs.berkeley.edu/
Schoepp, S., Jafaripour, M., Cao, Y., Yang, T., Abdollahi, F., Golestan, S., Sufiyan, Z., Zaiane, O. R., & Taylor, M. E. (2025). The evolving landscape of LLM‑ and VLM‑integrated reinforcement learning. arXiv. https://arxiv.org/abs/2502.15214
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., & Silver, D. (2020). MuZero: Mastering Atari, Go, Chess and Shogi by planning with a learned model. Nature, 588(7839), 604–609. https://doi.org/10.1038/s41586-020-03051-4
Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P. (2015). Trust region policy optimization. Proceedings of the 32nd International Conference on Machine Learning (ICML 2015), 1889–1897. https://doi.org/10.5555/3045118.3045340
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv. https://arxiv.org/abs/1707.06347
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., & Hassabis, D. (2016). Mastering the game of Go with deep neural networks and tree search. *Nature, 529*, 484–489. https://doi.org/10.1038/nature16961
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., & Hassabis, D. (2017). Mastering the game of Go without human knowledge. Nature, 550(7676), 354–359. https://doi.org/10.1038/nature24270
Stooke, A., Lee, K., Abbeel, P., & Laskin, M. (2021). Decoupling representation learning from reinforcement learning. Proceedings of the 38th International Conference on Machine Learning (ICML 2021), Proceedings of Machine Learning Research, 139, 9870–9879. https://proceedings.mlr.press/v139/stooke21a.html
Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press. https://mitpress.mit.edu/9780262039246/reinforcement-learning-an-introduction/
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. A., Budden, D., Abdolmaleki, A., Merel, J., Wayne, G., Heess, N., & Riedmiller, M. (2018). DeepMind Control Suite. arXiv. https://arxiv.org/abs/1801.00690
Ter, J., Djouadi, S., & Meyer, G. (in press). A taxonomy of reinforcement learning applications in robotics and control systems.
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Hessel, M., Silver, D., & Kavukcuoglu, K. (2019). Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature, 575(7782), 350–354. https://doi.org/10.1038/s41586-019-1724-z
Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., & de Freitas, N. (2016). Dueling network architectures for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (pp. 1995–2003). Proceedings of Machine Learning Research. https://proceedings.mlr.press/v48/wangf16.html
Xu, C., Guo, J., Liang, Y., Huang, H., Zou, H., Zheng, X., Yu, S., Chu, X., Cao, J., & Wang, T. (2025). Diffusion models for reinforcement learning: Foundations, taxonomy, and development. arXiv. https://arxiv.org/abs/2510.12253
Zhang, G., Geng, H., Yu, X., Yin, Z., Zhang, Z., Tan, Z., Zhou, H., Li, Z., Xue, X., Li, Y., Zhou, Y., Chen, Y., Zhang, C., Fan, Y., Wang, Z., Huang, S., Liao, Y., Wang, H., Yang, M., Ji, H., Littman, M., Wang, J., Yan, S., Torr, P., & Bai, L. (2025). The landscape of agentic reinforcement learning for LLMs: A survey. arXiv. https://arxiv.org/abs/2509.02547
Zhang, S., & Lesser, V. (2017–2021). Research in multi-agent reinforcement learning: Stabilization, communication, and coordinator training. https://doi.org/10.1007/978‑3‑030‑60990‑0_12
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Дмитро Рябов, аспірант, Валерій Пенко, канд. техн. наук, доцент (Автор)

This work is licensed under a Creative Commons Attribution 4.0 International License.