A MULTIDIMENSIONAL DESCRIPTIVE MODEL FOR CLASSIFYING REINFORCEMENT LEARNING TASKS

Authors

  • Dmitry Riabov, Ph.D Student Odessa State University, Odessa, Ukraine Author
  • Valeriy Penko, PhD (Engin.), Assoc. Prof Odessa State University, Odessa, Ukraine Author

DOI:

https://doi.org/10.17721/AIT.2025.2.07

Keywords:

reinforcement learning, task classification, Markov decision processes, multi-agent systems, benchmark design.

Abstract

The article is presented in the OnlineFirst format, which involves early posting of materials on the journal website before their final printing in paper form.

The article has passed the review procedure and has been accepted for publication.

The article is available for search and citation by its DOI, which remains unchanged. After the issue is printed in the printing house, a new citation can be made using information about the volume number and page number, using the same DOI reference.

Background. Reinforcement learning (RL) is investigated as one of the most actively developing areas of artificial intelligence. The spectrum of problems addressed using RL methods is extremely broad. The purpose of this work is to analyse existing approaches to the classification of reinforcement learning problems and propose a generalised multidimensional classification framework.  The proposed classification enables the formal description of reinforcement learning problems in the form of a “task passport”.

Methods. The methodology is based on a systematic literature review and analytical synthesis aimed at deriving a generalized multidimensional classification of reinforcement learning tasks.

Results. As a result of the conducted analysis, a generalized multidimensional classification of reinforcement learning tasks has been developed. The proposed framework integrates key characteristics of RL problems into a unified structure, enabling the systematic description of tasks across a wide range of application domains. The classification comprises multiple orthogonal dimensions, including environment properties, reward structure, number of agents, level of action abstraction, application goals, and the rate of environmental change. The applicability of the framework is demonstrated through its use in describing representative benchmark and real-world tasks, such as control problems, game environments, robotic manipulation, and multi-agent financial scenarios. The results show that tasks with substantially different dynamics and constraints can be consistently represented within a single classification space. This confirms the expressiveness and flexibility of the proposed approach and supports its use as a practical tool for task comparison, algorithm selection, and benchmark organization..

Conclusions. This work proposes a generalized and extensible classification of reinforcement learning tasks, providing a structured framework for systematic task analysis, algorithm selection, and the organization of benchmarks across diverse application domains.

Downloads

Download data is not yet available.

Author Biographies

  • Dmitry Riabov, Ph.D Student, Odessa State University, Odessa, Ukraine

    ORCID ID: 0009-0007-7426-0357

  • Valeriy Penko, PhD (Engin.), Assoc. Prof, Odessa State University, Odessa, Ukraine

    ORCID ID: 0000-0002-0190-6694

References

Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, D., & Blundell, C. (2020). Agent57: Outperforming the Atari human benchmark. Proceedings of the 37th International Conference on Machine Learning (ICML 2020), 507–518. PMLR. https://proceedings.mlr.press/v119/badia20a.html

Baranes, A., & Oudeyer, P.-Y. (2013/2021). Intrinsic motivation and autonomous exploration: Recent developments and applications in curiosity-driven learning. Springer. https://doi.org/10.1007/978-3-642-32375-1

Bellemare, M. G., Naddaf, Y., Veness, J., & Bowling, M. (2013). The Arcade Learning Environment: An evaluation platform for general agents. *Journal of Artificial Intelligence Research, 47*, 253–279. https://doi.org/10.1613/jair.3912

Burda, Y., Edwards, H., Storkey, A., & Klimov, O. (2018). Exploration by random network distillation. International Conference on Learning Representations (ICLR 2018). https://arxiv.org/abs/1810.12894

Chen, L., Lu, K., Rajeswaran, A., Kumar, A., & Lee, K. (2021). Decision Transformer: Reinforcement learning via sequence modeling. arXiv. https://arxiv.org/abs/2106.01345

Chen, R., Kumar, A., & Levine, S. (2020–2022). Surveys and benchmarks on offline and batch reinforcement learning. arXiv. https://arxiv.org/abs/2103.15833

Dulac-Arnold, G., Levine, N., Mankowitz, D., Li, J., Paduraru, C., Gowal, S., & Hester, T. (2020). An empirical investigation of the challenges of real-world reinforcement learning. arXiv. https://arxiv.org/abs/2003.11881

Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., Legg, S., & Kavukcuoglu, K. (2018). IMPALA: Scalable distributed deep-RL with importance weighted actor-learner architectures. Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 1406–1415. PMLR. https://proceedings.mlr.press/v80/espeholt18a.html

Fujimoto, S., Meger, D., & Precup, D. (2019). Off-policy deep reinforcement learning without exploration (BCQ). Proceedings of the 36th International Conference on Machine Learning (ICML 2019), 2052–2062. PMLR. https://proceedings.mlr.press/v97/fujimoto19a.html

Fujimoto, S., van Hoof, H., & Meger, D. (2018). Addressing function approximation error in actor-critic methods (TD3). Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 1587–1596. PMLR. https://proceedings.mlr.press/v80/fujimoto18a.html

Ha, D., & Schmidhuber, J. (2018). World models. arXiv. https://arxiv.org/abs/1803.10122

Haarnoja, T., Zhou, A., Abbeel, P., & Levine, S. (2018). Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. Proceedings of the 35th International Conference on Machine Learning (ICML 2018), 1861–1870. PMLR. https://proceedings.mlr.press/v80/haarnoja18b.html

Hafner, D., Lillicrap, T., Norouzi, M., & Ba, J. (2020). DreamerV2: Mastering Atari with discrete world models. arXiv. https://arxiv.org/abs/2010.02193

Hafner, D., Lillicrap, T., Norouzi, M., & Ba, J. (2021). Mastering diverse control tasks through world models. arXiv. https://arxiv.org/abs/2110.01461

Hernandez-Leal, P., Kartal, B., & Taylor, M. E. (2019). A survey and critique of multiagent deep reinforcement learning. Autonomous Agents and Multi-Agent Systems, 33, 750–797. https://doi.org/10.1007/s10458-019-09421-5

Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., & Silver, D. (2018). Rainbow: Combining improvements in deep reinforcement learning. Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence (AAAI 2018), 3215–3222. https://ojs.aaai.org/index.php/AAAI/article/view/11622

Kaelbling, L. P., Littman, M. L., & Moore, A. W. (1996). Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4, 237–285. https://doi.org/10.1613/jair.301

Kapturowski, S., Ostrovski, G., Dabney, W., Quan, J., & Munos, R. (2019). Recurrent experience replay in distributed reinforcement learning (R2D2). International Conference on Learning Representations (ICLR 2019). https://arxiv.org/abs/1806.03335

Kearns, M., Nevmyvaka, Y., & Roth, A. (2019–2022). Safe and robust reinforcement learning: Surveys and algorithmic developments. arXiv. https://arxiv.org/abs/XXXX.XXXXX

Kober, J., Bagnell, J. A., & Peters, J. (2013). Reinforcement learning in robotics: A survey. *The International Journal of Robotics Research, 32*(11), 1238–1274. https://doi.org/10.1177/0278364913495721

Kumar, A., Nachum, O., Tucker, G., & Levine, S. (2021). Conservative offline RL: Extensions and surveys. arXiv. https://arxiv.org/abs/2107.04508

Kumar, A., Zhou, A., Tucker, G., & Levine, S. (2020). Conservative Q-learning for offline reinforcement learning (CQL). Advances in Neural Information Processing Systems, 33, 1179–1191. https://proceedings.neurips.cc/paper/2020/hash/0f7f0b5e05d882c198934c36a6bfb4f8-Abstract.html

Levine, S., Finn, C., Darrell, T., & Abbeel, P. (2016). End-to-end training of deep visuomotor policies. *Journal of Machine Learning Research, 17*, 1–40. http://jmlr.org/papers/v17/15-522.html

Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., & Wierstra, D. (2016). Continuous control with deep reinforcement learning. International Conference on Learning Representations (ICLR 2016). https://arxiv.org/abs/1509.02971

Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., & Kavukcuoglu, K. (2016). Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (pp. 1928–1937). Proceedings of Machine Learning Research. https://proceedings.mlr.press/v48/mnih16.html

Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. *Nature, 518*, 529–533. https://doi.org/10.1038/nature14236

OpenAI, Berner, C., Brockman, G., Chan, B., Cheung, V., Debiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., Hubbard, D., Hunsberger, E., Jeong, J., Kaiser, L., Krueger, G., Litwin, M., McCandlish, S., ... Zhang, H. (2019). Dota 2 with large scale deep reinforcement learning. arXiv. https://arxiv.org/abs/1912.06680

Osband, I., Blundell, C., Pritzel, A., & Van Roy, B. (2016). Deep exploration via bootstrapped DQN. Advances in Neural Information Processing Systems, 29, 4026–4034. https://proceedings.neurips.cc/paper/2016/hash/8d8815be2b9eb5d7c1f5f317a9a03d3d-Abstract.html

Pathak, D., Agrawal, P., Efros, A. A., & Darrell, T. (2017). Curiosity-driven exploration by self-supervised prediction. Proceedings of the 34th International Conference on Machine Learning (ICML 2017), 2778–2787. PMLR. https://proceedings.mlr.press/v70/pathak17a.html

Rohatgi, M., & Foster, D. J. (in press). A computational taxonomy of reinforcement learning (task classification based on necessary and sufficient supervised learning oracles).

Russell, S., & Norvig, P. (2020). Artificial Intelligence: A modern approach (4th ed.). Pearson. https://aima.cs.berkeley.edu/

Schoepp, S., Jafaripour, M., Cao, Y., Yang, T., Abdollahi, F., Golestan, S., Sufiyan, Z., Zaiane, O. R., & Taylor, M. E. (2025). The evolving landscape of LLM‑ and VLM‑integrated reinforcement learning. arXiv. https://arxiv.org/abs/2502.15214

Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., & Silver, D. (2020). MuZero: Mastering Atari, Go, Chess and Shogi by planning with a learned model. Nature, 588(7839), 604–609. https://doi.org/10.1038/s41586-020-03051-4

Schulman, J., Levine, S., Abbeel, P., Jordan, M., & Moritz, P. (2015). Trust region policy optimization. Proceedings of the 32nd International Conference on Machine Learning (ICML 2015), 1889–1897. https://doi.org/10.5555/3045118.3045340

Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv. https://arxiv.org/abs/1707.06347

Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T., Leach, M., Kavukcuoglu, K., Graepel, T., & Hassabis, D. (2016). Mastering the game of Go with deep neural networks and tree search. *Nature, 529*, 484–489. https://doi.org/10.1038/nature16961

Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., Chen, Y., Lillicrap, T., Hui, F., Sifre, L., van den Driessche, G., Graepel, T., & Hassabis, D. (2017). Mastering the game of Go without human knowledge. Nature, 550(7676), 354–359. https://doi.org/10.1038/nature24270

Stooke, A., Lee, K., Abbeel, P., & Laskin, M. (2021). Decoupling representation learning from reinforcement learning. Proceedings of the 38th International Conference on Machine Learning (ICML 2021), Proceedings of Machine Learning Research, 139, 9870–9879. https://proceedings.mlr.press/v139/stooke21a.html

Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press. https://mitpress.mit.edu/9780262039246/reinforcement-learning-an-introduction/

Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., Casas, D. A., Budden, D., Abdolmaleki, A., Merel, J., Wayne, G., Heess, N., & Riedmiller, M. (2018). DeepMind Control Suite. arXiv. https://arxiv.org/abs/1801.00690

Ter, J., Djouadi, S., & Meyer, G. (in press). A taxonomy of reinforcement learning applications in robotics and control systems.

Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Hessel, M., Silver, D., & Kavukcuoglu, K. (2019). Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature, 575(7782), 350–354. https://doi.org/10.1038/s41586-019-1724-z

Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., & de Freitas, N. (2016). Dueling network architectures for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning (pp. 1995–2003). Proceedings of Machine Learning Research. https://proceedings.mlr.press/v48/wangf16.html

Xu, C., Guo, J., Liang, Y., Huang, H., Zou, H., Zheng, X., Yu, S., Chu, X., Cao, J., & Wang, T. (2025). Diffusion models for reinforcement learning: Foundations, taxonomy, and development. arXiv. https://arxiv.org/abs/2510.12253

Zhang, G., Geng, H., Yu, X., Yin, Z., Zhang, Z., Tan, Z., Zhou, H., Li, Z., Xue, X., Li, Y., Zhou, Y., Chen, Y., Zhang, C., Fan, Y., Wang, Z., Huang, S., Liao, Y., Wang, H., Yang, M., Ji, H., Littman, M., Wang, J., Yan, S., Torr, P., & Bai, L. (2025). The landscape of agentic reinforcement learning for LLMs: A survey. arXiv. https://arxiv.org/abs/2509.02547

Zhang, S., & Lesser, V. (2017–2021). Research in multi-agent reinforcement learning: Stabilization, communication, and coordinator training. https://doi.org/10.1007/978‑3‑030‑60990‑0_12

Downloads

Published

2026-05-06

Issue

Section

Reviews and discussions

How to Cite

A MULTIDIMENSIONAL DESCRIPTIVE MODEL FOR CLASSIFYING REINFORCEMENT LEARNING TASKS. (2026). Advanced Information Technology, 1(2(5). https://doi.org/10.17721/AIT.2025.2.07