Engineering Offline Reinforcement Learning for Safe Deployment of Real-World Autonomous Systems: Learning and Verification Strategies in High-Risk Environments

While the real-world deployment of autonomous systems holds immense potential, online learning approaches are virtually impossible in high-risk environments due to unpredictable safety issues. Offline Reinforcement Learning offers a powerful solution by leveraging pre-collected data to learn and verify safe and reliable policies, enabling the deployment of autonomous systems without costly real-world trial and error.

1. The Challenge / Context: The Reinforcement Learning Dilemma in High-Risk Real-World Environments

Despite remarkable advancements in various AI fields over the past decade, the commercialization of high-risk real-world autonomous systems such as self-driving cars, robotic surgery, and power plant control remains a significant challenge. The traditional online Reinforcement Learning (RL) paradigm finds optimal policies through trial and error via continuous interaction with the environment. However, this "Exploration" process carries the following critical problems:

  • Safety Issues: Incorrect actions can lead to casualties, equipment damage, or severe financial losses. For example, unsafe exploratory behavior by a self-driving car on the road is unacceptable.
  • Cost Issues: Collecting data and experimenting in real environments incurs enormous costs and time. It often requires resetting physical robots or operating expensive vehicles repeatedly.
  • Data Inefficiency: Online RL often demands vast amounts of interaction data, which can be impractical in