Reward Machines for Deep RL in Noisy and Uncertain Environments


報酬マシンは表形式とディープ RL の両方の設定で採用されてきましたが、通常は報酬関数の構成要素を形成するドメイン固有の語彙のグラウンドトゥルース解釈に依存していました。
このペーパーでは、騒がしく不確実な環境におけるディープ RL 用の報酬マシンの使用について検討します。
我々は、この問題を POMDP として特徴付け、ドメイン固有の語彙の不確実な解釈の下でタスク構造を活用する一連の RL アルゴリズムを提案します。


Reward Machines provide an automata-inspired structure for specifying instructions, safety constraints, and other temporally extended reward-worthy behaviour. By exposing complex reward function structure, they enable counterfactual learning updates that have resulted in impressive sample efficiency gains. While Reward Machines have been employed in both tabular and deep RL settings, they have typically relied on a ground-truth interpretation of the domain-specific vocabulary that form the building blocks of the reward function. Such ground-truth interpretations can be elusive in many real-world settings, due in part to partial observability or noisy sensing. In this paper, we explore the use of Reward Machines for Deep RL in noisy and uncertain environments. We characterize this problem as a POMDP and propose a suite of RL algorithms that leverage task structure under uncertain interpretation of domain-specific vocabulary. Theoretical analysis exposes pitfalls in naive approaches to this problem, while experimental results show that our algorithms successfully leverage task structure to improve performance under noisy interpretations of the vocabulary. Our results provide a general framework for exploiting Reward Machines in partially observable environments.


著者 Andrew C. Li,Zizhao Chen,Toryn Q. Klassen,Pashootan Vaezipoor,Rodrigo Toro Icarte,Sheila A. McIlraith
発行日 2024-06-17 16:39:08+00:00
arxivサイト arxiv_id(pdf)

提供元, 利用サービス, Google

カテゴリー: cs.AI, cs.FL, cs.LG, F.4.3 パーマリンク