-
-ergodic dynamics, such an average may differ arbitrarily from what the individual agent experiences as it lives out one trajectory over time. Furthermore, RL typically seeks a time-invariant policy that
-
reset, failures can be irreversible, and the real world keeps changing. Standard RL finds policies by optimizing over many hypothetical futures. Under non-ergodic dynamics, such an average may differ
Enter an email to receive alerts for computer-"https:"-"https:"-"https:" "https:" "KTH" "DIFFER" positions