Oscillation of Episode Q0 during DDPG training
Mostra commenti meno recenti
How do I interpret this kind of Episode Q0 oscillation?

The oscillation shows a pattern like up and down and the range also increases quite regularly.
According to other docs, they're saying the Q0 is supposed to approach actual discounted future reward as long as the critic network is designed properly.
Is this kind of Q0 oscillation just evidence that my critic network is not well-designed?
Is there any solution to work it out?
I'm not sure this question is acceptable to this community because I think it's more or less a theoretical issue.
1 Commento
Heesu Kim
il 6 Apr 2021
Risposte (0)
Categorie
Scopri di più su Reinforcement Learning Toolbox in Centro assistenza e File Exchange
Community Treasure Hunt
Find the treasures in MATLAB Central and discover how the community can help you!
Start Hunting!