Oscillation of Episode Q0 during DDPG training

Heesu Kim

6 Apr 2021

0 Risposte

Aggiornato 6 Apr 2021

6 Visualizzazioni (30 giorni)

Accedi per rispondere a questa domanda.

Follow Question

Accedi per rispondere a questa domanda.

Follow Question

Mostra commenti meno recenti

0 voti

How do I interpret this kind of Episode Q0 oscillation?

The oscillation shows a pattern like up and down and the range also increases quite regularly.

According to other docs, they're saying the Q0 is supposed to approach actual discounted future reward as long as the critic network is designed properly.

Is this kind of Q0 oscillation just evidence that my critic network is not well-designed?

Is there any solution to work it out?

I'm not sure this question is acceptable to this community because I think it's more or less a theoretical issue.

1 Commento
Mostra -1 commenti meno recenti Nascondi -1 commenti meno recenti

Heesu Kim il 6 Apr 2021

As a side note, I'm using DDPG + LSTM model that RL toolbox provides

Accedi per commentare.

Accedi per rispondere a questa domanda.

Follow Question

Risposte (0)

Accedi per rispondere a questa domanda.

Categorie

Scopri di più su Reinforcement Learning Toolbox in Centro assistenza e File Exchange

Prodotti

Reinforcement Learning Toolbox

Release

R2021a

Tag

il 6 Apr 2021

il 6 Apr 2021

Community Treasure Hunt

Find the treasures in MATLAB Central and discover how the community can help you!

Start Hunting!

Translated by

Oscillation of Episode Q0 during DDPG training

1 Commento Mostra -1 commenti meno recenti Nascondi -1 commenti meno recenti

Risposte (0)

Categorie

Prodotti

Release

Tag

Vedere anche

Community Treasure Hunt

1 Commento
Mostra -1 commenti meno recenti Nascondi -1 commenti meno recenti