why I get a different action result every new time with same sample observations after deploying trained RL policies?

Question

0 voti

load("agent0218_300016_40000.mat","agent");

obsInfo = getObservationInfo(agent);

actInfo = getActionInfo(agent);

ResetHandle = @() myResetFunction(test_sss);

StepHandle = @(Action,LoggedSignals) myStepFunction(Action,LoggedSignals,test_sss);

envT = rlFunctionEnv(obsInfo,actInfo,StepHandle,ResetHandle);

simOpts = rlSimulationOptions('MaxSteps',size(test_sss,1));

experience = sim(envT,agent,simOpts);

ac3=squeeze(experience.Action.bs.Data);

%******************************************************************************

generatePolicyFunction(agent);

%******************************************************************************

for iii=1:size(ac3,1)

observation1=test_sss{iii,:};

action1(iii,1) = evaluatePolicy(observation1);

end

sum(abs(ac3-action1))

0 Commenti
Mostra -2 commenti meno recenti Nascondi -2 commenti meno recenti

Accedi per commentare.

Accedi per rispondere a questa domanda.

Follow Question

Answer 1

Emmanouil Tzorakoleftherakis il 23 Feb 2021

0 voti

Which agent are you using? Some agents are stochastic, meaning that the output is sampled based on probability distributions so by construction they won't give you the same result.

Another possible reason is the reset function. It seems you are saving simulation data and running inference again, but every time you call 'sim', the reset function is called first. So if there are any components that randomize initial conditions/parameters, then you are not comparing with the same data.

1 Commento
Mostra -1 commenti meno recenti Nascondi -1 commenti meno recenti

liang zhang il 2 Mar 2022

Modificato: liang zhang il 2 Mar 2022

I also encountered the same problem when I used the DDPG agent for verification, my reset function doesn't randomize initial any conditions/parameters，I guess if the trained DDPG agent also has its own noise? Shouldn't a trained agent be a fixed set of neural network parameters?

Accedi per commentare.

Answer 2

de y il 24 Feb 2021

0 voti

Thanks @Emmanouil Tzorakoleftherakis a lot.

I am using PPO and SAC agent, the same question came out. My codes indicated the agent had trainned to a satisfied and balanced result, I want to use it to decide action. But my wonder is that SIM is one of simulation way,whereas generatePolicyFunction() and evaluatePolicy is another way, my observations of every step is the same,why every running evaluatePolicy with the same observations happened , the different action result with SIM() came out. It confused me because that there didn't had any components that randomize initial conditions/parameters

0 Commenti
Mostra -2 commenti meno recenti Nascondi -2 commenti meno recenti

Accedi per commentare.

why I get a different action result every new time with same sample observations after deploying trained RL policies?

0 Commenti
Mostra -2 commenti meno recenti Nascondi -2 commenti meno recenti

Risposta accettata

1 Commento
Mostra -1 commenti meno recenti Nascondi -1 commenti meno recenti

Più risposte (1)

0 Commenti
Mostra -2 commenti meno recenti Nascondi -2 commenti meno recenti

Categorie

Prodotti

Tag

Community Treasure Hunt

why I get a different action result every new time with same sample observations after deploying trained RL policies?

0 Commenti Mostra -2 commenti meno recenti Nascondi -2 commenti meno recenti

Risposta accettata

1 Commento Mostra -1 commenti meno recenti Nascondi -1 commenti meno recenti

Più risposte (1)

0 Commenti Mostra -2 commenti meno recenti Nascondi -2 commenti meno recenti

Categorie

Prodotti

Tag

Vedere anche

Community Treasure Hunt

0 Commenti
Mostra -2 commenti meno recenti Nascondi -2 commenti meno recenti

1 Commento
Mostra -1 commenti meno recenti Nascondi -1 commenti meno recenti

0 Commenti
Mostra -2 commenti meno recenti Nascondi -2 commenti meno recenti