Reproduction
CollectiveRewardWrapper.observation_spec() advertises a scalar float64 COLLECTIVE_REWARD, but _get_timestep uses an untyped np.sum.
On main at 992db43f45ec856d5f0d8cb2936cdb4501b8fb9b, integer rewards produce an integer observation, float32 rewards produce float32, and a standard dm_env.restart produces None for the collective reward. All fail the wrapper's own observation-spec validation.
The accumulation type can also change the numerical result: [2**62, 2**62] stored as int64 sums to a negative value, and two float16 values of 65504 overflow even though their total is representable in the advertised float64 type.
Expected behavior
Accumulate the derived collective reward in its declared float64 type. With no reward on a FIRST timestep, expose a zero collective observation while preserving the underlying reward=None, discount, observations and step type.
Keep the existing float64 calculation, reward ownership and reset/step argument forwarding. This only changes the derived observation, not the individual rewards. Regression coverage should validate the complete wrapper lifecycle against its actual specifications and retain existing float64 inputs as controls.
Reproduction
CollectiveRewardWrapper.observation_spec()advertises a scalar float64COLLECTIVE_REWARD, but_get_timestepuses an untypednp.sum.On main at
992db43f45ec856d5f0d8cb2936cdb4501b8fb9b, integer rewards produce an integer observation, float32 rewards produce float32, and a standarddm_env.restartproducesNonefor the collective reward. All fail the wrapper's own observation-spec validation.The accumulation type can also change the numerical result:
[2**62, 2**62]stored as int64 sums to a negative value, and two float16 values of 65504 overflow even though their total is representable in the advertised float64 type.Expected behavior
Accumulate the derived collective reward in its declared float64 type. With no reward on a FIRST timestep, expose a zero collective observation while preserving the underlying
reward=None, discount, observations and step type.Keep the existing float64 calculation, reward ownership and reset/step argument forwarding. This only changes the derived observation, not the individual rewards. Regression coverage should validate the complete wrapper lifecycle against its actual specifications and retain existing float64 inputs as controls.