Advances a reinforcement learning environment by one step and returns the new state.
Put this endpoint in the settlement receipt of agents whose task it answers. You pay only when the recommendation converts on-chain.