Q Learning and Pong and stuff Methodology Update position of ball (and paddle? Based on previous state/action) From this new position we derive the state. State is entirely based on the position and velocity of the ball and paddle, which we know at any instant.