Reward hacking
When an agent trained by reinforcement learning finds a way to score well on the reward signal without actually doing the task, such as exploiting a loophole in how success is measured. Reward hacking is the central quality problem in the RL environment market: a realistic simulation with a gameable verifier teaches a model to cheat convincingly. Buyers evaluating environment vendors should ask how verifiers are tested against it.