We need to establish clear objectives for the billiards task and for other observational tasks. Consider properties that an architecture designed according to the CP might exhibit and how we could measure the performance. As an example, after conditioning on prior video frames of bouncing billiards, one could simply generate future frames and measure MSE pixel-loss between predicted and truth. One could measure ability to generate distinct objects, maintain valid momentum conservation, etc. Of course, far simpler baselines will likely succeed on objectives like this, so we will likely have to generalize this task in order to measure relevant quantities.
We need to establish clear objectives for the billiards task and for other observational tasks. Consider properties that an architecture designed according to the CP might exhibit and how we could measure the performance. As an example, after conditioning on prior video frames of bouncing billiards, one could simply generate future frames and measure MSE pixel-loss between predicted and truth. One could measure ability to generate distinct objects, maintain valid momentum conservation, etc. Of course, far simpler baselines will likely succeed on objectives like this, so we will likely have to generalize this task in order to measure relevant quantities.