Description
Hi, thanks for the great work! I strictly followed the paper and official code settings to reproduce three tasks, but observed significant gaps on two of them:
| Task |
Paper |
Reproduced |
| Swap Blocks |
67% |
37% |
| Observe and Pick Up |
4% |
4% ✅ |
| Rearrange Blocks |
89% |
59% |
The correct reproduction on Observe and Pick Up suggests my environment is set up properly. The ~30% gaps on the other two tasks seem beyond normal seed variance.
Could you kindly clarify if there are any undocumented hyperparameters, preprocessing/evaluation details, or known hardware/library sensitivities for these two tasks? Training logs or checkpoints for cross-checking would also be greatly appreciated.
Thank you! 🙏
Description
Hi, thanks for the great work! I strictly followed the paper and official code settings to reproduce three tasks, but observed significant gaps on two of them:
The correct reproduction on Observe and Pick Up suggests my environment is set up properly. The ~30% gaps on the other two tasks seem beyond normal seed variance.
Could you kindly clarify if there are any undocumented hyperparameters, preprocessing/evaluation details, or known hardware/library sensitivities for these two tasks? Training logs or checkpoints for cross-checking would also be greatly appreciated.
Thank you! 🙏