In my full end-to-end evaluation, the reference trajectories contain
ObjectDetection in 21 samples. However, the model selected ObjectDetection in
only 6 samples, 5 of which were executed successfully, and none overlapped with
those reference samples. In these cases, the model generally selected
TextToBbox instead.
Would it be possible to make the functional descriptions of ObjectDetection and
TextToBbox more explicit, so that the model can distinguish between them more
reliably? I wonder whether such a clarification could improve the evaluation
results.
Thank you very much for your work!
In my full end-to-end evaluation, the reference trajectories contain
ObjectDetection in 21 samples. However, the model selected ObjectDetection in
only 6 samples, 5 of which were executed successfully, and none overlapped with
those reference samples. In these cases, the model generally selected
TextToBbox instead.
Would it be possible to make the functional descriptions of ObjectDetection and
TextToBbox more explicit, so that the model can distinguish between them more
reliably? I wonder whether such a clarification could improve the evaluation
results.
Thank you very much for your work!