Hi maintainers, thank you for releasing agent-as-a-judge. This is a very useful evaluation framework.
While integrating our own agentic system, we noticed that the repository currently includes only a single OpenHands example:
benchmark/workspaces/OpenHands/39_Drug_Response_Prediction_SVM_GDSC_ML
benchmark/trajectories/OpenHands/39_Drug_Response_Prediction_SVM_GDSC_ML.json
Could you please clarify:
- Whether additional OpenHands workspaces and trajectories exist?
- If so, are there plans to release them publicly?
Understanding this would help us align our own agent outputs and trajectory collection more accurately with the intended evaluation setup.
Thanks again for your great work!
Hi maintainers, thank you for releasing agent-as-a-judge. This is a very useful evaluation framework.
While integrating our own agentic system, we noticed that the repository currently includes only a single OpenHands example:
benchmark/workspaces/OpenHands/39_Drug_Response_Prediction_SVM_GDSC_ML
benchmark/trajectories/OpenHands/39_Drug_Response_Prediction_SVM_GDSC_ML.json
Could you please clarify:
Understanding this would help us align our own agent outputs and trajectory collection more accurately with the intended evaluation setup.
Thanks again for your great work!