Hello!
I’ve been reading the Agent-as-a-Judge paper and found it quite interesting. Thanks for the work!
One thing I’m unclear about is how this framework would generalize to real-world queries that don’t come with predefined requirements or dependencies (unlike in the DevAI benchmark).
The paper doesn’t go into much detail on this. How do you think requirements could be collected or inferred in a more open-ended, practical setting? Would it be an additional agent step, or maybe structured prompting?
Curious to hear what you think about this and if you have run some experiments.
Thanks,
Hello!
I’ve been reading the Agent-as-a-Judge paper and found it quite interesting. Thanks for the work!
One thing I’m unclear about is how this framework would generalize to real-world queries that don’t come with predefined requirements or dependencies (unlike in the DevAI benchmark).
The paper doesn’t go into much detail on this. How do you think requirements could be collected or inferred in a more open-ended, practical setting? Would it be an additional agent step, or maybe structured prompting?
Curious to hear what you think about this and if you have run some experiments.
Thanks,