We propose to build
fresh one-time questions to evaluate LLMs instead of relying
on static benchmarks.
This is one of your proposal in the paper. It might easy for coding/math problems as they can be generated from almost infinite combinations.
Is there an active community in pulling out such one-time evaluation for other domains?
This is one of your proposal in the paper. It might easy for coding/math problems as they can be generated from almost infinite combinations.
Is there an active community in pulling out such one-time evaluation for other domains?