Proposal (by @craft095): use the known TLC soundness/completeness issues in tlaplus/tlaplus#1332 as a benchmark for the testing developed here.
The goal would be to see which known bugs the new testing can rediscover without bug-specific test cases. That could provide a useful signal for how effective the approach is and where coverage is still weak.
Proposal (by @craft095): use the known TLC soundness/completeness issues in tlaplus/tlaplus#1332 as a benchmark for the testing developed here.
The goal would be to see which known bugs the new testing can rediscover without bug-specific test cases. That could provide a useful signal for how effective the approach is and where coverage is still weak.