I noticed that the CyberGym results currently displayed for MDASH do not appear to align with the updated results being referenced for MAI Cyber Flash 1.
Could we please investigate why there is a discrepancy between the published numbers? Seeing different performance figures in different locations makes it difficult to determine which results are the most current and authoritative.
It would be helpful to understand:
Which benchmark results should be considered the latest.
Whether the reported figures are based on different evaluation runs, model versions, or methodologies.
If any updates are planned to ensure consistency across published materials.
Thank you for looking into this.
https://microsoft.ai/news/introducing-mai-cyber-1-flash-inside-mdash/
I noticed that the CyberGym results currently displayed for MDASH do not appear to align with the updated results being referenced for MAI Cyber Flash 1.
Could we please investigate why there is a discrepancy between the published numbers? Seeing different performance figures in different locations makes it difficult to determine which results are the most current and authoritative.
It would be helpful to understand:
Which benchmark results should be considered the latest.
Whether the reported figures are based on different evaluation runs, model versions, or methodologies.
If any updates are planned to ensure consistency across published materials.
Thank you for looking into this.
https://microsoft.ai/news/introducing-mai-cyber-1-flash-inside-mdash/