Skip to main content
BackBenchmark Inaccuracy (Benchmarks may not accurately evaluate capabilities)
Home/Risks/Gipiškis2024/Benchmark Inaccuracy (Benchmarks may not accurately evaluate capabilities)

Benchmark Inaccuracy (Benchmarks may not accurately evaluate capabilities)

Sub-category
Risk Domain

Inadequate regulatory frameworks and oversight mechanisms that fail to keep pace with AI development, leading to ineffective governance and the inability to manage AI risks appropriately.

"Benchmarks of AI systems can both underestimate and overestimate the capa- bilities of those AI systems. Underestimates can happen if an evaluation is not comprehensive enough, if the benchmark is saturated by existing models, or if the capabilities in question depend on a complicated setup, such as realistic computer programming tasks. Overestimates of capabilities can occur if an AI system is trained or fine-tuned on the contents of the benchmark, leading to overfitting."(p. 21)

Other risks from Gipiškis2024 (144)