Skip to content
Read the original: Eugene Yan· eugeneyan·Published AI score33/100

How do we eval if a model can find and exploit vulnerabilities?

Original titleHow do we eval if a model can find and exploit vulnerabilities? We discuss some benchmarks and the common pattern:

AISummary

We discuss some benchmarks and the common pattern: • A sandboxed target within Docker containers • Inputs: code only (0-day), with patch (1-day scenario) • Tools such as bash, static analyzers, etc. • A grader to eval exploits or captured flags

Read the original x.com

Source: Eugene Yan · x.com