Explainers: benchmarks

HumanEval benchmark explained: problems, pass@k and limits
HumanEval tests generated Python functions against hidden tests. Its score is narrower than real coding work.

SWE-bench Explained: Issues, Patches and Test Results
How SWE-bench turns GitHub issues into tested patches, and why benchmark versions affect score comparisons.