HumanEval is a code-generation benchmark built from 164 handwritten Python problems. A model receives a function signature and docstring, generates code, and is judged by whether that code passes unit tests.
OpenAI introduced it in the 2021 paper Evaluating Large Language Models Trained on Code and published the harness in the openai/human-eval repository.
On this page
What is in HumanEval?
Each problem describes a compact programming task. The prompt includes a function signature, docstring and sometimes examples. The package contains canonical solutions and tests used for evaluation.
The tasks cover string handling, arithmetic, lists and simple algorithms. They do not reproduce the dependency graph, changing requirements or review process of a production repository. That narrowness helps when the question is “can the model synthesize a function that satisfies these tests?” It is a limitation when the score is presented as general software-engineering ability.
What does pass@k mean?
A model can generate several candidate solutions. Pass@k estimates the probability that at least one of the top k sampled candidates passes.
- pass@1 asks whether one sampled answer succeeds.
- pass@10 asks whether at least one of ten succeeds.
- Higher k gives the model more attempts.
Pass@1 and pass@10 answer different questions. Temperature, sample count and harness settings affect comparisons. The paper describes an unbiased estimator because simply counting any success can distort results when sample counts vary.
How is generated code evaluated?
The harness executes generated Python against tests. OpenAI’s repository warns users not to run untrusted model-generated code outside a robust security sandbox. Its execution helper is disabled by default for this reason.
A reproduction therefore needs isolation, resource controls and a disposable environment. This is an operational requirement, not a footnote.
Why published scores differ
A comparable result should disclose:
- exact model and checkpoint;
- original HumanEval or a modified variant;
- temperature and sample count;
- pass@1, pass@10 or another k;
- prompt formatting;
- timeouts and execution rules;
- contamination controls;
- evaluation harness version.
HumanEval+, EvalPlus and related suites may extend the tests. They are not interchangeable labels.
The contamination question
HumanEval has been public for years. Problems, solutions and discussions may appear in training data. A model can improve because it saw close variants, not only because it generalised the skill.
The original paper discussed filtering and evaluation design, but later comparisons still need a contamination statement. A score table without one leaves a material question unanswered.
What HumanEval does not measure
HumanEval does not directly test whether a model can navigate a large repository, choose a safe dependency, debug an integration, preserve an API contract across files, review a pull request or resolve an ambiguous requirement.
Repository-level benchmarks ask some of those questions but add their own environment and grading choices. No single benchmark covers “coding” as one skill.
How to read a result
Start with the exact benchmark name and metric. Read the model card or technical report for the harness and sampling setup. Prefer comparisons run by one evaluator under one protocol. Treat vendor-run results as company-reported unless independently reproduced.
A careful sentence says: “Model A recorded X pass@1 on original HumanEval under the vendor’s stated setup.” It should not become “Model A writes correct software X percent of the time.”
HumanEval remains useful because its task and success condition are legible. Its value comes from asking one narrow question consistently.
For other evaluation methods, browse the benchmark explainers.




