lm-evaluation-harness · Issues· 994 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #4187
`mmlu_*_generative` `get_response` compares the whole first line to the gold letter — models that prefix or explain their answer score exactly 0.000
Updated Sep 18, 2026 - #4185
Python API guide passes unsupported num_fewshot to evaluate()
Updated Sep 18, 2026 - #4183
Metadata n-shot override is ignored in collected results
Updated Sep 18, 2026 - #1473
Issue with `bigbench_gender_inclusive_sentences_german_multiple_choice`
Updated Sep 18, 2026 - #4178
An apostrophe in a CLI value silently swallows every following key=value argument
Updated Sep 17, 2026 - #1396
Acc vs acc_norm
asking questionsUpdated Sep 17, 2026 - #4175
longbench2_multi group omits longbench2_legal_multi, so it disagrees with the longbench2 tag roll-up
Updated Sep 16, 2026 - #4171
Shared tasks are double-counted in nested group aggregation
Updated Sep 16, 2026 - #4169
Finite retry exhaustion returns None instead of raising
Updated Sep 16, 2026 - #1501
Add New Lambada Translations
good first issueUpdated Sep 15, 2026 - #4159
Group rows keep a 0.0 stderr when every subtask is entirely on the boundary, so the row with the most evidence carries no bound
Updated Sep 14, 2026 - #4158
GGUF backend's echo=true handling doesn't work against current llama-server — every loglikelihood request fails
Updated Sep 14, 2026 - #4153
MELA group omits Icelandic and lists Arabic twice
Updated Sep 14, 2026 - #4070
At 0 (or 100%) accuracy the reported stderr is exactly 0.0 at any N, so null results claim infinite precision; a boundary-aware interval or flag would fix it
Updated Sep 13, 2026 - #4147
CLI key-value parser preserves whitespace in argument names
Updated Sep 12, 2026