Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
Back to tool

lm-evaluation-harness · Issues· 994 open

Open on GitHub

Locally synced open issues (discussions stay on GitHub)

  • #4187

    `mmlu_*_generative` `get_response` compares the whole first line to the gold letter — models that prefix or explain their answer score exactly 0.000

    Updated Sep 18, 2026
  • #4185

    Python API guide passes unsupported num_fewshot to evaluate()

    Updated Sep 18, 2026
  • #4183

    Metadata n-shot override is ignored in collected results

    Updated Sep 18, 2026
  • #1473

    Issue with `bigbench_gender_inclusive_sentences_german_multiple_choice`

    Updated Sep 18, 2026
  • #4178

    An apostrophe in a CLI value silently swallows every following key=value argument

    Updated Sep 17, 2026
  • #1396

    Acc vs acc_norm

    asking questionsUpdated Sep 17, 2026
  • #4175

    longbench2_multi group omits longbench2_legal_multi, so it disagrees with the longbench2 tag roll-up

    Updated Sep 16, 2026
  • #4171

    Shared tasks are double-counted in nested group aggregation

    Updated Sep 16, 2026
  • #4169

    Finite retry exhaustion returns None instead of raising

    Updated Sep 16, 2026
  • #1501

    Add New Lambada Translations

    good first issueUpdated Sep 15, 2026
  • #4159

    Group rows keep a 0.0 stderr when every subtask is entirely on the boundary, so the row with the most evidence carries no bound

    Updated Sep 14, 2026
  • #4158

    GGUF backend's echo=true handling doesn't work against current llama-server — every loglikelihood request fails

    Updated Sep 14, 2026
  • #4153

    MELA group omits Icelandic and lists Arabic twice

    Updated Sep 14, 2026
  • #4070

    At 0 (or 100%) accuracy the reported stderr is exactly 0.0 at any N, so null results claim infinite precision; a boundary-aware interval or flag would fix it

    Updated Sep 13, 2026
  • #4147

    CLI key-value parser preserves whitespace in argument names

    Updated Sep 12, 2026