Baike.dev
All toolsAI codingTrendingOpen sourceNewsSubmit
Log in
Back to tool

opencompass · Issues· 404 open

Open on GitHub

Locally synced open issues (discussions stay on GitHub)

  • #2647

    `gsm8k_postprocess` breaks numeric answers at thousands separators (`\boxed{$9{,}500}` → `500`, `\boxed{8,000}` → `000`)

    Updated Sep 18, 2026
  • #2635

    TopkRetriever crashes on transformers 5.x: tokenizer.encode_plus was removed

    Updated Sep 8, 2026
  • #2631

    Grading integrity: model code rewrites ground truth in three benchmarks; CRUXEval splices prediction into assertion; unparseable verdicts exit denominator; LCEvaluator eval() on raw prediction (follow-up to #2535)

    Updated Sep 3, 2026
  • #2627

    [Bug] ToxicEvaluator: expected_max_toxicity uses built-in max() on a NaN-containing array, so the metric depends on sample order

    Updated Sep 1, 2026
  • #2615

    OpenAI-compat: OpenAISDK openai_api_base /v1, not the requests OpenAI full-path default

    Updated Aug 28, 2026
  • #2604

    Feature request: EvalPort import/export for CustomDataset (adapter already built and tested)

    Updated Aug 21, 2026
  • #2508

    [Feature] DIK-structured Turkish syntactic dependency annotation dataset for structured linguistic evaluation

    Updated Aug 8, 2026
  • #2574

    CompassAcademic: 4 of 7 objective boards have an unseparated #1, two are exact ties broken by display order

    Updated Aug 3, 2026
  • #2573

    [Bug] Qwen3.5 cannot run when using HuggingFacewithChatTemplate

    Updated Jul 31, 2026
  • #2516

    MedFailBench integration proposal for clinical AI safety evaluation

    Updated Jul 18, 2026
  • #1678

    [Bug] TypeError: Expected a datasets.Dataset or a datasets.DatasetDict object, but got {'train': <modelscope.msdatasets.ms_dataset.MsDataset object at xxx>, 'test': <modelscope.msdatasets.ms_dataset.MsDataset object at xxx>}

    Updated Jul 17, 2026
  • #1875

    [Feature] 目前是否有适配Codeforces、SWE Verified、Aider-Polyglot这些在R1中出现的数据集的计划呢?

    Updated Jul 16, 2026
  • #2159

    [Feature] 请问主观评测脚本支持用本地模型作为judge模型吗?

    Updated Jul 15, 2026
  • #1856

    [Bug] 在对DeepSeek-R1-Distill-Qwen-1.5B模型评测livecodebench数据集时,lcb_test_output为什么为0呢?

    Updated Jul 15, 2026
  • #2386

    [Feature] Add another OpenAISDK model using prompt completion API

    Updated Jul 7, 2026