gorilla · Issues· 280 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #1367
[BFCL] Several multi_turn_base items have ground truth that cannot be produced from the environment
Updated Sep 9, 2026 - #1338
[BFCL] Eager import of qwen_agent forces audio deps on Claude-only users
Updated Sep 4, 2026 - #1362
Irrelevance detection credits a tool call that failed to parse as a correct abstention
Updated Aug 26, 2026 - #1360
Would an EvalPort-format export of BFCL be welcome in EvalPort's benchmarks/?
Updated Aug 22, 2026 - #1333
[BFCL] DeepSeek v4?
Updated Aug 19, 2026 - #1352
[BFCL] Clarify default and ablation behavior of synthetic request failures in BFCL V4 Web Search
Updated Aug 12, 2026 - #1355
[BFCL] Strict regex-based evaluator gives false negatives even on semantically accurate model responses
Updated Aug 7, 2026 - #1354
[BFCL] simple_python_363: release-commit answer key expects `find_closest`, but the item only presents `restaurant_search.find_closest`
Updated Aug 5, 2026 - #1353
BFCL: deliberately skipped categories render as 0.00% in score CSVs and dilute Overall Acc
Updated Aug 5, 2026 - #1349
EvalScope integration: one-command evaluation support for BFCL-v3 and BFCL-v4
Updated Jul 24, 2026 - #1172
[BFCL] SerpApi is so expensive!is there any cheaper search api to alternative? And how many calls are needed to complete the web searches for these 200 pieces of web search data?
Updated Jun 26, 2026 - #1343
[BFCL] Question regarding Custom Handler implementation for Qwen2.5-Instruct fine-tuned on Rlla-4k (BFCL-v4)
Updated Jun 10, 2026 - #1339
[BFCL] OpenAICompletionsHandler: stale keep-alive connections cause spurious 400 errors
Updated May 31, 2026 - #1331
[BFCL] Function-calling parser returns empty response when tool_calls=[]
Updated Apr 27, 2026 - #1301
[BFCL] The most unreadable codes I have ever seen
Updated Apr 17, 2026