deepeval · Issues· 614 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #3150
[TS] Published 0.9.13 ships a stale dist/telemetry.js that require()s the removed @sentry/node, every CLI command crashes
Updated Sep 18, 2026 - #3303
LangChain CallbackHandler implements on_llm_error/on_tool_error/on_retriever_error but not on_chain_error, so an errored chain span is never exited and its trace is never posted
Updated Sep 16, 2026 - #3299
EvaluationDataset.save_as('csv'/'jsonl') splits context items that contain '|' when reloaded
Updated Sep 16, 2026 - #262
Add JudgeLM as a way to evaluate and compare historical test runs
enhancementhelp wantedUpdated Sep 15, 2026 - #3287
KimiModel.__init__ raises TypeError float(None) whenever the selected model has no registry pricing
Updated Sep 15, 2026 - #3283
Verbose judge verdicts are scored as pass in role_violation, prompt_alignment, conversation_completeness
Updated Sep 14, 2026 - #3133
OpenAIModel rejects unknown catalog ids; LocalModel base_url is the Chat Completions path
Updated Sep 14, 2026 - #3208
HumanEval benchmark validates n/k with a bare assert and accepts invalid values
Updated Sep 14, 2026 - #2748
Add retail support evaluation example using synthetic customer support cases
Updated Sep 14, 2026 - #3220
fix(metrics): whitespace-only actual_output/expected_output pass the empty-param guard
Updated Sep 13, 2026 - #3126
Question: stable cache format for reading per-test-case metric scores
Updated Sep 12, 2026 - #3263
OpenInference TOOL spans drop tool record on non-dict tool args and leave tools_called output empty (Strands parity)
Updated Sep 10, 2026 - #3261
AgentCore TOOL spans drop gen_ai.tool.call.result and crash on non-dict tool args (Strands parity)
Updated Sep 10, 2026 - #3186
fix(scorer): string scorers treat whitespace-only predictions as non-empty
Updated Sep 8, 2026 - #3257
[First-time contributor] Looking for beginner-friendly issues to contribute
Updated Sep 8, 2026