evals · Issues· 340 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #1833
[Architecture Proposal] Hardware-Gated Containment Framework (Genesis Protocol V5.0) for Recursive Self-Improvement
Updated Sep 13, 2026 - #1832
Proposal: harness-level failure attribution in eval reports (model vs harness vs eval definition)
Updated Sep 13, 2026 - #1825
Proposal: detect tool-trajectory regressions before deployment
Updated Sep 13, 2026 - #1829
Proposal: RES actor-indexing behavioral eval (20 samples, no custom code)
Updated Sep 9, 2026 - #1826
False-positive "distillation" ban on $200/mo Pro subscriber — automated appeal rejected by bot, requesting human review from Trust & Safety
bugUpdated Sep 6, 2026 - #1827
Proposal: deterministic eval for agent action-boundary violations
Updated Sep 5, 2026 - #1824
Proposal: Shared Sparse Latent Neural Bus for Local ↔ Cloud Models
Updated Sep 5, 2026 - #1819
HttpRecorder silently discards logged eval events; HTTP error responses never trigger the fallback at any failure rate
Updated Aug 29, 2026 - #1818
API retry helper retries forever (no max_tries/max_time) and the error-detection branch in its callers is dead code
Updated Aug 29, 2026 - #1712
Model-graded classify: unparseable judge output silently becomes minimum score
Updated Aug 29, 2026 - #1817
[Bug] Context/Token Explosion in Geometric Script Generation (CadQuery/CalculiX)
bugUpdated Aug 26, 2026 - #1811
LangChain completion wrappers drop runtime kwargs
Updated Aug 21, 2026 - #1810
Steganography does not flag missing required protocol fields
Updated Aug 21, 2026 - #1809
SchellingPoint mutates process-wide random state
Updated Aug 21, 2026 - #1805
MakeMeSay get_content fails on its documented dict response type
Updated Aug 21, 2026