data-juicer · Issues· 61 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #1031
[Tracking] Reduce peak memory usage across Data-Juicer operators / 降低算子峰值内存
Updated Aug 6, 2026 - #1006
I have an idea to implement lineage capabilities for unstructured data—similar to the structured data lineage found in Apache Atlas or OpenLineage—within the current project
enhancementUpdated Jul 26, 2026 - #920
为什么datajuicer在Ray模式下,不支持groupby算子呢?
questionUpdated Feb 25, 2026 - #909
ETL过程的咨询
questionUpdated Feb 15, 2026 - #695
data-juicer 有计划支持流式或微批计算么
enhancementUpdated Feb 13, 2026 - #915
Multi-Branch Execution (DAG feature enhancment)
enhancementdj:coreUpdated Feb 13, 2026 - #667
如何实现两个算子的逻辑运算关系 How to implement the logical operation between two operators?
good first issuequestionUpdated Feb 13, 2026 - #732
什么时候data-juicer支持使用对象存储的数据集?
questiondj:distdj:coredj:datalakeUpdated Feb 11, 2026 - #734
How to reduce the memory overhead when processing millions of samples ?
questionUpdated Feb 11, 2026 - #792
能把所有用的模型放到一个地方?
good first issueUpdated Feb 11, 2026 - #670
Improve GPU utilization in ray mode
enhancementdj:distUpdated Jan 23, 2026 - #878
perplexity 算子,在计算中文数据集时,都特别大
questionUpdated Jan 16, 2026 - #879
去重算子CPU利用率非常低
questionUpdated Jan 14, 2026 - #813
数据增强是不是只能用于单字段的json,不能用于多字段的
Updated Jan 13, 2026 - #797
Is there any issue with the return types of the read, read_json, and read_webdataset functions in the RayDataSet class?
Updated Dec 24, 2025