A legacy .xls cell that says N/A or NULL is dropped from the index
Self Checks
- I have read the Contributing Guide and Language Policy.
- This is only for bug report, if you would like to ask a question, please head to Discussions.
- I have searched for existing issues search for existing issues, including closed ones.
- I confirm that I am using English to submit this report, otherwise it will be closed.
- 【中文用户 & Non English User】请使用英语提交,否则会被关闭 :)
- Please do not modify this template :) and fill in all the required fields.
Dify version
main (9c6c48b50b)
Cloud or Self Hosted
Self Hosted (Source)
Steps to reproduce
Upload a legacy .xls to a knowledge base in which some cells literally say N/A, NA, n/a, NULL, None, NaN or nan — "not applicable" in a hand-written sheet, NULL in a database export.
api/core/rag/extractor/excel_extractor.py reads the .xls branch through pandas:
df = excel_file.parse(sheet_name=sheet_name)
pandas treats those seven strings as missing values by default, so the cell arrives as NaN and is skipped together with its column name.
A three-row sheet, written with xlwt and read exactly as the extractor does:
Code | Status | Units indexed today
A1 | N/A | 12 -> "Code":"A1";"Units":"12"
A2 | NULL | 7 "Code":"A2";"Units":"7"
A3 | ok | 3 "Code":"A3";"Status":"ok";"Units":"3"
✔️ Expected Behavior
A cell keeps the text it holds, as the .xlsx branch of the same extractor already does (it walks openpyxl cells and gets the string):
"Code":"A1";"Status":"N/A";"Units":"12"
"Code":"A2";"Status":"NULL";"Units":"7"
"Code":"A3";"Status":"ok";"Units":"3"
❌ Actual Behavior
The Status key is absent from the first two rows, not merely empty. A question such as "what is the status of A1?" finds no evidence in the index, and the chunk looks complete, so nothing signals the loss. dtype=object does not change this; only keep_default_na=False does.
A fix with tests is open in #42371.
Source: langgenius/dify