OptimizeFile/MergeCreateFile fail (and DecryptFile silently corrupts) an empty-user-password RC4-encrypted PDF when startxref needs reconstruction
Summary
For an RC4-encrypted PDF (V=1/R=2, 40-bit, empty user password -- i.e. openable without a password, just permission-restricted) whose startxref value is off by a small number of bytes (so the reader has to fall back to xref reconstruction), api.OptimizeFile and api.MergeCreateFile fail to process content that api.ReadContextFile/api.Validate handle fine, and api.DecryptFile silently produces a corrupted file -- it reports success but the output stream is not actually decrypted correctly.
This is not a hypothetical edge case: it's exactly the shape of file produced by at least one real-world PDF generator (iTextSharp, used by some e-invoicing platforms) when the ciphertext bytes of the /Info dictionary's encrypted strings happen to need PDF-literal-string escaping (open-paren, close-paren, backslash) and the generator computes xref offsets before accounting for that escaping -- producing a file whose startxref points a few bytes short of the real xref keyword.
Minimal reproducer
1.2KB, fully synthetic (built programmatically -- no real-world document data). Two files:
- A: a normal RC4-encrypted PDF, single page, one placeholder text stream. startxref correctly points at the xref keyword.
- B: byte-identical to A except startxref's value is decremented by 15 (so it now points into the middle of the preceding object instead of at xref).
$ qpdf --password= --check B.pdf
WARNING: B.pdf: file is damaged
WARNING: B.pdf (offset ...): xref not found
WARNING: B.pdf: Attempting to reconstruct cross-reference table
checking B.pdf
PDF Version: 1.4
R = 2
P = -60
User password =
Supplied password is user password
...
qpdf: operation succeeded with warningsqpdf reconstructs the xref table and decrypts/reads the file with no further issue -- confirming file B is a legitimate (if slightly malformed) encrypted PDF that a robust reader can recover from, not something structurally invalid.
Against this repo (tested on v0.15.0, and separately on a downstream fork pinned at v0.12.1 -- same result on both):
api.OptimizeFile(pathB, out, nil)
// -> optimize: optimize context: optimize resources: page 1: resource dict:
// page 1 content decode: stream filter[0] "FlateDecode": decode: zlib: invalid header
api.DecryptFile(pathB, out, nil)
// -> nil (reports success)
// but "out" is NOT a correctly decrypted file: its page-content stream fails
// to decompress (same "zlib: invalid header"), i.e. DecryptFile silently
// wrote back mis-decrypted bytes instead of returning an error.
api.MergeCreateFile([]string{pathB}, out, false, nil)
// -> merge destination B.pdf: read and validate: read context:
// encryption status: this file is encryptedFor comparison, file A (clean startxref) succeeds on all three calls with no error.
ReadContextFile and Validate on file B both succeed and correctly report Encrypted=true -- so the xref-reconstruction path itself finds the right encryption dictionary and derives what looks like the correct file key (in the real-world file I traced this from, I independently decrypted its content streams using my own from-spec implementation of the RC4 key derivation, and the derived key matched what this library computed internally). The corruption specifically shows up when the reconstructed xref table is later relied on again during optimize/decrypt-to-file, which is consistent with the same object's stream being decrypted twice somewhere in the xref-reconstruction-then-reprocess path (RC4 being self-inverse, decrypting an already-decrypted RC4 stream reproduces exactly this "corrupted, non-zlib-header" result) -- but I have not traced further inside the library to confirm the exact call site.
Why this matters beyond the synthetic case
I hit this while investigating a real customer-support bug: a real invoice PDF (RC4/R2, empty user password, iTextSharp-produced, startxref off by exactly 15 bytes for the reason above) was being merged into a multi-file export. The merge step used this exact detection pattern (probe-merge a single file to classify it as broken/encrypted) and got the exact same zlib: invalid header error, which was then surfaced to an end user as "this file is corrupted" -- which is wrong; the file opens fine everywhere including this library's own ReadContextFile/Validate, and the actual content is intact (verified via manual RC4 decryption). I don't want to attach the real file since it's a customer document. I confirmed the same failure reproduces with the real file's xref defect plus placeholder content plus the Google Noto font (OFL-licensed) standing in for the original's embedded CJK font, then reduced that down to the minimal synthetic repro above.
Environment
- Reproduced on github.com/pdfcpu/pdfcpu v0.15.0
- Also reproduced on a downstream fork pinned at v0.12.1 (same encrypt/decrypt/optimize code paths, no relevant patches in the fork)
- go1.26, darwin/arm64
Files (base64, decode with base64 -d)
A.pdf (clean startxref, all three calls succeed):
JVBERi0xLjQKJeLjz9MKMSAwIG9iago8PCAvVHlwZSAvQ2F0YWxvZyAvUGFnZXMgMiAwIFIgPj4K
ZW5kb2JqCjIgMCBvYmoKPDwgL1R5cGUgL1BhZ2VzIC9LaWRzIFszIDAgUl0gL0NvdW50IDEgPj4K
ZW5kb2JqCjMgMCBvYmoKPDwgL1R5cGUgL1BhZ2UgL1BhcmVudCAyIDAgUiAvTWVkaWFCb3ggWzAg
MCA1OTUgODQyXSAvQ29udGVudHMgNCAwIFIgL1Jlc291cmNlcyA8PCAvRm9udCA8PCAvRjEgNSAw
IFIgPj4gPj4gPj4KZW5kb2JqCjQgMCBvYmoKPDwgL0xlbmd0aCA2MyAvRmlsdGVyIC9GbGF0ZURl
Y29kZSA+PgpzdHJlYW0Km18SQc4WXrbl+bn78RV101hgBduotGf1z2O/QSf8hOKPLSWdaGK0mq0d
tBSHfAJ4MIngCBjXso93t0lAME6QCmVuZHN0cmVhbQplbmRvYmoKNSAwIG9iago8PCAvVHlwZSAv
Rm9udCAvU3VidHlwZSAvVHlwZTEgL0Jhc2VGb250IC9IZWx2ZXRpY2EgPj4KZW5kb2JqCjYgMCBv
YmoKPDwgL1Byb2R1Y2VyIDxhZWZhNTE5MjFiZmQ2ZjQzOGE2M2IxMjM0MDFmMjI0MWE3MDY5Mzhm
YmEyYjM4MTAwYz4gL0NyZWF0aW9uRGF0ZSA8OTliOTBkZDY0MWFlMmIxMmRiNzNmMjc0MDA1ZDdk
NTFlYjUzYzVjZGY4N2E2Yj4gPj4KZW5kb2JqCjcgMCBvYmoKPDwgL0xlbmd0aCAzOCAvRmlsdGVy
IC9GbGF0ZURlY29kZSAvTGVuZ3RoMSAyMDAwID4+CnN0cmVhbQpxUbKOSihSvS3yU8N2QaxfpfHW
dwLEOoUvvSoXvOVfPDELBywPRgplbmRzdHJlYW0KZW5kb2JqCjggMCBvYmoKPDwgL0ZpbHRlciAv
U3RhbmRhcmQgL1YgMSAvUiAyIC9PIDxhMGIxMzBlZDY5ZmU1NTkwMmRjMzNkOGM2MDIzNzljODU2
ZmU3ZWQxNWU5ZDhlMGJhY2Q3MzY3YTVmNjdiNmRjPiAvVSA8ODkyOTgwN2NlOWIzYjNiMmY3MzZh
NDNlMTFmNTdhOTIwYzgzNDRhZTAxN2E0Nzc2ZDI1NTQzYmZkN2RhYTFlMT4gL1AgLTYwID4+CmVu
ZG9iagp4cmVmCjAgOQowMDAwMDAwMDAwIDY1NTM1IGYgCjAwMDAwMDAwMTUgMDAwMDAgbiAKMDAw
MDAwMDA2NCAwMDAwMCBuIAowMDAwMDAwMTIxIDAwMDAwIG4gCjAwMDAwMDAyNDcgMDAwMDAgbiAK
MDAwMDAwMDM4MSAwMDAwMCBuIAowMDAwMDAwNDUxIDAwMDAwIG4gCjAwMDAwMDA1OTggMDAwMDAg
biAKMDAwMDAwMDcyMSAwMDAwMCBuIAp0cmFpbGVyCjw8IC9TaXplIDkgL1Jvb3QgMSAwIFIgL0lu
Zm8gNiAwIFIgL0VuY3J5cHQgOCAwIFIgL0lEIFs8MzAzMTMyMzMzNDM1MzYzNzM4Mzk2MTYyNjM2
NDY1NjY+IDwzMDMxMzIzMzM0MzUzNjM3MzgzOTYxNjI2MzY0NjU2Nj5dID4+CnN0YXJ0eHJlZgo5
MTcKJSVFT0YKB.pdf (startxref off by 15 bytes, reproduces the bug):
JVBERi0xLjQKJeLjz9MKMSAwIG9iago8PCAvVHlwZSAvQ2F0YWxvZyAvUGFnZXMgMiAwIFIgPj4K
ZW5kb2JqCjIgMCBvYmoKPDwgL1R5cGUgL1BhZ2VzIC9LaWRzIFszIDAgUl0gL0NvdW50IDEgPj4K
ZW5kb2JqCjMgMCBvYmoKPDwgL1R5cGUgL1BhZ2UgL1BhcmVudCAyIDAgUiAvTWVkaWFCb3ggWzAg
MCA1OTUgODQyXSAvQ29udGVudHMgNCAwIFIgL1Jlc291cmNlcyA8PCAvRm9udCA8PCAvRjEgNSAw
IFIgPj4gPj4gPj4KZW5kb2JqCjQgMCBvYmoKPDwgL0xlbmd0aCA2MyAvRmlsdGVyIC9GbGF0ZURl
Y29kZSA+PgpzdHJlYW0Km18SQc4WXrbl+bn78RV101hgBduotGf1z2O/QSf8hOKPLSWdaGK0mq0d
tBSHfAJ4MIngCBjXso93t0lAME6QCmVuZHN0cmVhbQplbmRvYmoKNSAwIG9iago8PCAvVHlwZSAv
Rm9udCAvU3VidHlwZSAvVHlwZTEgL0Jhc2VGb250IC9IZWx2ZXRpY2EgPj4KZW5kb2JqCjYgMCBv
YmoKPDwgL1Byb2R1Y2VyIDxhZWZhNTE5MjFiZmQ2ZjQzOGE2M2IxMjM0MDFmMjI0MWE3MDY5Mzhm
YmEyYjM4MTAwYz4gL0NyZWF0aW9uRGF0ZSA8OTliOTBkZDY0MWFlMmIxMmRiNzNmMjc0MDA1ZDdk
NTFlYjUzYzVjZGY4N2E2Yj4gPj4KZW5kb2JqCjcgMCBvYmoKPDwgL0xlbmd0aCAzOCAvRmlsdGVy
IC9GbGF0ZURlY29kZSAvTGVuZ3RoMSAyMDAwID4+CnN0cmVhbQpxUbKOSihSvS3yU8N2QaxfpfHW
dwLEOoUvvSoXvOVfPDELBywPRgplbmRzdHJlYW0KZW5kb2JqCjggMCBvYmoKPDwgL0ZpbHRlciAv
U3RhbmRhcmQgL1YgMSAvUiAyIC9PIDxhMGIxMzBlZDY5ZmU1NTkwMmRjMzNkOGM2MDIzNzljODU2
ZmU3ZWQxNWU5ZDhlMGJhY2Q3MzY3YTVmNjdiNmRjPiAvVSA8ODkyOTgwN2NlOWIzYjNiMmY3MzZh
NDNlMTFmNTdhOTIwYzgzNDRhZTAxN2E0Nzc2ZDI1NTQzYmZkN2RhYTFlMT4gL1AgLTYwID4+CmVu
ZG9iagp4cmVmCjAgOQowMDAwMDAwMDAwIDY1NTM1IGYgCjAwMDAwMDAwMTUgMDAwMDAgbiAKMDAw
MDAwMDA2NCAwMDAwMCBuIAowMDAwMDAwMTIxIDAwMDAwIG4gCjAwMDAwMDAyNDcgMDAwMDAgbiAK
MDAwMDAwMDM4MSAwMDAwMCBuIAowMDAwMDAwNDUxIDAwMDAwIG4gCjAwMDAwMDA1OTggMDAwMDAg
biAKMDAwMDAwMDcyMSAwMDAwMCBuIAp0cmFpbGVyCjw8IC9TaXplIDkgL1Jvb3QgMSAwIFIgL0lu
Zm8gNiAwIFIgL0VuY3J5cHQgOCAwIFIgL0lEIFs8MzAzMTMyMzMzNDM1MzYzNzM4Mzk2MTYyNjM2
NDY1NjY+IDwzMDMxMzIzMzM0MzUzNjM3MzgzOTYxNjI2MzY0NjU2Nj5dID4+CnN0YXJ0eHJlZgo5
MDIKJSVFT0YKHappy to share the larger repro (with the real xref defect + placeholder content + Noto font, ~4MB) or dig further into the exact double-decrypt call site if useful -- just let me know.
Source: pdfcpu/pdfcpu