从大型巴西基金监管 PDF 中提取结构化数据的挑战(230 千字符,22 种实体类型)
作者: reichaves创建于 2026年2月11日更新于 2026年8月20日
Hi! I'm Reinaldo Chaves, a data journalist from Brazil. I'm building an open-source tool to extract structured information from Brazilian investment fund regulation PDFs using LangExtract + Gemini. These are standardized documents (CVM Resolution 175/2022) with ~100 pages / 230K characters each, containing 22 entity types (fund name, CNPJ, administrator, manager, fees, duration, risk factors, liquidation events, legal forum, etc.). **Project repo:** https://GitHub.com/reichaves/langextract-fundos The goal is to enable investigative journalists to systematically analyze thousands of fund regulations for transparency and accountability purposes.
内容来源: google/langextract