Lightning-fast file system indexer and search tool
Demo: sist2.shyy.io
Community URL: Discord
sist2 (Simple incremental search tool)
Warning: sist2 is in early development
* See format support
** See Archive files
*** See OCR
…
Navigate to http://localhost:8080/ to configure sist2-admin.
Choose search backend (See comparison):
docker run -d -p 9200:9200 -e "discovery.type=single-node" elasticsearch:7.17.9
Download the latest sist2 release.
Select the file corresponding to your platform. On Linux, mark the binary as executable with
chmod +x; on Windows, run sist2-x64-windows.exe from a terminal.
See usage guide for command line usage.
Example usage:
sist2 scan ~/Documents --output ./documents.sist2sist2 index --es-url http://localhost:9200 ./documents.sist2sist2 sqlite-index --search-index ./search.sist2 ./documents.sist2sist2 web ./documents.sist2sist2 web --search-index ./search.sist2 ./documents.sist2| File type | Library | Content | Thumbnail | Metadata |
|---|---|---|---|---|
| pdf,xps,fb2,epub | MuPDF | text+ocr | yes | author, title |
| cbz,cbr | libscan | - | yes | - |
audio/* |
ffmpeg | - | yes | ID3 tags |
video/* |
ffmpeg | - | yes | title, comment, artist |
image/* |
ffmpeg | ocr | yes | Common EXIF tags, GPS tags |
| raw, rw2, dng, cr2, crw, dcr, k25, kdc, mrw, pef, xf3, arw, sr2, srf, erf | LibRaw | no | yes | Common EXIF tags, GPS tags |
| ttf,ttc,cff,woff,fnt,otf | Freetype2 | - | yes, bmp |
Name & style |
text/plain |
libscan | yes | no | - |
| html, xml | libscan | yes | no | - |
| tar, zip, rar, 7z, ar ... | Libarchive | yes* | - | no |
| docx, xlsx, pptx | libscan | yes | if embedded | creator, modified_by, title |
| doc (MS Word 1-2003, incl. DOS and Macintosh) | libantiword2 | yes | no | author, title, modified_by |
| mobi, azw, azw3 | libmobi | yes | yes | author, title |
| wpd (WordPerfect) | libwpd | yes | no | planned |
| json, jsonl, ndjson | libscan | yes | - | - |
* See Archive files
sist2 will scan files stored into archive files (zip, tar, 7z...) as if they were directly in the file system. Recursive (archives inside archives) scan is also supported.
Limitations:
.gif, .mp4 w/ fragmented metadata etc.)
is limitted (see --mem-buffer option)You can enable OCR support for ebook (pdf,xps,fb2,epub) or image file types with the
--ocr-lang option in combination with --ocr-images and/or --ocr-ebooks.
Download the language data files with your package manager (apt install tesseract-ocr-eng) or
directly from Github.
The sist2app/sist2 image comes with common languages
(hin, jpn, eng, fra, rus, spa, chi_sim, deu, pol) pre-installed.
You can use the + separator to specify multiple languages. The language
name must be identical to the *.traineddata file installed on your system
(use chi_sim rather than chi-sim).
Examples:
sist2 scan --ocr-ebooks --ocr-lang jpn ~/Books/Manga/
sist2 scan --ocr-images --ocr-lang eng ~/Images/Screenshots/
sist2 scan --ocr-ebooks --ocr-images --ocr-lang eng+chi_sim ~/Chinese-Bilingual/
sist2 v3.0.7+ supports SQLite search backend. The SQLite search backend has fewer features and generally comparable query performance for medium-size indices, but it uses much less memory and is easier to set up.
| SQLite | Elasticsearch | |
|---|---|---|
| Requires separate search engine installation | ✓ | |
| Memory footprint | ~20MB | >500MB |
| Query syntax | fts5 | query_string |
| Fuzzy search | ✓ (spellfix) | ✓ (3-grams) |
| Media Types tree real-time updating | ✓ | |
| Manual tagging | ✓ | ✓ |
| User scripts | ✓ | ✓ |
| Media Type breakdown for search results | ✓ | |
| Embeddings search | ✓ O(n) | ✓ O(logn) |
| Per-chunk embeddings search | ✓ | ✓ (ES 8.11+) |
No open issues yet, or sync has not completed.