Document Search Index
Build an in-memory index of selectable text from multiple local documents by PDF page or text line. Search an exact phrase or all terms, read a context snippet and open the corresponding source PDF page. Unlike a single-text replacement or a predefined book-term index, this explores an ad hoc file collection.
Key features
- In-memory inverted index for up to eight documents with PDF pages and TXT/MD lines
- Korean and English exact phrase or all-term search with context snippets
- Open a PDF result at its original local page or inspect a text line number
- Removing a document deletes its posting lists and resets stale results
- Explicit file, page, character and result bounds with a JSON search manifest
How to use
- Choose multiple local PDF, TXT or MD files, or load the example.
- Review indexed documents and page or line counts.
- Enter terms, choose phrase or all-word matching, and search.
- Inspect snippets and positions; open an original PDF page when needed.
- Remove an unneeded document, export JSON or clear the whole index.
Use cases
- Find a contract phrase across several meeting PDFs
- Search Korean notes and a PDF report in one session
- Remove one document and compare refreshed results
- Export document, location and excerpt evidence as JSON
Frequently asked questions
Does it search image-only scanned PDFs?
No. It reads selectable text already embedded in the PDF. Run OCR to make a scanned PDF searchable first, then check OCR accuracy against the source.
Where are my documents and index stored?
Only in this tab's memory. Closing or refreshing the tab removes the index. This tool does not upload your documents.
Can removed documents still appear in results?
No. Removal clears that document's page or line segments and inverted-index postings, and invalidates the previous result list.
Can extracted PDF text order differ from the visible page?
Yes. Multi-column layouts, tables and vertical writing can produce a different extraction order. Use the page and snippet as clues and check the original page.
What formats and limits apply?
Up to eight PDF, UTF-8 TXT or MD files. Each may be 10 MiB, all files together 24 MiB; PDFs are limited to 50 pages each and 120 total. Extracted text is bounded too. Encrypted and image-only PDFs cannot be searched.
Privacy
Documents and postings stay in this tab's browser memory; they are not uploaded or saved in browser storage. PDF code and worker load from this site's static assets. Exported JSON contains document names and excerpts; review it before sharing.
Comments & questions