Corpus Concordancer
Turn your own documents into a small searchable corpus. Each hit shows its document, original line and column, and the surrounding left/right context. Neighboring tokens are counted by side, and document counts show where usage differs. This goes beyond a simple word-frequency list by preserving the actual occurrences and positions.
Key features
- Strict UTF-8 import of multiple .txt and .md files, plus typed documents
- Prevent double counting of equal names or normalized document bodies
- Exact boundary matching of one-to-six-token phrases over Unicode text
- KWIC rows with original UTF-16 offsets, line, column and surrounding text
- Left and right neighbor-token counts in a two-, four- or eight-token window
- Separate formula-safe UTF-8 CSV exports for hits, neighbors and documents
How to use
- Load the example, choose several UTF-8 text or Markdown files, or add a named passage yourself.
- Review the document list and token counts. A repeated name or identical normalized body is rejected.
- Enter a one-to-six-token query, choose a left/right context window and search.
- Read per-document counts, left/right neighbor counts and KWIC lines with source positions.
- Download the desired CSV reports separately.
Use cases
- Compare how a concept is used across several class notes
- Find expressions frequently used before or after a term in your writing
- Compare a phrase's context in several translation drafts
Frequently asked questions
Does this group Korean particles or English inflections?
No. There is no morphological analysis or stemming. A run of Unicode letters and digits is one token, so ‘강’ and ‘강은’ differ. NFC spelling and letter case are normalized for comparison.
How does phrase search treat punctuation?
Punctuation and spaces separate source tokens. Enter one to six query tokens separated by spaces. Consecutive source tokens match even if a comma or newline lies between them, and the KWIC row retains that original punctuation.
What exactly do neighboring-token counts mean?
For every hit, the selected number of tokens before and after the phrase is counted separately by side. The same word in multiple positions around one hit contributes multiple counts. This is not a statistical association score or a causal claim.
What happens if I upload the same document twice?
An equal name or a body equal after Unicode NFC and line-ending normalization is rejected. The tool does not infer whether different editions are equivalent or merge them automatically.
Are documents uploaded or automatically saved?
Documents and results stay in this browser tab's memory. This tool has no automatic upload or save. A local CSV is created only when you choose an export.
Privacy
Documents, queries and results are processed in this browser tab only. This tool has no automatic server submission or account storage. Review sensitive material before exporting or sharing reports.
Comments & questions