PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
-
Updated
Aug 1, 2026 - Python
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
⚡ From finding text to search and replace, from sorting to beautifying text and more 🎨
Diff Match Patch is a high-performance library in multiple languages that manipulates plain text.
Intuitive find & replace CLI (sed alternative)
fastNLP: A Modularized and Extensible NLP Framework. Currently still in incubation.
Python library for creating PEG parsers
Text Classification Algorithms: A Survey
Open-source humanize text toolkit. Documents 4 humanization methodologies with reference implementations, plus a production pipeline combining LLM rewriting with a cross-engine translation chain. Python, OpenAI-compatible API.
A fast and convenient fuzzy matcher library for rust
Persian NLP Toolkit
The most accurate natural language detection library for Go, suitable for short text and mixed-language text
A fast implementation of Aho-Corasick in Rust.
Program to convert lines of text into a tree structure.
Thai natural language processing in Python
High performance and CommonMark compliant HTML to Markdown converter. Maintained by the Kreuzberg team. Kreuzberg is a fast, polyglot document intelligence engine with a Rust core. It extracts structured data from 98+ document formats using streaming parsers and built-in OCR.
Text Normalization & Inverse Text Normalization
All-in-one text de-duplication
A simple Python module for parsing human names into their individual components
Add a description, image, and links to the text-processing topic page so that developers can more easily learn about it.
To associate your repository with the text-processing topic, visit your repo's landing page and select "manage topics."