pdfsizeopt
tesseract-ocr-for-php
pdfsizeopt | tesseract-ocr-for-php | |
---|---|---|
6 | 4 | |
715 | 2,792 | |
- | - | |
0.0 | 4.4 | |
4 months ago | 7 months ago | |
Python | PHP | |
GNU General Public License v3.0 only | MIT License |
Stars - the number of stars that a project has on GitHub. Growth - month over month growth in stars.
Activity is a relative number indicating how actively a project is being developed. Recent commits have higher weight than older ones.
For example, an activity of 9.0 indicates that a project is amongst the top 10% of the most actively developed projects that we are tracking.
pdfsizeopt
-
PostScript’s Sudden Death in Sonoma
> ...tools like pdftk have been able to losslessly compress them...
I have had good luck with pdfsizeopt.
https://github.com/pts/pdfsizeopt
-
PDF/A-3, PDF for Long-Term Preservation, Use of ISO 32000-1, with Embedded Files
The big restriction is that the classic Postscript typefaces are not available (no Times, Helvetica, or Zapf Dingbats), and the PDF file must bundle any fonts it uses.
The pdfsizeopt package will make any PDF smaller, and I think it deletes letters/characters from the included font that are not used.
https://github.com/pts/pdfsizeopt
-
PDF processing and analysis with open-source tools
This is missing the "pdfsizeopt" suite, that bundles statically compiled utilities to reduce size.
Static compilation means that it will run on most Linux platforms without extra required software.
I believe one aspect of it will remove characters from included fonts that are not used.
It really is quite impressive.
https://github.com/pts/pdfsizeopt
- Compressing bloated PDFs - pdfcompressor.com
-
Reducing the Size of Large PDFs
There is a general PDF shrinker, known as "pdfsizeopt" that is bundled with static builds of gs and a number of other utilities.
It cuts some of our PDFs to 10x smaller, mostly by removing unused fonts (but doubtless also some other magic).
The developer asks for donations for production use from those who can afford it.
https://github.com/pts/pdfsizeopt
Send donations to the author of pdfsizeopt:
https://flattr.com/submit/auto?user_id=pts&url=https://githu...
-
What are some good plataforms to build your own tabletop system?
Christian Mehrstram, creator of Whitehack, uses emacs to write LaTex documents, then converts those into PDFs using a script called pdfsizeopt.
tesseract-ocr-for-php
-
PDF processing and analysis with open-source tools
There’s even a library for php (https://github.com/thiagoalessio/tesseract-ocr-for-php). Haven’t used it. I did used python Pytesseract & works fairly well.
- Laravel OCR?
-
What are my options for extracting text from photos? I've already got ImageMagick installed, and assume there's a handful of PHP libraries for this task? Which are most performant and most likely to be maintained?
Depends on how consistent and legible the images are. If you've got, say, a scanned page with black-on-white text, it will work fairly well with PHPOCR (http://phpocr.sourceforge.net/) or https://github.com/thiagoalessio/tesseract-ocr-for-php.
-
Processing Identity Documents in Laravel
The next step is to use Tesseract in our PHP class, to do that we'll use this excellent package
What are some alternatives?
pdf-diff - A tool for visualizing differences between two pdf files.
react-native-tesseract-ocr - Tesseract OCR wrapper for React Native
chai - chai - Experience Zero Trust security with Chai! Convert and view documents as vivid images right in your browser. No mandatory downloads, no hassle—just pure, joyful security! 🌈
Laravel - Laravel is a web application framework with expressive, elegant syntax. We’ve already laid the foundation for your next big idea — freeing you to create without sweating the small things.
TCPDF - Official clone of PHP library to generate PDF documents and barcodes
Symfony - The Symfony PHP framework
author-tools - Author Tools
identitydocuments - A Laravel package for parsing and processing Identity Documents
WeasyPrint - The awesome document factory
tessdata - Trained models with fast variant of the "best" LSTM models + legacy models
CUPS - Apple CUPS Sources
tesseract-ocr - Tesseract Open Source OCR Engine (main repository)