Home | Determine the OS and version of PHP »
OCR for PDF in Ubuntu
By admin | July 8, 2008
To get OCR in Ubuntu, you need to use the Open-Source Tesseract OCR engine. However, it can only perform OCR on the TIFF format. In order to allow it to do PDF, we also need the Evince PDF reader to allow us to export a page to TIFF to feed Tesseract with.
Install tesseract in Ubuntu:
$ sudo apt-get install tesseract-ocr
Now, get the TIFF from Evince; Right click the page and click “Save image as…” and export the file. Then, to get a good OCR we need to convert it to monochrome.
$ tesseract foo.tif bar
Unable to load unicharset file /usr/share/tesseract-ocr/tessdata/eng.unicharset
If the above bug happens, fix it by:
$ sudo apt-get install tesseract-ocr-eng
This is caused by a nasty bug in Ubuntu.
And now finally, do:
$ tesseract foo.tif bar# and without the .txt extension, or you will end up with bar.txt.txt
to complete it!
If you found this article helpful or interesting, please help Compdigitec spread the word. Don’t forget to subscribe to Compdigitec Labs for more useful and interesting articles!
Topics: Linux | 104 Comments »

August 12th, 2026 at 19:56
ก่อนหน้านี้เข้าใจว่าเลือกรุ่นไหนก็ใช้งานคล้ายกัน แต่จริงๆ มีรายละเอียดต่างกันพอสมควรเลยครับ
Also visit my webpage – ติดตั้งกลอนดิจิทัล,
http://Yu856.com,
August 12th, 2026 at 20:17
เรื่องมาตรฐานการติดตั้งเป็นจุดที่หลายคนอาจมองข้าม แต่จริงๆ สำคัญมากครับ
Also visitt mmy bllg post; บริการจาก U Smart Lock
August 12th, 2026 at 20:51
รายละเอียดเกี่ยวกับสีและจำนวนแผ่นที่พิมพ์ได้ช่วยให้เปรียบเทียบสินค้าได้ง่ายครับ
Visit my homepage – สั่งซื้อตลับหมึกแท้
August 12th, 2026 at 21:08
ขอบคุณสำหรับข้อมูลครับ การตรวจรหัสตลับหมึกให้ตรงกับรุ่นเครื่องก่อนสั่งซื้อเป็นเรื่องสำคัญมาก
Feel free to visit my web page; ร้านขายตลับหมึกแท้