Sites Web : Convert Scanned PDF Documents to Text with Google OCR - La veille technologique de Pyrat.net

Publié le lundi 27 juin 2005

⇒ https://veille.pyrat.net/

La création de site web implique de se tenir à jour et au courant des évolutions techniques.

Vous retrouverez ici les pages que nous avons visitées et dont nous avons estimées digne d’intérêt.

Il n’y a d’ailleurs pas que de la technique…

Convert Scanned PDF Documents to Text with Google OCR

Octobre 2008, Par jpyrat

Now if you have bunch of scanned PDF files on your hard drive and no OCR software, here’s what you can do to convert them into recognizable text.

Create a folder in your website (say abc.com/pdf) and upload all the PDF images to that folder. Now create a public web page that links to all the PDF files. Wait for the Google bots to spider your stuff.

Once done, type the query « site:abc.com/pdf filetype:pdf » to see the PDF documents as HTML.

→ Lire la suite sur le site d’origine…


Revenir en haut