You are viewing a plain text version of this content. The canonical link for it is here.
Posted to dev@tika.apache.org by "Nick Burch (JIRA)" <ji...@apache.org> on 2013/11/21 15:53:35 UTC

[jira] [Commented] (TIKA-1199) Tika extracts weird signs instead of text

    [ https://issues.apache.org/jira/browse/TIKA-1199?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13828985#comment-13828985 ] 

Nick Burch commented on TIKA-1199:
----------------------------------

I've just tried this file with the Apache PDFBox tool "org.apache.pdfbox.PDFBox" using the option "ExtractText". I get the same binary stuff out too. However, the linux tool pdftotext is able to extract the text out just fine, so it doesn't appear to be a corrupt file.

This will therefore need reporting upstream to the Apache PDFBox project

> Tika extracts weird signs instead of text
> -----------------------------------------
>
>                 Key: TIKA-1199
>                 URL: https://issues.apache.org/jira/browse/TIKA-1199
>             Project: Tika
>          Issue Type: Bug
>          Components: parser
>    Affects Versions: 1.4
>         Environment: MacOSX, Linux
>            Reporter: Marc Teutelink
>         Attachments: gaat fout.pdf, plain_text_tika_output_from_gaat_fout_pdf.txt, structured_text_tika_output_from_gaat_fout_pdf.xml
>
>
> Tika extracts complete bogus text from the attached document. I have attached the .PDF in question and also added the plain and structured text output from Tika.



--
This message was sent by Atlassian JIRA
(v6.1#6144)