You are viewing a plain text version of this content. The canonical link for it is here.
Posted to dev@pdfbox.apache.org by "Bernard (JIRA)" <ji...@apache.org> on 2010/06/01 01:31:37 UTC
[jira] Commented: (PDFBOX-586) Text Extraction Regression ?
[ https://issues.apache.org/jira/browse/PDFBOX-586?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12873825#action_12873825 ]
Bernard commented on PDFBOX-586:
--------------------------------
Hi,
The previous version which was OK was 0.7.3.
The newest 1.1.0 :
- I took the source
- compile it using Eclipse & created a .jar : and 60% of my PDF are not extracted.
Andreas : how do you get those results ? From a .jar provided by PDFBox team ? or did you recompile the sources ? Did you commented something in the sources (encrypted PFD lib. use for instance ?) Are you in Windows or Linux or Mac ?
I've spent a lot of time evaluating this 1.1.0, and stopped adding new features to my app. I may go back to PDFBox 0.7.3 (and all it's slowness problems) ....
> Text Extraction Regression ?
> ----------------------------
>
> Key: PDFBOX-586
> URL: https://issues.apache.org/jira/browse/PDFBOX-586
> Project: PDFBox
> Issue Type: Bug
> Components: Text extraction
> Affects Versions: 1.1.0
> Environment: Windows XP + Eclipse + PDFBox sources
> Reporter: Bernard
> Attachments: ASEB-Camping_Car_ou_Bateau.pdf, Eval.pdf, internals.pdf, PDFBOX586-ASEB-Camping_Car_ou_Bateau.txt, PDFBOX586-Eval.txt, PDFBOX586-internals.txt
>
>
> Hi,
> I have noticed that I can extract text some PDF files in PDFBox 0.7.4 but for the same file, the same page, PDFBox 1.1.0 doesn't retreive any text, or the extraction is worst.
> Am I the only only one who think there is a regression in text extraction ?
> My code is like this :
> PDDocument document = PDDocument.load("/sdcard/internals.pdf");
> int numberOfPages = document.getNumberOfPages();
> resources = this.getResources();
>
> android.util.Log.d(TEST_PDFBOX, "readerPDF() resources : "+resources); // ANDROID code here to get file
> resourceGlyphList = R.raw.glyphlist;
> InputStream rawResource = resources.openRawResource(R.raw.pdftextstripper); // PDFBOX property file
> android.util.Log.d(TEST_PDFBOX, "readerPDF() rawResource : "+rawResource);
> Properties properties = new Properties();
> properties.load(rawResource);
>
> PDFTextStripper stripper = new PDFTextStripper(properties );
>
> stripper.setStartPage(pageNumber ); // 1 or any other page
> stripper.setEndPage(pageNumber ); // same page as above
> String s = "Page : "+pageNumber+"<br><br>"+stripper.getText(document);
> android.util.Log.d(TEST_PDFBOX, "readerPDF() stripper extract pages text : "+s);
> Maybe I should use page.getContents().getStream() or stripper.getTextForRegion( "class1" ) or stripper.writeText(doc, outputStream)
> I want the text as a String, not as a newly created file....
--
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.