You are viewing a plain text version of this content. The canonical link for it is here.
Posted to dev@nutch.apache.org by "ASF GitHub Bot (Jira)" <ji...@apache.org> on 2019/09/27 13:04:00 UTC

[jira] [Commented] (NUTCH-2457) Embedded documents likely not correctly parsed by Tika

    [ https://issues.apache.org/jira/browse/NUTCH-2457?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16939422#comment-16939422 ] 

ASF GitHub Bot commented on NUTCH-2457:
---------------------------------------

sebastian-nagel commented on pull request #474: NUTCH-2457 Embedded documents likely not correctly parsed by Tika
URL: https://github.com/apache/nutch/pull/474
 
 
   - add unit test for embedded documents
 
----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
users@infra.apache.org


> Embedded documents likely not correctly parsed by Tika
> ------------------------------------------------------
>
>                 Key: NUTCH-2457
>                 URL: https://issues.apache.org/jira/browse/NUTCH-2457
>             Project: Nutch
>          Issue Type: Bug
>    Affects Versions: 1.14
>            Reporter: Tim Allison
>            Priority: Major
>             Fix For: 1.16
>
>
> While working on TIKA-2490, I think I found that Nutch's current method of requesting a mime-specific parser for each file will fail to parse embedded files, e.g. https://github.com/apache/tika/blob/master/tika-server/src/test/resources/test_recursive_embedded.docx
> The fix should be straightforward, and I'll submit a PR once I can get Nutch up and running in my dev environment. 



--
This message was sent by Atlassian Jira
(v8.3.4#803005)