You are viewing a plain text version of this content. The canonical link for it is here.
Posted to dev@tika.apache.org by "Tyler Palsulich (JIRA)" <ji...@apache.org> on 2015/03/15 04:25:39 UTC

[jira] [Resolved] (TIKA-1138) Empty body and empty title with some TXT documents

     [ https://issues.apache.org/jira/browse/TIKA-1138?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]

Tyler Palsulich resolved TIKA-1138.
-----------------------------------
    Resolution: Fixed

Closing as fixed, since the linked files now have their content extracted.

> Empty body and empty title with some TXT documents
> --------------------------------------------------
>
>                 Key: TIKA-1138
>                 URL: https://issues.apache.org/jira/browse/TIKA-1138
>             Project: Tika
>          Issue Type: Bug
>          Components: parser
>    Affects Versions: 1.3
>         Environment: Windows 7
>            Reporter: Koutsoulis Philippe
>
> *No error in logs*
> *+Extract from my "Structured Text":+*
> {noformat}
> <?xml version="1.0" encoding="UTF-8"?><html xmlns="http://www.w3.org/1999/xhtml">
> <head>
> ...
> <title/>
> </head>
> <body/></html>
> {noformat}
> *+Files to reproduce+*
> [http://top1000.anthologeek.net/participants.current.txt]
> [http://www.gregdonner.org/workbench/wb_31rev.txt]



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)