You are viewing a plain text version of this content. The canonical link for it is here.
Posted to dev@nutch.apache.org by "Sebastian Nagel (JIRA)" <ji...@apache.org> on 2018/06/12 19:32:00 UTC
[jira] [Updated] (NUTCH-2032) Plugin to index the raw content of a
readable document.
[ https://issues.apache.org/jira/browse/NUTCH-2032?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]
Sebastian Nagel updated NUTCH-2032:
-----------------------------------
Fix Version/s: (was: 1.15)
> Plugin to index the raw content of a readable document.
> --------------------------------------------------------
>
> Key: NUTCH-2032
> URL: https://issues.apache.org/jira/browse/NUTCH-2032
> Project: Nutch
> Issue Type: New Feature
> Components: indexer, parser
> Affects Versions: 1.10
> Reporter: Luis Lopez
> Assignee: Lewis John McGibbney
> Priority: Major
> Labels: content, index, index-rawcontent, parser, raw
>
> This is related to https://issues.apache.org/jira/browse/NUTCH-1785 and
> https://issues.apache.org/jira/browse/NUTCH-1458
> We created a couple plugins to index the raw content of readable documents. If we include these plugins in the plugin chain we'll index the raw content of a readable document, i.e. XML, HTML, CSV, TXT etc. The index-rawcontent plugin is not designed to index binary files, however having the full content of an HTML/XML or a CSV document is really critical for some of us.
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)