You are viewing a plain text version of this content. The canonical link for it is here.

Posted to dev@lucene.apache.org by "Steven Parkes (JIRA)" <ji...@apache.org> on 2007/07/31 23:07:52 UTC

[jira] Created: (LUCENE-971) Create enwiki indexable data as line-per-article rather than file-per-article

Create enwiki indexable data as line-per-article rather than file-per-article
-----------------------------------------------------------------------------

                 Key: LUCENE-971
                 URL: https://issues.apache.org/jira/browse/LUCENE-971
             Project: Lucene - Java
          Issue Type: Improvement
            Reporter: Steven Parkes


Create a line per article rather than a file. Consume with indexLineFile task.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


---------------------------------------------------------------------
To unsubscribe, e-mail: java-dev-unsubscribe@lucene.apache.org
For additional commands, e-mail: java-dev-help@lucene.apache.org

[jira] Resolved: (LUCENE-971) Create enwiki indexable data as line-per-article rather than file-per-article

Posted by "Michael McCandless (JIRA)" <ji...@apache.org>.

     [ https://issues.apache.org/jira/browse/LUCENE-971?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]

Michael McCandless resolved LUCENE-971.
---------------------------------------

    Resolution: Fixed


OK, committed with these small changes:

  * Replaced conf/wikipedia.alg -> conf/extractWikipedia.alg in the
    comment in that file.

  * Moved doc.maker line up under the "# Where to get documents from:"
    comment

  * In build.xml, removed the extract-enwiki target so that "ant
    enwiki" does the right thing.

Thanks Steve!


> Create enwiki indexable data as line-per-article rather than file-per-article
> -----------------------------------------------------------------------------
>
>                 Key: LUCENE-971
>                 URL: https://issues.apache.org/jira/browse/LUCENE-971
>             Project: Lucene - Java
>          Issue Type: Improvement
>            Reporter: Steven Parkes
>            Assignee: Steven Parkes
>         Attachments: LUCENE-971.patch.txt, LUCENE-971.patch.txt, LUCENE-971.patch.txt
>
>
> Create a line per article rather than a file. Consume with indexLineFile task.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


---------------------------------------------------------------------
To unsubscribe, e-mail: java-dev-unsubscribe@lucene.apache.org
For additional commands, e-mail: java-dev-help@lucene.apache.org

[jira] Commented: (LUCENE-971) Create enwiki indexable data as line-per-article rather than file-per-article

Posted by "Doron Cohen (JIRA)" <ji...@apache.org>.

    [ https://issues.apache.org/jira/browse/LUCENE-971?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel#action_12517048 ] 

Doron Cohen commented on LUCENE-971:
------------------------------------

Mmm... an additional advantage of this is not needing to extract 
the entire enwiki collection in order to index it - setting the 
repetition count to 100 for AddDocTask in alternative 1 or for 
WriteLineDocTask in alternative 2 would  mean that only 100 
docs from the huge file are extracted.

> Create enwiki indexable data as line-per-article rather than file-per-article
> -----------------------------------------------------------------------------
>
>                 Key: LUCENE-971
>                 URL: https://issues.apache.org/jira/browse/LUCENE-971
>             Project: Lucene - Java
>          Issue Type: Improvement
>            Reporter: Steven Parkes
>         Attachments: LUCENE-971.patch.txt
>
>
> Create a line per article rather than a file. Consume with indexLineFile task.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


---------------------------------------------------------------------
To unsubscribe, e-mail: java-dev-unsubscribe@lucene.apache.org
For additional commands, e-mail: java-dev-help@lucene.apache.org