You are viewing a plain text version of this content. The canonical link for it is here.
Posted to dev@lucene.apache.org by "Ryadh Dahimene (JIRA)" <ji...@apache.org> on 2018/09/11 16:06:00 UTC

[jira] [Commented] (LUCENE-8462) New Arabic snowball stemmer

    [ https://issues.apache.org/jira/browse/LUCENE-8462?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16610847#comment-16610847 ] 

Ryadh Dahimene commented on LUCENE-8462:
----------------------------------------

Hi [~rcmuir], what do you think of the new dataset for the test vocabulary? Should I look for other alternatives?

[~thetaphi], following your comment I've started looking at the ant task `{{ant patch-snowball`}}  and yes it needs to be updated for the new snowball generated java classes. I will try to work on that and submit a change on a separate issue.

> New Arabic snowball stemmer
> ---------------------------
>
>                 Key: LUCENE-8462
>                 URL: https://issues.apache.org/jira/browse/LUCENE-8462
>             Project: Lucene - Core
>          Issue Type: Improvement
>            Reporter: Ryadh Dahimene
>            Priority: Trivial
>              Labels: Arabic, snowball, stemmer
>          Time Spent: 10m
>  Remaining Estimate: 0h
>
> Added a new Arabic snowball stemmer based on [https://github.com/snowballstem/snowball/blob/master/algorithms/arabic.sbl]
> As well an Arabic test dataset in `TestSnowballVocabData.zip` from the -snowball-data- generated from the input file available here -[https://github.com/snowballstem/snowball-data/tree/master/arabic]-
> [https://github.com/ibnmalik/golden-corpus-arabic/blob/develop/core/words.txt]
>  
> Link to the corresponding Github PR:
>  [https://github.com/apache/lucene-solr/pull/439]
>  Edited: updated the corpus link
>  



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

---------------------------------------------------------------------
To unsubscribe, e-mail: dev-unsubscribe@lucene.apache.org
For additional commands, e-mail: dev-help@lucene.apache.org