You are viewing a plain text version of this content. The canonical link for it is here.
Posted to dev@nutch.apache.org by "Sebastian Nagel (JIRA)" <ji...@apache.org> on 2012/11/01 09:57:12 UTC

[jira] [Comment Edited] (NUTCH-1483) Can't crawl filesystem with protocol-file plugin

    [ https://issues.apache.org/jira/browse/NUTCH-1483?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13488558#comment-13488558 ] 

Sebastian Nagel edited comment on NUTCH-1483 at 11/1/12 8:55 AM:
-----------------------------------------------------------------

Thanks!
Issue with un-reversing URLs pulled out to NUTCH-1484 since it's more critical (no work-around).
Fixing the URL normalizers (and filters, see last comment) will take more time. 
Btw., {{file://localhost/Documents/}} is the only legal for according to [RFC 1738|http://tools.ietf.org/html/rfc1738] (1994) while {{file:///Documents/}} is allowed by [RFC 3986|http://tools.ietf.org/html/rfc3986] (2005):
{quote}
the "file" URI scheme is defined so that no authority, an empty host, and "localhost" all mean the end-user's machine
{quote}
Maybe we could also make protocol-file more lazy.
                
      was (Author: wastl-nagel):
    Thanks!
Issue with un-reversing URLs pulled out to NUTCH-1484 since it's more critical (no work-around).
Fixing the URL normalizers (and filters, see last comment) will take more time. 
Btw., {{file://localhost/Documents/}} is the only legal for according to [RFC 1738|http://tools.ietf.org/html/rfc1738] (1994) while {{file:///Documents/}} is allowed by [RFC 3986|http://tools.ietf.org/html/rfc3986]:
{quote}
the "file" URI scheme is defined so that no authority, an empty host, and "localhost" all mean the end-user's machine
{quote}
Maybe we could also make protocol-file more lazy.
                  
> Can't crawl filesystem with protocol-file plugin
> ------------------------------------------------
>
>                 Key: NUTCH-1483
>                 URL: https://issues.apache.org/jira/browse/NUTCH-1483
>             Project: Nutch
>          Issue Type: Bug
>          Components: protocol
>    Affects Versions: 1.6, 2.1
>         Environment: OpenSUSE 12.1, OpenJDK 1.6.0, HBase 0.90.4
>            Reporter: Rogério Pereira Araújo
>         Attachments: NUTCH-1483.patch
>
>
> I tried to follow the same steps described in this wiki page:
> http://wiki.apache.org/nutch/IntranetDocumentSearch
> I made all required changes on regex-urlfilter.txt and added the following entry in my seed file:
> file:///home/rogerio/Documents/
> The permissions are ok, I'm running nutch with the same user as folder owner, so nutch has all the required permissions, unfortunately I'm getting the following error:
> org.apache.nutch.protocol.file.FileError: File Error: 404
>         at org.apache.nutch.protocol.file.File.getProtocolOutput(File.java:105)
>         at org.apache.nutch.fetcher.FetcherReducer$FetcherThread.run(FetcherReducer.java:514)
> fetch of file://home/rogerio/Documents/ failed with: org.apache.nutch.protocol.file.FileError: File Error: 404
> Why the logs are showing file://home/rogerio/Documents/ instead of file:///home/rogerio/Documents/ ???
> Note: The regex-urlfilter entry only works as expected if I add the entry 
> +^file://home/rogerio/Documents/ instead of +^file:///home/rogerio/Documents/ as wiki says.

--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira