Class URLNormalizers
plugin.include property).
There is one global scope defined by default, which consists of all active normalizers. The order in which these normalizers are executed may be defined in "urlnormalizer.order" property, which lists space-separated implementation classes (if this property is missing normalizers will be run in random order). If there are more normalizers activated than explicitly named on this list, the remaining ones will be run in random order after the ones specified on the list are executed.
You can define a set of contexts (or scopes) in which normalizers may be called. Each scope can have its own list of normalizers (defined in "urlnormalizer.scope.<scope_name>" property) and its own order (defined in "urlnormalizer.order.<scope_name>" property). If any of these properties are missing, default settings are used for the global scope.
In case no normalizers are required for any given scope, a
org.apache.nutch.net.urlnormalizer.pass.PassURLNormalizer should
be used.
Each normalizer may further select among many configurations, depending on the scope in which it is called, because the scope name is passed as a parameter to each normalizer. You can also use the same normalizer for many scopes.
Several scopes have been defined, and various Nutch tools will attempt using scope-specific normalizers first (and fall back to default config if scope-specific configuration is missing).
Normalizers may be run several times, to ensure that modifications introduced
by normalizers at the end of the list can be further reduced by normalizers
executed at the beginning. By default this loop is executed just once - if
you want to ensure that all possible combinations have been applied you may
want to run this loop up to the number of activated normalizers. This loop
count can be configured through urlnormalizer.loop.count property.
As soon as the url is unchanged the loop will stop and return the result.
- Author:
- Andrzej Bialecki
-
Field Summary
FieldsModifier and TypeFieldDescriptionstatic final StringScope used when updating the CrawlDb with new URLs.static final StringDefault scope.static final StringScope used byFetcherwhen processing redirect URLs.static final StringScope used byGenerator.static final StringScope used when indexing URLs.static final StringScope used byInjector.static final StringScope used when updating the LinkDb with new URLs.static final StringScope used when constructing newOutlinkinstances.static final StringScope used byURLPartitioner. -
Constructor Summary
Constructors -
Method Summary
-
Field Details
-
SCOPE_DEFAULT
Default scope. If no scope properties are defined then the configuration for this scope will be used.- See Also:
-
SCOPE_PARTITION
Scope used byURLPartitioner.- See Also:
-
SCOPE_GENERATE_HOST_COUNT
Scope used byGenerator.- See Also:
-
SCOPE_FETCHER
Scope used byFetcherwhen processing redirect URLs.- See Also:
-
SCOPE_CRAWLDB
Scope used when updating the CrawlDb with new URLs.- See Also:
-
SCOPE_LINKDB
Scope used when updating the LinkDb with new URLs.- See Also:
-
SCOPE_INJECT
Scope used byInjector.- See Also:
-
SCOPE_OUTLINK
Scope used when constructing newOutlinkinstances.- See Also:
-
SCOPE_INDEXER
Scope used when indexing URLs.- See Also:
-
-
Constructor Details
-
URLNormalizers
-
-
Method Details
-
normalize
Normalize- Parameters:
urlString- The URL string to normalize.scope- The given scope.- Returns:
- A normalized String, using the given
scope - Throws:
MalformedURLException- If the given URL string is malformed.
-