Class LinkAnalysisScoringFilter

java.lang.Object
org.apache.nutch.scoring.AbstractScoringFilter
org.apache.nutch.scoring.link.LinkAnalysisScoringFilter
All Implemented Interfaces:
Configurable, Pluggable, ScoringFilter

public class LinkAnalysisScoringFilter extends AbstractScoringFilter
  • Constructor Details

    • LinkAnalysisScoringFilter

      public LinkAnalysisScoringFilter()
  • Method Details

    • setConf

      public void setConf(Configuration conf)
      Specified by:
      setConf in interface Configurable
      Overrides:
      setConf in class AbstractScoringFilter
    • generatorSortValue

      public float generatorSortValue(Text url, CrawlDatum datum, float initSort) throws ScoringFilterException
      Description copied from interface: ScoringFilter
      This method prepares a sort value for the purpose of sorting and selecting top N scoring pages during fetchlist generation.
      Specified by:
      generatorSortValue in interface ScoringFilter
      Overrides:
      generatorSortValue in class AbstractScoringFilter
      Parameters:
      url - url of the page
      datum - page's datum, should not be modified
      initSort - initial sort value, or a value from previous filters in chain
      Returns:
      a sort value for use in sorting and selecting the top N scoring pages during fetchlist generation
      Throws:
      ScoringFilterException - if there is a fatal error preparing the sort value
    • indexerScore

      public float indexerScore(Text url, NutchDocument doc, CrawlDatum dbDatum, CrawlDatum fetchDatum, Parse parse, Inlinks inlinks, float initScore) throws ScoringFilterException
      Description copied from interface: ScoringFilter
      This method calculates a indexed document score/boost.
      Specified by:
      indexerScore in interface ScoringFilter
      Overrides:
      indexerScore in class AbstractScoringFilter
      Parameters:
      url - url of the page
      doc - indexed document. NOTE: this already contains all information collected by indexing filters. Implementations may modify this instance, in order to store/remove some information.
      dbDatum - current page from CrawlDb. NOTE:
      • changes made to this instance are not persisted
      • may be null if indexing is done without CrawlDb or if the segment is generated not from the CrawlDb (via FreeGenerator).
      fetchDatum - datum from FetcherOutput (containing among others the fetching status)
      parse - parsing result. NOTE: changes made to this instance are not persisted.
      inlinks - current inlinks from LinkDb. NOTE: changes made to this instance are not persisted.
      initScore - initial boost value for the indexed document.
      Returns:
      boost value for the indexed document. This value is passed as an argument to the next scoring filter in chain. NOTE: implementations may also express other scoring strategies by modifying the indexed document directly.
      Throws:
      ScoringFilterException - if there is a fatal error whilst calculating the indexed document score/boost
    • initialScore

      public void initialScore(Text url, CrawlDatum datum) throws ScoringFilterException
      Description copied from interface: ScoringFilter
      Set an initial score for newly discovered pages. Note: newly discovered pages have at least one inlink with its score contribution, so filter implementations may choose to set initial score to zero (unknown value), and then the inlink score contribution will set the "real" value of the new page.
      Specified by:
      initialScore in interface ScoringFilter
      Overrides:
      initialScore in class AbstractScoringFilter
      Parameters:
      url - url of the page
      datum - new datum. Filters will modify it in-place.
      Throws:
      ScoringFilterException - if there is a fatal error setting an initial score for newly discovered pages
    • passScoreAfterParsing

      public void passScoreAfterParsing(Text url, Content content, Parse parse) throws ScoringFilterException
      Description copied from interface: ScoringFilter
      Currently a part of score distribution is performed using only data coming from the parsing process. We need this method in order to ensure the presence of score data in these steps.
      Specified by:
      passScoreAfterParsing in interface ScoringFilter
      Overrides:
      passScoreAfterParsing in class AbstractScoringFilter
      Parameters:
      url - page url
      content - original content. NOTE: modifications to this value are not persisted.
      parse - target instance to copy the score information to. Implementations may modify this in-place, primarily by setting some metadata properties.
      Throws:
      ScoringFilterException - if there is a fatal error processing score data in subsequent steps after parsing
    • passScoreBeforeParsing

      public void passScoreBeforeParsing(Text url, CrawlDatum datum, Content content) throws ScoringFilterException
      Description copied from interface: ScoringFilter
      This method takes all relevant score information from the current datum (coming from a generated fetchlist) and stores it into Content metadata. This is needed in order to pass this value(s) to the mechanism that distributes it to outlinked pages.
      Specified by:
      passScoreBeforeParsing in interface ScoringFilter
      Overrides:
      passScoreBeforeParsing in class AbstractScoringFilter
      Parameters:
      url - url of the page
      datum - source datum. NOTE: modifications to this value are not persisted.
      content - instance of content. Implementations may modify this in-place, primarily by setting some metadata properties.
      Throws:
      ScoringFilterException - if there is a fatal error injecting score information from the current datum into Content metadata