Class OCRParams

java.lang.Object
com.datalogics.PDFL.OCRParams

public class OCRParams extends Object
  • Constructor Details

    • OCRParams

      public OCRParams()
      Create an OCRParams structure with defaults.
  • Method Details

    • delete

      public void delete()
    • getTesseract4Engine

      public static String getTesseract4Engine()
      Specify the Tesseract v4 engine. Pass this to the engine property.
    • getResolution

      public double getResolution()
      Each image will use this resolution by default.
    • setResolution

      public void setResolution(double resolution)
      Each image will use this resolution by default.
    • getPageSegmentationMode

      public PageSegmentationMode getPageSegmentationMode()
      The page segmentation mode.
    • setPageSegmentationMode

      public void setPageSegmentationMode(PageSegmentationMode mode)
      The page segmentation mode.
    • getPerformance

      public Performance getPerformance()
      The desired engine performance.
    • setPerformance

      public void setPerformance(Performance performance)
      The desired engine performance.
    • getEngine

      public String getEngine()
      The engine in use.
    • setEngine

      public void setEngine(String engine)
      The engine in use.
    • getLanguages

      public List<LanguageSetting> getLanguages()
      The list of languages to use.
    • setLanguages

      public void setLanguages(List<LanguageSetting> languages)
      The list of languages to use.
    • getCandidateFontNames

      public List<String> getCandidateFontNames()
      The names of candidate fonts for placing text under an image. The default list should work well in most cases. If you're using text that isn't represented by Latin fonts, or by Chinese, Japanese, or Korean fonts, then retrieve this list, add the font that can represent that text, then set that new list on this object.

      Enough font names must be supplied to cover the expected languages/scripts in use.

      The code selects a font to represent each word. If a word code-switches between different scripts, for instance, if it contains non-Latin text and Arabic numerals, then make sure to supply the name of a font family that can handle both the text and the numerals.

      The quality of the results depends on the font choice. The list is searched in order until a font works for a particular word. To make the text fit better, it's recommended to list proportional fonts before fixed-width fonts. Decorative fonts with flourishes, like Zapf Chancery, deliver poor results. Generally, supply a font that would be used in block text, such as in a newspaper or work of literature, such as Times Roman, or a font already in the list, like MinionPro.

      If the PlaceTextUnder method in OCREngine can't identify a font that covers the whole text of a word, an exception will be thrown.

    • setCandidateFontNames

      public void setCandidateFontNames(List<String> candidateFontNames)
      The names of candidate fonts for placing text under an image. The default list should work well in most cases. If you're using text that isn't represented by Latin fonts, or by Chinese, Japanese, or Korean fonts, then retrieve this list, add the font that can represent that text, then set that new list on this object.

      Enough font names must be supplied to cover the expected languages/scripts in use.

      The code selects a font to represent each word. If a word code-switches between different scripts, for instance, if it contains non-Latin text and Arabic numerals, then make sure to supply the name of a font family that can handle both the text and the numerals.

      The quality of the results depends on the font choice. The list is searched in order until a font works for a particular word. To make the text fit better, it's recommended to list proportional fonts before fixed-width fonts. Decorative fonts with flourishes, like Zapf Chancery, deliver poor results. Generally, supply a font that would be used in block text, such as in a newspaper or work of literature, such as Times Roman, or a font already in the list, like MinionPro.

      If the PlaceTextUnder method in OCREngine can't identify a font that covers the whole text of a word, an exception will be thrown.

    • setEnableImagePreprocessing

      public void setEnableImagePreprocessing(boolean enable)
      Enable all image preprocessing, default enabled. Note that once the OCREngine is initialized, this setting is permanent.
    • getEnableImagePreprocessing

      public boolean getEnableImagePreprocessing()
      Get the image preprocessing enable state. True if any image preprocessing is on.
    • getConfigurationParameters

      public Map<String,String> getConfigurationParameters()
      Get the configuration parameters. Note: Reserved for internal use. Do not use unless directed to by Datalogics Support.
    • setConfigurationParameters

      public void setConfigurationParameters(Map<String,String> configurationParameters)
      Set the configuration parameters. Note: Reserved for internal use. Do not use unless directed to by Datalogics Support.