tesseract

mirror of https://github.com/tesseract-ocr/tesseract.git synced 2024-12-01 07:59:05 +08:00

Author	SHA1	Message	Date
Stefan Weil	2e73c9d5ea	classify: Fix typos in comments and strings All of them were found by codespell. Signed-off-by: Stefan Weil <sw@weilnetz.de>	2016-02-05 10:54:26 +01:00
Ray Smith	b1d99dfe23	Added a backup adaptive classifier to take over from primary when it fills on a large document	2015-06-12 11:10:53 -07:00
Ray Smith	d74c625e52	Fixed blob division params to fix CJK training speed.	2015-06-12 10:59:26 -07:00
Ray Smith	5bb0d89291	Improved debug of class pruner	2015-05-13 17:07:11 -07:00
Ray Smith	84920b92b3	Font and classifier output structure cleanup. Font recognition was poor, due to forcing a 1st and 2nd choice at a character level, when the total score for the correct font is often correct at the word level, so allowed the propagation of a full set of fonts and scores to the word recognizer, which can now decide word level fonts using the scores instead of simple votes. Change precipitated a cleanup of output data structures for classifier results, eliminating ScoredClass and INT_RESULT_STRUCT, with a few extra elements going in UnicharRating, and using that wherever possible. That added the extra complexity of 1-rating due to a flip between 0 is good and 0 is bad for the internal classifier scores before they are converted to rating and certainty.	2015-05-12 17:24:34 -07:00
Ray Smith	53fc4456cc	Fixed issue 1252: Refactored LearnBlob and its call hierarchy to make it a member of Classify. Eliminated the flexfx scheme for calling global feature extractor functions through an array of function pointers. Deleted dead code I found as a by-product. This CL does not change BlobToTrainingSample or ExtractFeatures to be full members of Classify (the eventual goal) as that would make it even bigger, since there are a lot of callers to these functions. When ExtractFeatures and BlobToTrainingSample are members of Classify they will be able to access control parameters in Classify, which will greatly simplify developing variations to the feature extraction process.	2015-05-12 15:22:34 -07:00
theraysmith@gmail.com	1a487252f4	Fixed slow-down that was caused by upping MAX_NUM_CLASSES git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@1013 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2014-01-24 21:12:35 +00:00
theraysmith@gmail.com	7ec4fd7a56	Refactorerd control functions to enable parallel blob classification git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@904 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2013-11-08 20:30:56 +00:00
theraysmith@gmail.com	99edf4ccbd	Refactored classifier to make it easier to add new ones and generalized feature extractor to allow fx from grey git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@873 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2013-09-23 15:15:06 +00:00
theraysmith@gmail.com	5bc5e2a0b4	Added simultaneous multi-language capability, Added support for ShapeTable in classifier and training, Refactored class pruner, Added new uniform classifier API, Added new training error counter git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@650 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2012-02-02 02:57:42 +00:00
theraysmith	c86a0f6892	Various fixes, including memory leak in fixspace, font labels on output, removed some annoying debug output, fixes to initialization of parameters, general cleanup, and added Hindi git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@570 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2011-03-21 21:45:36 +00:00
theraysmith	df738bb9a4	Deleted lots of dead code, including PBLOB git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@559 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2011-03-18 21:53:11 +00:00
theraysmith	eba04e7c5b	Fixed debug display, training on fragments git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@533 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2010-11-30 01:00:17 +00:00
zdenop@gmail.com	4523ce9f7d	3.01 code from http://github.com/jimregan/tesseract-ocr with addaptions related to Linux and Windows (VC2008) compile process git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@526 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2010-11-23 18:34:14 +00:00
theraysmith	694d3f2c20	Changes to classify for 3.00 git-svn-id: https://tesseract-ocr.googlecode.com/svn/trunk@291 d0cd1f9f-072b-0410-8dd7-cf729c803f20	2009-07-11 02:17:36 +00:00

15 Commits