Semi-Automated Text Categorization Using Demonstration and Integration Based Term Set
Semi-Automated Text Categorization Using Demonstration and Integration Based Term Set
Manual Analysis of massive amounts of textual data requires incredible amount of processing time and effort in the interpretation of the text and organizing them in required format. In the current scenario, the major problem is with text or document categorization because of the high dimensionality of feature space. Now-a-days there are many methods available to deal with text feature selection. This paper aims at one such semi-automated text categorization feature selection methodology to deal with a enormous data using two phases of David Merrill’s First principles of instruction (FPI). It uses a pre-defined category group by providing them with the proper training set based on the demonstration and integration phase of FPI. The methodology involves the text tokenization, text categorization and text analysis.
#text #mining #characterization #feature #selection #and #phase