Text mining is the process of extracting useful information from text data. Text data is
often referred to as unstructured data because it’s in raw form. To be analyzed, text data needs
to be converted to structured data so that the tools of descriptive statistics, data visualization
and data mining can be applied. Data mining with text data is more challenging than data
mining with traditional numerical data, as it requires more preprocessing to convert the text to
a format amenable for analysis.
The text-mining process converts unstructured text into numerical data and applies
quantitative techniques. Some of the ways to resolves the issues are tokenization, stemming
and frequency term document matrix. Tokenization is the process of dividing text into separate
terms, referred to as tokens where symbols and punctuations to be removed from the
document and all letters should be converted to lower case. Stemming is the process of
converting a word to its stem or root word, would drop the “ing” and “ed” and place only the
root word. A frequency term-document matrix is a matrix whose rows represent documents
and columns represent tokens, and the entries in the matrix are the frequency of occurrence of
each token in each document.
Camm, J.D., Cochran, J.J. Fry, M.J., Ohlmann, J.W., Anderson, D.R., Sweeney, D.J., &
Willians, T.A. (2019). Business Analytics (3rd ed.).
Text mining is the process of extracting useful information from text data. The text data are: words,
phrases, sentences, and paragraphs. Text data can be difficult because it is in its raw form and cannot be
stored in a traditional structured database. Because of this, it is often referred to as unstructured data.
Data mining with text is more challenging than data mining with traditional numerical data, because it
requires preprocessing to convert the text to a format amenable for analysis. Once the text data is
converted to numerical data, the analytical methods used for descriptive text mining are the same used
for numerical data mining. One way to resolve some issues resulted from text mining is to use
tokenization. Tokenization is the process of dividing text into separate terms, referred to as tokens. The
tokens are a unique identification symbol that retain all the essential information about the data.