Questions to test a Data Scientist on
Natural Language Processing With All
Correct & Verified Answers
Which of the following techniques can be used for the purpose of keyword normalization, the
process of converting a keyword into its base form?
Lemmatization
Levenshtein
Stemming
Soundex
A) 1 and 2
B) 2 and 4
C) 1 and 3
D) 1, 2 and 3
E) 2, 3 and 4
F) 1, 2, 3 and 4 Correct answer-(C)
Lemmatization and stemming are the techniques of keyword normalization, while Levenshtein and
Soundex are techniques of string matching.
N-grams are defined as the combination of N keywords together. How many bi-grams can be
generated from given sentence:
"Analytics Vidhya is a great source to learn data science"
A) 7
B) 8
C) 9
D) 10
E) 11 Correct answer-(C)
Bigrams: Analytics Vidhya, Vidhya is, is a, a great, great source, source to, To learn, learn data, data
science
How many trigrams phrases can be generated from the following sentence, after performing
following text cleaning steps:
Stopword Removal
Replacing punctuations by a single space
"#Analytics-vidhya is a great source to learn @data_science."
A) 3
B) 4
C) 5
D) 6
E) 7 Correct answer-(C)
, After performing stopword removal and punctuation replacement the text becomes: "Analytics
vidhya great source learn data science"
Trigrams - Analytics vidhya great, vidhya great source, great source learn, source learn data, learn
data science
Which of the following regular expression can be used to identify date(s) present in the text object:
"The next meetup on data science will be held on 2017-09-21, previously it happened on 31/03,
2016"
A) \d{4}-\d{2}-\d{2}
B) (19|20)\d{2}-(0[1-9]|1[0-2])-[0-2][1-9] C) (19|20)\d{2}-(0[1-9]|1[0-2])-([0-2][1-9]|3[0-1])
D) None of the above Correct answer-(D)
None if these expressions would be able to identify the dates in this text object.
You have collected a data of about 10,000 rows of tweet text and no other information. You want to
create a tweet classification model that categorizes each of the tweets in three buckets - positive,
negative and neutral.
5) Which of the following models can perform tweet classification with regards to context mentioned
above?
A) Naive Bayes
B) SVM
C) None of the above Correct answer-(C)
Since, you are given only the data of tweets and no other information, which means there is no
target variable present. One cannot train a supervised learning model, both svm and naive bayes are
supervised learning techniques.
You have created a document term matrix of the data, treating every tweet as one document. Which
of the following is correct, in regards to document term matrix?
Removal of stopwords from the data will affect the dimensionality of data
Normalization of words in the data will reduce the dimensionality of data
Converting all the words in lowercase will not affect the dimensionality of the data
A) Only 1
B) Only 2
C) Only 3
D) 1 and 2
E) 2 and 3
F) 1, 2 and 3 Correct answer-(D)
Choices A and B are correct because stopword removal will decrease the number of features in the
matrix, normalization of words will also reduce redundant features, and, converting all words to
lowercase will also decrease the dimensionality.
Which of the following features can be used for accuracy improvement of a classification model?
Natural Language Processing With All
Correct & Verified Answers
Which of the following techniques can be used for the purpose of keyword normalization, the
process of converting a keyword into its base form?
Lemmatization
Levenshtein
Stemming
Soundex
A) 1 and 2
B) 2 and 4
C) 1 and 3
D) 1, 2 and 3
E) 2, 3 and 4
F) 1, 2, 3 and 4 Correct answer-(C)
Lemmatization and stemming are the techniques of keyword normalization, while Levenshtein and
Soundex are techniques of string matching.
N-grams are defined as the combination of N keywords together. How many bi-grams can be
generated from given sentence:
"Analytics Vidhya is a great source to learn data science"
A) 7
B) 8
C) 9
D) 10
E) 11 Correct answer-(C)
Bigrams: Analytics Vidhya, Vidhya is, is a, a great, great source, source to, To learn, learn data, data
science
How many trigrams phrases can be generated from the following sentence, after performing
following text cleaning steps:
Stopword Removal
Replacing punctuations by a single space
"#Analytics-vidhya is a great source to learn @data_science."
A) 3
B) 4
C) 5
D) 6
E) 7 Correct answer-(C)
, After performing stopword removal and punctuation replacement the text becomes: "Analytics
vidhya great source learn data science"
Trigrams - Analytics vidhya great, vidhya great source, great source learn, source learn data, learn
data science
Which of the following regular expression can be used to identify date(s) present in the text object:
"The next meetup on data science will be held on 2017-09-21, previously it happened on 31/03,
2016"
A) \d{4}-\d{2}-\d{2}
B) (19|20)\d{2}-(0[1-9]|1[0-2])-[0-2][1-9] C) (19|20)\d{2}-(0[1-9]|1[0-2])-([0-2][1-9]|3[0-1])
D) None of the above Correct answer-(D)
None if these expressions would be able to identify the dates in this text object.
You have collected a data of about 10,000 rows of tweet text and no other information. You want to
create a tweet classification model that categorizes each of the tweets in three buckets - positive,
negative and neutral.
5) Which of the following models can perform tweet classification with regards to context mentioned
above?
A) Naive Bayes
B) SVM
C) None of the above Correct answer-(C)
Since, you are given only the data of tweets and no other information, which means there is no
target variable present. One cannot train a supervised learning model, both svm and naive bayes are
supervised learning techniques.
You have created a document term matrix of the data, treating every tweet as one document. Which
of the following is correct, in regards to document term matrix?
Removal of stopwords from the data will affect the dimensionality of data
Normalization of words in the data will reduce the dimensionality of data
Converting all the words in lowercase will not affect the dimensionality of the data
A) Only 1
B) Only 2
C) Only 3
D) 1 and 2
E) 2 and 3
F) 1, 2 and 3 Correct answer-(D)
Choices A and B are correct because stopword removal will decrease the number of features in the
matrix, normalization of words will also reduce redundant features, and, converting all words to
lowercase will also decrease the dimensionality.
Which of the following features can be used for accuracy improvement of a classification model?