• Wrong document? Swap it for free
  • Written by students who passed
  • Immediately available after payment
  • Read online or as PDF
Sell
Where do you study
Your language
Document preview thumbnail
Preview 2 out of 8 pages
Exam (elaborations)

Questions to test a Data Scientist on Natural Language Processing With All Correct & Verified Answers

Document preview thumbnail
Preview 2 out of 8 pages

Questions to test a Data Scientist on Natural Language Processing With All Correct & Verified Answers

Content preview

Questions to test a Data Scientist on
Natural Language Processing With All
Correct & Verified Answers
Which of the following techniques can be used for the purpose of keyword normalization, the
process of converting a keyword into its base form?

Lemmatization
Levenshtein
Stemming
Soundex
A) 1 and 2
B) 2 and 4
C) 1 and 3
D) 1, 2 and 3
E) 2, 3 and 4
F) 1, 2, 3 and 4 Correct answer-(C)

Lemmatization and stemming are the techniques of keyword normalization, while Levenshtein and
Soundex are techniques of string matching.

N-grams are defined as the combination of N keywords together. How many bi-grams can be
generated from given sentence:

"Analytics Vidhya is a great source to learn data science"

A) 7
B) 8
C) 9
D) 10
E) 11 Correct answer-(C)

Bigrams: Analytics Vidhya, Vidhya is, is a, a great, great source, source to, To learn, learn data, data
science

How many trigrams phrases can be generated from the following sentence, after performing
following text cleaning steps:

Stopword Removal
Replacing punctuations by a single space
"#Analytics-vidhya is a great source to learn @data_science."

A) 3
B) 4
C) 5
D) 6
E) 7 Correct answer-(C)

, After performing stopword removal and punctuation replacement the text becomes: "Analytics
vidhya great source learn data science"

Trigrams - Analytics vidhya great, vidhya great source, great source learn, source learn data, learn
data science

Which of the following regular expression can be used to identify date(s) present in the text object:

"The next meetup on data science will be held on 2017-09-21, previously it happened on 31/03,
2016"

A) \d{4}-\d{2}-\d{2}
B) (19|20)\d{2}-(0[1-9]|1[0-2])-[0-2][1-9] C) (19|20)\d{2}-(0[1-9]|1[0-2])-([0-2][1-9]|3[0-1])
D) None of the above Correct answer-(D)

None if these expressions would be able to identify the dates in this text object.

You have collected a data of about 10,000 rows of tweet text and no other information. You want to
create a tweet classification model that categorizes each of the tweets in three buckets - positive,
negative and neutral.

5) Which of the following models can perform tweet classification with regards to context mentioned
above?

A) Naive Bayes
B) SVM
C) None of the above Correct answer-(C)

Since, you are given only the data of tweets and no other information, which means there is no
target variable present. One cannot train a supervised learning model, both svm and naive bayes are
supervised learning techniques.

You have created a document term matrix of the data, treating every tweet as one document. Which
of the following is correct, in regards to document term matrix?

Removal of stopwords from the data will affect the dimensionality of data
Normalization of words in the data will reduce the dimensionality of data
Converting all the words in lowercase will not affect the dimensionality of the data
A) Only 1
B) Only 2
C) Only 3
D) 1 and 2
E) 2 and 3
F) 1, 2 and 3 Correct answer-(D)

Choices A and B are correct because stopword removal will decrease the number of features in the
matrix, normalization of words will also reduce redundant features, and, converting all words to
lowercase will also decrease the dimensionality.

Which of the following features can be used for accuracy improvement of a classification model?

Document information

Uploaded on
May 6, 2025
Number of pages
8
Written in
2024/2025
Type
Exam (elaborations)
Contains
Questions & answers
$16.99

Wrong document? Swap it for free Within 14 days of purchase and before downloading, you can choose a different document. You can simply spend the amount again.
Written by students who passed
Immediately available after payment
Read online or as PDF

Seller avatar
Reputation scores are based on the amount of documents a seller has sold for a fee and the reviews they have received for those documents. There are three levels: Bronze, Silver and Gold. The better the reputation, the more your can rely on the quality of the sellers work.
Studyclub
3.6
(14)
Sold
70
Followers
1
Items
13387
Last sold
1 day ago



Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

Student with book image

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Working on your references?

Create accurate citations in APA, MLA and Harvard with our free citation generator.

Working on your references?

Frequently asked questions