Exam (elaborations)

Natural Language Processing (NLP), Top Exam Questions and answers, verified.

Rating

Sold

Pages

Grade

A+

Uploaded on

13-06-2023

Written in

2022/2023

Natural Language Processing (NLP), Top Exam Questions and answers, verified. Artificial Intelligence - -A computer performing tasks that a human can do NLP Sentiment analysis is a form of... - -classification NLP topic modeling is a form of... - -Dimensionality reduction Tokenization - -Splitting raw text into small, indivisible units for processing. Units can be words, sentences, n-grams (n-word combos), other characters defined by regex Stop words - -Words that have very little semantic value Stemming and Lemmatization - -Cut word down to base form Stemming- uses rough heuristics to reduce words to base Lemmatization- uses vocabulary and morphological analysis (makes run, runs, running, and ran all the same) Named Entity Recognition - -Identifies and tags named entities in text (people, places, organizations, phone numbers, emails, etc) aka entity extraction Compound term extraction - -extracting and tagging compound words or phrases in text Levenshtein distance - -Minimum number of operations to get from one word to another. One way of quantifying word similarity Levenshtein operations - -Deletions (delete a character) Insertions (insert a character) Mutation (change a character) Corpus - -Collection of texts Bag of words model - -- Simplified representation of text, where each document is recognized as a bag of its words - Grammar and word order are disregarded, but multiplicity is kept Cosine similarity - -Way to quantify the similarity between documents 1. Put each document in vector format 2. Find the cosine of the angle between the documents Term frequency-inverse document frequency - -(term frequency) * (inverse document frequency) Term frequency - -Term count/total terms Inverse document frequency - -- Considers how common a word is among all the documents - Rare words get additional weight Which classification models suffer from curse of dimensionality? - -KNN, SVM, linear models (linear/logisitic regression), decision trees Distance-based models feature selection - -Removes features that aren't helpful (might not be predictive of target and may not have a lot of variation) Art (try fitting with some features and changing it and comparing, regularization, feature importance scores) Feature extraction - -Uses information from all features, but creates artificial new features that are composites (uses information from all features, may put more weight on certain features) SVD - -Singular value decomposition, type of matrix decomposition (just multiplication) easy to compute & doesn't require square matrix matrices are our data, rows are observations, columns are features High level, what does SVD do? - -Generalization of eigendecomposition for rectangular matrices PCA - -Principal Components Analysis unsupervised technique care about the direction of maximal variation b/c that represents the differences in our observations and that helps us in our clustering/classification tasks What is PCA used for? - -Dropping the components that explain the least variance (uses SVD behind the scenes). Dimensionality reduction Document-term matrix - -rows = documents

Show more Read less

Institution

Course

Whoops! We can’t load your doc right now. Try again or contact support.

Report Copyright Violation

Written for

Course: NLP

All documents for this subject (76)

Document information

Uploaded on: June 13, 2023
Number of pages: 7
Written in: 2022/2023
Type: Exam (elaborations)
Contains: Questions & answers

Subjects

top exam questi
natural language processing nlp

Content preview

Natural Language Processing (NLP), Top
Exam Questions and answers, verified.

Artificial Intelligence - ✔✔-A computer performing tasks that a human can do

NLP Sentiment analysis is a form of... - ✔✔-classification

NLP topic modeling is a form of... - ✔✔-Dimensionality reduction

Tokenization - ✔✔-Splitting raw text into small, indivisible units for processing. Units can be words,
sentences, n-grams (n-word combos), other characters defined by regex

Stop words - ✔✔-Words that have very little semantic value

Stemming and Lemmatization - ✔✔-Cut word down to base form

Stemming- uses rough heuristics to reduce words to base

Lemmatization- uses vocabulary and morphological analysis (makes run, runs, running, and ran all the
same)

Named Entity Recognition - ✔✔-Identifies and tags named entities in text (people, places, organizations,
phone numbers, emails, etc)

aka entity extraction

Compound term extraction - ✔✔-extracting and tagging compound words or phrases in text

, Levenshtein distance - ✔✔-Minimum number of operations to get from one word to another. One way
of quantifying word similarity

Levenshtein operations - ✔✔-Deletions (delete a character)

Insertions (insert a character)

Mutation (change a character)

Corpus - ✔✔-Collection of texts

Bag of words model - ✔✔-- Simplified representation of text, where each document is recognized as a
bag of its words

- Grammar and word order are disregarded, but multiplicity is kept

Cosine similarity - ✔✔-Way to quantify the similarity between documents

1. Put each document in vector format

2. Find the cosine of the angle between the documents

Term frequency-inverse document frequency - ✔✔-(term frequency) * (inverse document frequency)

Term frequency - ✔✔-Term count/total terms

Inverse document frequency - ✔✔-- Considers how common a word is among all the documents

- Rare words get additional weight

Which classification models suffer from curse of dimensionality? - ✔✔-KNN, SVM, linear models
(linear/logisitic regression), decision trees

Distance-based models

$8.99

Get access to the full document:

100% satisfaction guarantee

Immediately available after payment

Both online and in PDF

No strings attached

Get to know the seller

PassPoint02

4.1

(39)

Get to know the seller

PassPoint02 Chamberlain School Of Nursing

View profile

Sold

173

Member since

3 year

Number of followers

105

Documents

4552

Last sold

4 weeks ago

4.1

39 reviews

Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Frequently asked questions

What do I get when I buy this document?

You get a PDF, available immediately after your purchase. The purchased document is accessible anytime, anywhere and indefinitely through your profile.

Satisfaction guarantee: how does it work?

Our satisfaction guarantee ensures that you always find a study document that suits you well. You fill out a form, and our customer service team takes care of the rest.

Who am I buying these notes from?

Stuvia is a marketplace, so you are not buying this document from us, but from seller PassPoint02. Stuvia facilitates payment to the seller.

Will I be stuck with a subscription?

No, you only buy these notes for $8.99. You're not tied to anything after your purchase.

Can Stuvia be trusted?

4.6 stars on Google & Trustpilot (+1000 reviews) 48586 documents were sold in the last 30 days Founded in 2010, the go-to place to buy study notes for 16 years now

Natural Language Processing (NLP), Top Exam Questions and answers, verified.

Written for

Document information

Subjects

Content preview

Get to know the seller

Recently viewed by you

Why students choose Stuvia

Created by fellow students, verified by reviews

Didn't get what you expected? Choose another document

Pay as you like, start learning right away

Frequently asked questions

What do I get when I buy this document?

Satisfaction guarantee: how does it work?

Who am I buying these notes from?

Will I be stuck with a subscription?

Can Stuvia be trusted?