Summary

Summary of paper End-to-end Object Detection with Transformers

Rating

Sold

Pages

Uploaded on

05-07-2024

Written in

2023/2024

This is a summary of the paper End-to-end Object Detection with Transformers for the course Seminar of Computer Vision by Deep Learning in TU Delft

Institution

Course

Whoops! We can’t load your doc right now. Try again or contact support.

Report Copyright Violation

Written for

Institution: Technische Universiteit Delft (TU Delft)
Study: Computer Science And Engineering
Course: CS4245

All documents for this subject (10)

Document information

Uploaded on: July 5, 2024
Number of pages: 7
Written in: 2023/2024
Type: Summary

Subjects

transformers
deep learning
computer vision
tu delft
summary

Content preview

End-to-end Object Detection
with Transformers
Abstract
This approach removes the need for many hand-designed components like
non-maximum suppression procedure or anchor generation DETR doesn’t need
that! The main ingredients of the new framework, called DEtection TRansformer
or DETR are a set-based global loss that forces unique predictions via bipartite
matching and a transformer encoder-decoder architecture.
Prior methods: Current object detection pipelines include hand-crafted
components like spatial anchor generation and non-max suppression (NMS).
Each of these components is tuned specifically for a given task. For example,
NMS is threshold-based and requires an IOU (intersection over union) and
confidence threshold tuning to be able to effectively discard the overlapping
bounding boxes.

Introduction
Modern detectors address this set prediction task in an indirect way, by
defining surrogate regression and classification problems on a large set of
proposals, anchors or window centers. Their performances are significantly
influenced by postprocessing steps to collapse near-duplicate predictions.

DETR directly predicts (in parallel) the final set of detections by combining a common CNN
with a transformer architecture. During training, bipartite matching uniquely assigns
predictions with ground truth boxes.

Our DEtection TRansformer predicts all objects at once, and is trained end-to-
end with a set loss function which performs bipartite matching between
predicted and ground truth objects.

End-to-end Object Detection with Transformers 1

, Compared to most previous work on direct set prediction, the main features of
DETR are the conjunction of the bipartite matching loss and transformers with
(non-autoregressive) parallel decoding.

Related Work
Set Prediction
A task where a model predicts multiple elements whose ordering is not relevant
for correctness. (Essentially predicting multiple objects in an image).
The way this is solved now however is by introducing relationship or pre
defined knowledge into the model. For instance, the predicted bounding boxes
should not overlap significantly and should cover all detected objects.
Avoiding Near-Duplicates: In object classification sometimes there are the
same bounding boxes for the same predicition, this is solved by using NMS
however set prediction is set to resolve that.

Transformers and Parallel Decoding
Transformers introduced self-attention layers, which, similarly to Non-Local
Neural Networks, scan through each element of a sequence and update it by
aggregating information from the whole sequence.

Object Detection
Set-based loss: Several object detectors used the bipartite matching loss.
Recurrent detectors: Closest to our approach are end-to-end set predictions
for object detection and instance segmentation. Similarly to us, they use
bipartite-matching losses with encoder-decoder architectures based on CNN
activation to directly produce a set of bounding boxes. These approaches,
however, were only evaluated on small datasets and not against modern
baselines. In particular, they are based on autoregressive models (more
precisely RNNs), so they do not leverage the recent transformers with parallel
decoding.

The DETR model
Object Detection set prediction loss

End-to-end Object Detection with Transformers 2

$8.67

Get access to the full document:

100% satisfaction guarantee

Immediately available after payment

Both online and in PDF

No strings attached

Get to know the seller

guillemribes

Also available in package deal

Get to know the seller

guillemribes Technische Universiteit Delft

View profile

Sold

Member since

1 year

Number of followers

Documents

Last sold

0.0

0 reviews

Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Frequently asked questions

What do I get when I buy this document?

You get a PDF, available immediately after your purchase. The purchased document is accessible anytime, anywhere and indefinitely through your profile.

Satisfaction guarantee: how does it work?

Our satisfaction guarantee ensures that you always find a study document that suits you well. You fill out a form, and our customer service team takes care of the rest.

Who am I buying these notes from?

Stuvia is a marketplace, so you are not buying this document from us, but from seller guillemribes. Stuvia facilitates payment to the seller.

Will I be stuck with a subscription?

No, you only buy these notes for $8.67. You're not tied to anything after your purchase.

Can Stuvia be trusted?

4.6 stars on Google & Trustpilot (+1000 reviews) 46231 documents were sold in the last 30 days Founded in 2010, the go-to place to buy study notes for 15 years now

Summary of paper End-to-end Object Detection with Transformers

Written for

Document information

Subjects

Content preview

Also available in package deal

Get to know the seller

Recently viewed by you

Why students choose Stuvia

Created by fellow students, verified by reviews

Didn't get what you expected? Choose another document

Pay as you like, start learning right away

Frequently asked questions

What do I get when I buy this document?

Satisfaction guarantee: how does it work?

Who am I buying these notes from?

Will I be stuck with a subscription?

Can Stuvia be trusted?