Repository navigation
Expand file tree
/
Copy pathfinal-exam.html
More file actions
216 lines (188 loc) · 16.2 KB
/
Copy pathfinal-exam.html
File metadata and controls
216 lines (188 loc) · 16.2 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta http-equiv="X-UA-Compatible" content="IE=edge">
<meta name="viewport" content="width=device-width, initial-scale=1">
<meta name="description" content="">
<meta name="author" content="">
<title>CS5785 Final Exam</title>
<!-- Bootstrap Core CSS -->
<link href="css/bootstrap.min.css" rel="stylesheet">
<!-- Custom CSS -->
<link href="css/clean-blog.min.css" rel="stylesheet">
<!-- Custom Fonts -->
<!-- <link href="http://maxcdn.bootstrapcdn.com/font-awesome/4.1.0/css/font-awesome.min.css" rel="stylesheet" type="text/css">
<link href='http://fonts.googleapis.com/css?family=Lora:400,700,400italic,700italic' rel='stylesheet' type='text/css'>
<link href='http://fonts.googleapis.com/css?family=Open+Sans:300italic,400italic,600italic,700italic,800italic,400,300,600,700,800' rel='stylesheet' type='text/css'> -->
<!-- HTML5 Shim and Respond.js IE8 support of HTML5 elements and media queries -->
<!-- WARNING: Respond.js doesn't work if you view the page via file:// -->
<!--[if lt IE 9]>
<script src="https://oss.maxcdn.com/libs/html5shiv/3.7.0/html5shiv.js"></script>
<script src="https://oss.maxcdn.com/libs/respond.js/1.4.2/respond.min.js"></script>
<![endif]-->
</head>
<body>
<!-- Navigation -->
<nav class="navbar navbar-default navbar-custom navbar-fixed-top">
<div class="container-fluid">
<!-- Brand and toggle get grouped for better mobile display -->
<div class="navbar-header page-scroll">
<button type="button" class="navbar-toggle" data-toggle="collapse" data-target="#bs-example-navbar-collapse-1">
<span class="sr-only">Toggle navigation</span>
<span class="icon-bar"></span>
<span class="icon-bar"></span>
<span class="icon-bar"></span>
</button>
<a class="navbar-brand" href="http://tech.cornell.edu/">Cornell Tech</a>
</div>
<!-- Collect the nav links, forms, and other content for toggling -->
<div class="collapse navbar-collapse" id="bs-example-navbar-collapse-1">
<ul class="nav navbar-nav navbar-right">
<li>
<a href="index.html">Home</a>
</li>
<li>
<a href="lectures.html">Lectures</a>
</li>
<li>
<a href="assignments.html">Assignments</a>
</li>
<li>
<a href="contact.html">Contact</a>
</li>
</ul>
</div>
<!-- /.navbar-collapse -->
</div>
<!-- /.container -->
</nav>
<!-- Page Header -->
<!-- Set your background image for this header on the line below. -->
<header class="intro-header" style="background-image: url('img/banner.png'); background-color: #777;">
<div class="container">
<div class="row">
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
<div class="page-heading">
<h1>Final Exam</h1>
<hr class="small">
<span class="subheading">Instructions</span>
</div>
</div>
</div>
</div>
</header>
<!-- Main Content -->
<div class="container">
<div class="row">
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
<br /><br /><br /><br />
<h1>Please see the exam page on <a href="https://inclass.kaggle.com/c/cs5785-spring-2017-final">Kaggle</a>!</h1>
<br /><br /><br /><br />
<br /><br /><br /><br />
<div>
<h1>About The Exam</h1>
<p>The format of the exam is a mock peer-reviewed conference. You will develop an algorithm, prepare a professional paper, submit an anonymized version to the EasyChair conference system, and peer-review the work from other groups. There are three deliverables:</p>
<ul>
<li><strong>Phase 1: Develop your algorithm and write your writeup</strong></li>
<ul>
<li>Develop your algorithm and submit your results to the Kaggle leaderboard. <strong>Deadline: Monday, May 15, 11:59 PM EST.</strong></li>
<li>Write your paper submission and submit to CMS. <strong>Deadline: Monday, May 15, 11:59 PM EST</strong>. There should be only one submission from each team.</li>
</ul>
<li><strong>Phase 2: Peer-review other groups' work!</strong></li>
<ul>
<li>Write reviews and submit them to CMS. <strong>Deadline: Tuesday, May 16, 11:59 PM EST</strong>. (Will open after Phase 1 ends)</li>
<li><em>Everybody</em> on each team must complete peer reviews individually. (They are not "per-team")</li>
</ul>
</ul>
<h1>The Challenge: Build a large-scale image search engine!</h1>
<p>You and your team of <strong>two Cornell Tech students</strong> are surely on the path to fame and fortune! You have been recruited by Google to disrupt Google Image Search by building a better search engine using novel statistical learning techniques.</p>
<p>The specifications are simple: We need a way to <strong>search for relevant images</strong> given a natural language query. For instance, if a user types "dog jumping to catch frisbee," your system will <strong>rank-order the most relevant images</strong> from a large database.</p>
<p>Here are some details:</p>
<ul>
<li><strong>During training,</strong> you have a dataset of 10,000 samples. Each sample has the following data available for learning:</li>
<ul>
<li>A 224x224 JPG image.</li>
<li>A list of tags indicating objects appeared in the image.</li>
<li>Feature vectors extracted using <a href="https://arxiv.org/abs/1512.03385">ResNet</a>, a state-of-the-art Deep-learned CNN (You don't have to train or run ResNet -- we are providing the features for you). See <a href="http://ethereon.github.io/netscope/#/gist/b21e2aae116dc1ac7b50">here</a> for the illustration of the ResNet-101 architecture. The features are extracted from <em>pool5</em> and <em>fc1000</em> layer.</li>
<li>A five-sentence description, used to train your search engine.</li>
</ul>
<li><strong>During testing,</strong> your system matches a <em>single five-sentence description</em> against a pool of 2,000 <em>candidate samples</em> from the test set. Each sample has:</li>
<ul>
<li>A 224x224 JPEG image.</li>
<li>A list of tags for that image.</li>
<li>ResNet feature vectors for that image.</li>
</ul>
</ul>
<ul>
<li><strong>Output</strong>: For each description, your system must <strong>rank-score</strong> each testing image with the likelihood of that image matches the given sentence. Your system then returns the name of the top 20 relevant images, delimited by space. See "sample_submission.csv" on the data page for more details on the output format.</li>
<li><strong>Evaluation metric</strong>: There are 2,000 descriptions, and for each description, you must compare against the entire 2,000-image test set. That is, rank-order test images for each test description. We will use <strong>MAP@20</strong> as the evaluation metric. If the corresponding image of a description is among your algorithm's 20 highest scoring images, this metric gives you a certain score based on the ranking of the corresponding image. Please refer to the evaluation page for more details. </li>
</ul>
<p>Use all of your skills, tools, and experience. It is OK to use libraries like numpy, scikit-learn, pandas, etc., as long as you cite them. Use cross-validation on training set to debug your algorithm. Submit your results to the Kaggle leaderboard and send your complete writeup to EasyChair. The data you use --- and the way you use the data --- is completely up to you.</p>
<p>The best teams (of <strong>two Cornell Tech students</strong>, recall) might use visualization techniques for debugging (e.g., show top images retrieved by your algorithm and see whether they make sense or not), preprocessing, a nice way to compare tags and descriptions, leveraging visual features and combining them with tags and descriptions, supervised and/or unsupervised learning to best understand how to best take advantage of each data source available to them.</p>
<p>The best reports will be professionally written. We suggest including an <strong>Abstract,</strong> followed by an <strong>Introduction</strong> section including motivations of your method and a detailed description of each step in your pipeline / framework you used for this task. Then, include a detailed <strong>Experiment</strong> section and a <strong>Results</strong> section. If time permits, a <strong>Background / Related Work</strong> section can help to point out the relevant work in the image retrieval literature.</p>
<p>The best reports use a professional style adopted by a major academic conference. NIPS is a good choice. Teams can download a template from Section 4 Paper Format of the <a href="https://nips.cc/Conferences/2015/PaperInformation/AuthorSubmissionInstructions">NIPS Author Guidelines</a>. We recommend you to use the LaTeX template, but this is not a requirement. If you're more familiar with Word, use Word.</p>
<p>The best peer reviews will be detailed and thorough, pointing out potential areas of improvement as well as complimenting their peers' strengths.</p>
<p>The best teams understand how to divide the precious time and energy among all members of the group. They might split into roles, preventing programming students from messing up the writing and ensuring writing gurus don't screw up the code. Close contact is always maintained within top-performing teams. When hardship happens within the best teams, blame is not assigned---it is overcome.</p>
<p>May the best team win!</p>
<h2>Rules</h2>
<p>You are strongly encouraged to work in a group of <strong>two</strong> students.</p>
<p>Groups may not collaborate with other groups. You cannot cite other teams' work. This includes sharing data, intermediate and final results, writing, or discussing your detailed method with other groups. You <strong>may not</strong> view reports from other teams until <strong>after</strong> the peer-review phase begins!</p>
<h1>How to anonymize your paper</h1>
<p>The "Peer Review" phase will be double-blind. This means authors will not know who the reviewers are, and the reviewers should not know who the authors are. This protects both sides.</p>
<p>To maintain anonymity, <strong>don't</strong> include team member names in the actual paper that you submit to Dropbox!</p>
<p>However, we need to know who you are. So please use the following file name:
<code>TeamNameOnKaggle_Author1lastname_Author2lastname_report.pdf</code>
This will be hidden from your reviewers.</p>
<p><img src="http://i.imgur.com/F255rFT.png" alt="" width="368" height="309"></p>
<h2>Peer Review Tips</h2>
<p>See the following resources for tips on how to write a professional review:</p>
<ul>
<li><span>Matthew Might’s notes on </span><a href="http://matt.might.net/articles/how-to-peer-review/">how to peer review</a><span> and </span><a href="http://matt.might.net/articles/peer-fortress/">Peer Fortress: The Scientific Battlefield</a><span> (the last one provides examples NOT to follow)</span></li>
<li><a href="http://www.pamitc.org/cvpr16/reviewer_guidelines.php">CVPR 2016 Reviewer’s Guidelines</a><span>.</span></li>
<li><a href="http://mobilehci.acm.org/2015/download/ExcellenceInReviewsforHCICommunity.pdf" class="pdf-link">So You're a Program Committee Member Now: On Excellence in Reviews and Meta-Reviews and Championing Submitted Work That Has Merit</a></li>
</ul>
<h1>Score Breakdown</h1>
<ul>
<li><span>30% on Kaggle performance (score on the private test set).</span></li>
<li><span>30% on report: is the proposed method clearly described? is it professionally written? does it include enough detail for a professional (ie. skilled graduate student) to re-implement the results? A good report should at least contain a brief introduction of what you did, a detailed description of your method, an experimental evaluation part that shows the experimental results and analysis of your results. </span></li>
<li><span>20% on method: does it make sense? is it suited to the task? does it make use of the data in appropriate ways?</span></li>
<li><span>10% on peer reviews from other classmates.</span></li>
<li><span>10% on the quality of peer review: is it well-written and thoughtful? does it provide insight to the authors?</span></li>
</ul>
</div>
<div>
<h2>Data files: Download <a href="https://mjw-cornell-se3-public.s3.amazonaws.com/cs5785-final-spring2017-data.zip">here!</a></h2>
<ul>
<li><strong>images_train</strong> - 10,000 training images of size 224x224.</li>
<li><strong>images_test</strong> - 2,000 test images of size 224x224.</li>
<li><strong>tags_train</strong> - image tags correspond to training images. Each image have several tags indicating the human-labeled object categories appear in the image, in the form of "supercategory:category".</li>
<li><strong>tags_test</strong> - image tags correspond to test images. Each image have several tags indicating the human-labeled object categories appear in the image, in the form of "supercategory:category".</li>
<li><strong>features_train</strong> - features extracted from a pre-trained Residual Network (ResNet) on training set, including 1,000 dimensional feature from classification layer (fc1000) and 2,048 dimensional feature from final convolution layer (pool5). Each dimension of the fc1000 feature corresponds to a WordNet synset <a href="https://gist.github.com/yrevar/942d3a0ac09ec9e5eb3a">here</a>.</li>
<li><strong>features_test</strong> - features extracted from the same Residual Network (ResNet) on test set, including 1,000 dimensional feature from classification layer (fc1000) and 2,048 dimensional feature from final convolution layer (pool5).</li>
<li><strong>descriptions_train</strong> - image descriptions correspond to training images. Each image have 5 sentences for describing the image content.</li>
<li><strong>descriptions_test</strong> - image descriptions for test images. Each image have 5 sentences for describing the image content. <em><strong>Notice that one test description corresponds to one test image. The task you need to do is to return top 20 images in test set for each test description.</strong></em></li>
<li><strong>sample_submission.csv</strong> - a sample submission file in the correct format.</li>
</ul>
</div>
</div>
</div>
</div>
<hr>
<!-- Footer -->
<footer>
<div class="container">
<div class="row">
<div class="col-lg-8 col-lg-offset-2 col-md-10 col-md-offset-1">
<p class="copyright text-muted">Copyright © Applied Machine Learning 2017</p>
</div>
</div>
</footer>
<!-- jQuery -->
<script src="js/jquery.js"></script>
<!-- Bootstrap Core JavaScript -->
<script src="js/bootstrap.min.js"></script>
<!-- Custom Theme JavaScript -->
<script src="js/clean-blog.min.js"></script>
</body>
</html>