ORIE 4741/5741: Learning with Big Messy Data

Professor Madeleine Udell, Cornell University

Logistics

COVID-19: In Fall 2021, this course will be available both in-person and online. Students will be able to complete the course exclusively online for full credit. If you are unvaccinated, feel sick, or have a known exposure, please attend class online.

Schedule:

Class on Tuesday and Thursday 9:40 – 10:55am on Zoom or in Morrison Hall room 146
Discussion sections are held T 2:40pm - 3:30pm and W 11:20am - 12:10pm, on Zoom or in Rhodes 253 (Tuesday) or Upson 222 (Wednesday). Attendance is optional. You can enroll in one for a 4 credit option, or omit for 3 credits.
All course events (lectures, sections, and office hours) are listed on the course calendar with their times and location info.

Lectures and discussion sections will be held in-person and streamed online. Office hours will be mixed, some in-person and some online. Students are strongly encouraged to attend synchronously (at the scheduled times), but all course components (lectures and section) will be recorded. To receive full participation credit for the lectures, students attending lectures asynchronously must respond to questions on each lecture before the following lecture; see the participation tab for details.

Discussion forum: zulip.

Enrollment: We expect that everyone who want to take the class will be able to enroll. Note that you can enroll without the discussion (for 3 credits), and that you can enroll in a discussion section even if you plan to attend it asynchronously. You're also welcome to attend the discussion even if you can't enroll in it.

4741 vs 5741: ORIE 4741 and 5741 are substantively similar. ORIE 5741 is designed for graduate students, and has different project requirements, including a more business-oriented project and a final project presentation. The courses are graded separately.

Overview

Modern data sets, whether collected by scientists, engineers, physicians, bureaucrats, financiers, or tech billionaires, are often big, messy, and extremely useful. This course addresses scalable robust methods for learning from big messy data. We will cover techniques for learning with data that is messy — consisting of measurements that are continuous, discrete, boolean, categorical, or ordinal, or of more complex data such as graphs, texts, or sets, with missing entries and with outliers — and that is big — which means we can only use algorithms whose complexity scales linearly in the size of the data. We will cover techniques for cleaning data, supervised and unsupervised learning, finding similar items, model validation, and feature engineering. The course will culminate in a final project in which students extract useful information from a big messy data set.

Prerequisites: Familiarity with linear algebra and matrix notation, a modern scripting language (such as Python, Matlab, Julia, R), and basic complexity and O(n) notation. More formally, we strongly recommend

Linear Algebra (MATH 2940 or equivalent). Important topics: inner products, matrix multiplication, singular value decomposition.
Probability (ENGRD 2700 or equivalent). Important topics: random sampling, maximum likelihood estimation.
Programming (ENGRD/CS 2110 or equivalent). Important topics: basic comfort in a scripting language, iteration, functions.

Students familiar with these prerequisites in past years found that homework assignments took about about ten hours a week. Students without these prerequisites found that homework assignments took up to thirty hours a week as they caught up on background knowledge. Hence if you are tempted to ignore these prerequisites, please make sure you can afford to spend thirty hours a week on this class.

Announcements

Sign up for the class by completing the course survey, or ask questions about the course on our discussion forum, zulip.
On the waitlist? We expect that you'll be able to get in to the class. Please be patient and keep up with work in the meantime.
Want to get a head start on learning with big messy data? You might try learning Python, reviewing linear algebra, or reading the book Learning from Data. See about for more ideas.
Course materials on this website reflect the Fall 2020 course. Lectures and topics may change slightly in Fall 2021 to reflect student interest.