Homework Assignments

Homework assignments in this course are individual tasks designed to enhance student learning and provide opportunities for independent practice. Detailed instructions for completing each assignment, along with submission guidelines, will be made available on the class website.

It is essential for students to adhere strictly to these submission instructions, as assignments that fail to comply or are submitted past the deadline will not be considered for grading.

This policy ensures fairness and consistency in the evaluation process while emphasizing the importance of personal responsibility in meeting academic expectations.

The instructions and grading criteria are subject to change. Please check this site for updates and follow announcement in class and on iCollege


Assignments

Table with Homework Assignments and Due Dates
AssignmentDue DateStatus
Homework 1: Text Analysis of the 20 Newsgroups CorpusWednesday, September 16, 2026 at 23:59Posted
Homework 2: Bike-Share Trip AnalysisWednesday, September 23, 2026 at 23:59Posted
Homework 3: Ranking Scholarship ApplicationsWednesday, September 30, 2026 at 23:59Posted
Homework 4: A Pipeline Through Five Gigabytes of ReviewsWednesday, October 7, 2026 at 23:59Scheduled
Homework 5: A Query Engine Made of PipesWednesday, October 14, 2026 at 23:59Scheduled
Homework 6: Git and Reproducible WorkflowWednesday, October 21, 2026 at 23:59Scheduled
Homework 7: Python for Data AnalysisWednesday, October 28, 2026 at 23:59Scheduled
Homework 8: Visualization and Descriptive StatisticsWednesday, November 4, 2026 at 23:59Scheduled
Homework 9: Simulation and RandomnessWednesday, November 11, 2026 at 23:59Scheduled
Homework 10: Search and HashingWednesday, November 18, 2026 at 23:59Scheduled
Homework 11: Graphs and Shortest PathWednesday, November 25, 2026 at 23:59Scheduled
Homework 12: Introductory Machine LearningWednesday, December 2, 2026 at 23:59Scheduled

Please review instructions for completing and submitting each of the assignments.

Homework 1: Text Analysis of the 20 Newsgroups Corpus

Due: Wednesday, September 16, 2026 at 23:59
Build a complete command-line program that analyzes a corpus of newsgroup messages using nothing but core Python — no machine learning libraries, no regular expressions. Across ten test-driven phases you will read and tokenize text files, score each article for sentiment against the AFINN-111 lexicon, count terms with nested dictionaries, and rank the most distinctive keywords of each newsgroup by TF–IDF, then assemble everything into a single deterministic report. Every phase ships with its own tests, so you find out whether the piece you just wrote works before you build the next one on top of it.

Homework 2: Bike-Share Trip Analysis

Due: Wednesday, September 23, 2026 at 23:59
Analyze a month of campus bike-share trips using nothing but core Python — no pandas, no NumPy, no file reading, just a list of dictionaries and the loop patterns from

Homework 3: Ranking Scholarship Applications

Due: Wednesday, September 30, 2026 at 23:59
Order a list of scholarship applications from the lowest review score to the highest, without sorted(), .sort(), min(), or max(). Across seven test-driven phases you will build the solution out of small functions with clear contracts: one comparison rule that settles what a tie means, a function that combines two ordered groups, a function that splits a group in two, a function that orders the groups too small to split, and a function that uses all of them to order a group of any size. Then you will add a policy for applications with missing or unusable scores and print an exactly formatted ranking report. Each phase ships with its own tests, and only at the end are you asked to name the method you have built.

Homework 4: A Pipeline Through Five Gigabytes of Reviews

Due: Wednesday, October 7, 2026 at 23:59
Find the restaurants Yelp reviewers disagreed about most in 2018 and 2019, working from the 5.3 GB review file on the cluster, which is far too large to load into memory. Across nine test-driven phases you will build a pipeline of stages in core Python, each reading the file the stage before it wrote, one record at a time. You will filter the business and review files, join reviews to a small lookup table, compute count, mean, and standard deviation for thousands of restaurants without storing the values, and sort a file of any size by reusing your HW03 merge sort on files. Finally you will run the whole pipeline from the command line and print a ranked report. Every phase is tested on a small sample of the dataset before you run it on the real thing.

Homework 5: A Query Engine Made of Pipes

Due: Wednesday, October 14, 2026 at 23:59
Answer one question about 863,077 US Forest Service records of invasive plants — which ones cover the most ground, and where — by building the machine that runs the SQL rather than writing the SQL. Across eight test-driven phases you will write a single Python file that becomes seven command-line tools, each reading CSV on stdin and writing CSV on stdout, then compose them into a pipeline with Unix pipes: filter, select, group-by, aggregate, order-by, head. Along the way you will parse CSV correctly including quoted fields with embedded commas and newlines, decide what a text field actually means and handle nulls the way SQL does, stream rows so memory does not grow with the file, sort by several keys in either direction, and compute count, sum, avg, min, and max for a group without ever holding the group’s rows.

Homework 6: Git and Reproducible Workflow

Due: Wednesday, October 21, 2026 at 23:59
Put your course work under version control and build a readable history. You will create a repository, stage and commit changes in meaningful increments, inspect what changed and when, exclude files that do not belong in history, and synchronize your work with a remote.

Homework 7: Python for Data Analysis

Due: Wednesday, October 28, 2026 at 23:59
Analyze a real dataset with NumPy and Pandas, treating the table as a first-class structure rather than a collection of loops. You will load and inspect data, select and filter rows and columns, apply split-apply-combine to answer grouped questions, and decide — explicitly — how to handle missing values.

Homework 8: Visualization and Descriptive Statistics

Due: Wednesday, November 4, 2026 at 23:59
Describe a dataset both numerically and visually, and confront what each summary hides. You will compute measures of center and spread, characterize distribution shape and outliers, and choose chart types that match the question you are asking — with honest encoding, complete labels, and figures saved for reuse.

Homework 9: Simulation and Randomness

Due: Wednesday, November 11, 2026 at 23:59
Estimate an answer by running the experiment many times instead of solving it analytically. You will build a Monte Carlo simulation, use seeds to make pseudorandom results reproducible, and examine how much your estimate varies from run to run as the number of trials grows.

Homework 10: Search and Hashing

Due: Wednesday, November 18, 2026 at 23:59
Implement search from scratch and measure what it costs. You will write linear and binary search, identify the assumption binary search depends on, and time lookups across lists, dictionaries, and sets to see for yourself why hash-based lookup wins — your first empirical look at efficiency tradeoffs.

Homework 11: Graphs and Shortest Path

Due: Wednesday, November 25, 2026 at 23:59
Model a set of relationships as a graph and find your way through it. You will represent nodes and edges with dictionaries, traverse neighbors, implement breadth-first search with a queue, and reconstruct the shortest path between two points — the core machinery behind routing, recommendation, and network analysis.

Homework 12: Introductory Machine Learning

Due: Wednesday, December 2, 2026 at 23:59
Train and evaluate your first supervised model with scikit-learn. You will prepare features and labels, hold out test data, fit a model, and judge its predictions against a sensible baseline — building the intuition for why held-out data is necessary and what overfitting looks like when it happens.