Homework Assignments
Homework assignments in this course are individual tasks designed to enhance student learning and
provide opportunities for independent practice. Detailed instructions for completing each assignment,
along with submission guidelines, will be made available on the class website.
It is essential for students to adhere strictly to these submission instructions, as assignments that fail to comply or are submitted
past the deadline will not be considered for grading.
This policy ensures fairness and consistency in the
evaluation process while emphasizing the importance of personal responsibility in meeting academic
expectations.
The instructions and grading criteria are subject to change. Please check this site for updates and follow announcement in class and on iCollege
Assignments
Table with Homework Assignments and Due Dates| Assignment | Due Date | Status |
|---|
| Homework 1: Text Analysis of the 20 Newsgroups Corpus | Wednesday, September 16, 2026 at 23:59 | Posted |
| Homework 2: Bike-Share Trip Analysis | Wednesday, September 23, 2026 at 23:59 | Posted |
| Homework 3: Ranking Scholarship Applications | Wednesday, September 30, 2026 at 23:59 | Posted |
| Homework 4: A Pipeline Through Five Gigabytes of Reviews | Wednesday, October 7, 2026 at 23:59 | Scheduled |
| Homework 5: A Query Engine Made of Pipes | Wednesday, October 14, 2026 at 23:59 | Scheduled |
| Homework 6: Git and Reproducible Workflow | Wednesday, October 21, 2026 at 23:59 | Scheduled |
| Homework 7: Python for Data Analysis | Wednesday, October 28, 2026 at 23:59 | Scheduled |
| Homework 8: Visualization and Descriptive Statistics | Wednesday, November 4, 2026 at 23:59 | Scheduled |
| Homework 9: Simulation and Randomness | Wednesday, November 11, 2026 at 23:59 | Scheduled |
| Homework 10: Search and Hashing | Wednesday, November 18, 2026 at 23:59 | Scheduled |
| Homework 11: Graphs and Shortest Path | Wednesday, November 25, 2026 at 23:59 | Scheduled |
| Homework 12: Introductory Machine Learning | Wednesday, December 2, 2026 at 23:59 | Scheduled |
Please review instructions for completing and submitting each of the assignments.
Due: Wednesday, September 16, 2026 at 23:59
Build a complete command-line program that analyzes a corpus of newsgroup messages using nothing but core Python — no machine learning libraries, no regular expressions. Across ten test-driven phases you will read and tokenize text files, score each article for sentiment against the AFINN-111 lexicon, count terms with nested dictionaries, and rank the most distinctive keywords of each newsgroup by TF–IDF, then assemble everything into a single deterministic report. Every phase ships with its own tests, so you find out whether the piece you just wrote works before you build the next one on top of it.
Due: Wednesday, September 23, 2026 at 23:59
Analyze a month of campus bike-share trips using nothing but core Python — no pandas, no NumPy, no file reading, just a list of dictionaries and the loop patterns from
Due: Wednesday, September 30, 2026 at 23:59
Order a list of scholarship applications from the lowest review score to the highest, without sorted(), .sort(), min(), or max(). Across seven test-driven phases you will build the solution out of small functions with clear contracts: one comparison rule that settles what a tie means, a function that combines two ordered groups, a function that splits a group in two, a function that orders the groups too small to split, and a function that uses all of them to order a group of any size. Then you will add a policy for applications with missing or unusable scores and print an exactly formatted ranking report. Each phase ships with its own tests, and only at the end are you asked to name the method you have built.
Due: Wednesday, October 7, 2026 at 23:59
Find the restaurants Yelp reviewers disagreed about most in 2018 and 2019, working from the 5.3 GB review file on the cluster, which is far too large to load into memory. Across nine test-driven phases you will build a pipeline of stages in core Python, each reading the file the stage before it wrote, one record at a time. You will filter the business and review files, join reviews to a small lookup table, compute count, mean, and standard deviation for thousands of restaurants without storing the values, and sort a file of any size by reusing your HW03 merge sort on files. Finally you will run the whole pipeline from the command line and print a ranked report. Every phase is tested on a small sample of the dataset before you run it on the real thing.
Due: Wednesday, October 14, 2026 at 23:59
Answer one question about 863,077 US Forest Service records of invasive plants — which ones cover the most ground, and where — by building the machine that runs the SQL rather than writing the SQL. Across eight test-driven phases you will write a single Python file that becomes seven command-line tools, each reading CSV on stdin and writing CSV on stdout, then compose them into a pipeline with Unix pipes: filter, select, group-by, aggregate, order-by, head. Along the way you will parse CSV correctly including quoted fields with embedded commas and newlines, decide what a text field actually means and handle nulls the way SQL does, stream rows so memory does not grow with the file, sort by several keys in either direction, and compute count, sum, avg, min, and max for a group without ever holding the group’s rows.
Due: Wednesday, October 21, 2026 at 23:59
Put your course work under version control and build a readable history. You will create a repository, stage and commit changes in meaningful increments, inspect what changed and when, exclude files that do not belong in history, and synchronize your work with a remote.
Due: Wednesday, October 28, 2026 at 23:59
Analyze a real dataset with NumPy and Pandas, treating the table as a first-class structure rather than a collection of loops. You will load and inspect data, select and filter rows and columns, apply split-apply-combine to answer grouped questions, and decide — explicitly — how to handle missing values.
Due: Wednesday, November 4, 2026 at 23:59
Describe a dataset both numerically and visually, and confront what each summary hides. You will compute measures of center and spread, characterize distribution shape and outliers, and choose chart types that match the question you are asking — with honest encoding, complete labels, and figures saved for reuse.
Due: Wednesday, November 11, 2026 at 23:59
Estimate an answer by running the experiment many times instead of solving it analytically. You will build a Monte Carlo simulation, use seeds to make pseudorandom results reproducible, and examine how much your estimate varies from run to run as the number of trials grows.
Due: Wednesday, November 18, 2026 at 23:59
Implement search from scratch and measure what it costs. You will write linear and binary search, identify the assumption binary search depends on, and time lookups across lists, dictionaries, and sets to see for yourself why hash-based lookup wins — your first empirical look at efficiency tradeoffs.
Due: Wednesday, November 25, 2026 at 23:59
Model a set of relationships as a graph and find your way through it. You will represent nodes and edges with dictionaries, traverse neighbors, implement breadth-first search with a queue, and reconstruct the shortest path between two points — the core machinery behind routing, recommendation, and network analysis.
Due: Wednesday, December 2, 2026 at 23:59
Train and evaluate your first supervised model with scikit-learn. You will prepare features and labels, hold out test data, fit a model, and judge its predictions against a sensible baseline — building the intuition for why held-out data is necessary and what overfitting looks like when it happens.