Topics
Total pages: 14
File: topics/topic-01.md - Weight: 10
File: topics/topic-02.md - Weight: 20
File: topics/topic-03.md - Weight: 30
File: topics/topic-04.md - Weight: 40
File: topics/topic-05.md - Weight: 50
File: topics/topic-06.md - Weight: 60
File: topics/topic-07.md - Weight: 70
File: topics/topic-08.md - Weight: 80
File: topics/topic-09.md - Weight: 90
File: topics/topic-10.md - Weight: 100
File: topics/topic-11.md - Weight: 110
File: topics/topic-12.md - Weight: 120
File: topics/topic-13.md - Weight: 130
File: topics/topic-14.md - Weight: 140
Overview of class topics. Click on the titles for details.
Class Dates: 2026-08-26
This session introduces Python and the environment you will be working in for the rest of the term.
Class Dates: 2026-09-02
This session introduces Python as a language with its own syntax and conventions, establishing the habits of readable, consistent code before any complexity is added.
Class Dates: 2026-09-09
This session covers fundamental Python data structures, iteration techniques, and practical loop patterns, using a campus coffee cart sales dataset as a practical working example.
Class Dates: 2026-09-16
Decomposition is how a problem too large to hold in your head becomes a handful of steps you can name, write, and check one at a time. In Python those steps are functions: each one a contract with a single responsibility and clear expectations about its inputs, its return value, and its edge cases.
Class Dates: 2026-09-23
A file outlives the session that created it, which is what allows one program’s output to become another’s input. Safe file handling rests on a few habits: a with block that closes the file whatever happens, an explicit UTF-8 encoding, input and output locations that are named rather than assumed, and a read-back of anything written. Which format to write in depends on the shape of the data. CSV suits simple rectangular tables and moves easily between spreadsheets and programs, at the cost of careful handling of headers, quoting, embedded commas, missing fields, and the conversion of every value from text. JSON carries structured documents — nested objects and lists, numbers, booleans, and null — which makes it a fit for configuration, metadata, API payloads, and records that are not flat. JSONL puts one complete JSON record on each line, so large collections of independent records such as logs, events, and model outputs can be appended and processed one record at a time.
Class Dates: 2026-09-30
Unix presents every disk, every home directory, and every data file as one tree rooted at /, and gives you a text interface for moving through it. Learn where you are (pwd), what is there (ls), and how to get elsewhere (cd), and the rest — absolute versus relative paths, . and .., ~ — stops being syntax to memorize and becomes a map you can read.
Class Dates: 2026-10-07
Motivate version control from problems students already have: files named final_v3_actually_final, a change that broke working code with no way back, and no record of what was done when. Present a commit as a snapshot of the project with an author, a timestamp, and a message, and the repository as the sequence of those snapshots. Develop the three-area model — working tree, staging area, repository — because most beginner confusion with Git comes from not knowing which area a command acts on. Distinguish local history from a remote copy, and explain what pushing and pulling actually move. Frame commit messages as documentation written for a future reader, and connect version control to reproducibility more broadly: a project is reproducible when the code, the data references, and the sequence of changes are all recoverable, which is also what makes authorship verifiable.
Class Dates: 2026-10-14
Introduce the table — rows as observations, columns as variables — as the central data structure of analytics, and contrast it with the lists and dictionaries students have been assembling by hand. Explain vectorized thinking: expressing an operation over a whole column at once instead of looping over elements, which is both more readable and dramatically faster because the work happens in compiled code. Present split-apply-combine as the conceptual model behind grouped computation: partition rows by a key, compute a summary per group, and reassemble the results. Treat missing data as a substantive question rather than a technical nuisance — a blank cell may mean zero, unknown, or not applicable, and dropping versus filling changes what the resulting numbers mean. Close on the habit of inspecting a dataset (shape, types, ranges, missingness) before drawing any conclusion from it.
Class Dates: 2026-10-21
Develop summary statistics and visualization as two views of the same question, each covering the other’s blind spots. Define mean, median, variance, and standard deviation, and show why the mean and median diverge under skew and why a single number can describe very different distributions identically. Introduce distribution shape, spread, and outliers, and treat an outlier as something to investigate rather than automatically discard. Connect chart type to question type: distribution of one variable, comparison across categories, relationship between two variables, change over time. Discuss honest encoding — truncated axes, misleading aspect ratios, missing labels, and reading correlation in a scatter plot as if it were causation — and establish that a figure without axis labels and units is not yet a result.
Class Dates: 2026-10-28
Introduce simulation as a way to answer a question by running an experiment many times instead of solving it analytically, which makes problems accessible long before the corresponding mathematics is. Explain pseudorandomness: the generator is deterministic given its seed, so results are reproducible on demand — a property that is essential for grading, debugging, and scientific reporting, not a limitation. Build Monte Carlo intuition through the estimate-as-average idea, and emphasize the point beginners most often miss: a simulated result is itself uncertain, it varies from run to run, and more trials shrink that variability in a predictable way. Discuss how to report a simulation honestly by stating the number of trials and the seed, and how the same machinery underlies resampling, sensitivity analysis, and probabilistic reasoning about real data.
Class Dates: 2026-11-04
Use search as the entry point to algorithmic thinking: the same question — is this item present, and where — admits strategies with very different costs. Develop linear search as the general method that always works, then binary search as a much faster method that buys its speed with an assumption (the data must be sorted) and works by halving the remaining range. Introduce the intuition behind hashing: a hash function maps a key to a location, so a dictionary can go more or less directly to the value instead of scanning, which is why lookup cost barely grows with size. Present efficiency informally but honestly, in terms of how the work scales as data grows rather than formal notation, and name the tradeoffs: sorting has an upfront cost, dictionaries and sets use extra memory, and the fastest approach depends on how many lookups you will perform.
Class Dates: 2026-11-11
Introduce the graph as a model for anything defined by relationships rather than by rows: road networks, social connections, dependencies, web links. Define nodes and edges, and distinguish undirected from directed edges and unweighted from weighted ones, since each distinction changes what a “shortest” path means. Compare adjacency representations — an edge list, an adjacency matrix, and an adjacency list built from a dictionary of neighbor lists — in terms of memory and the cost of asking who is adjacent to a given node. Develop breadth-first search as exploration in rings of increasing distance, and derive the key result intuitively: because BFS reaches nodes in order of hop count, the first time it reaches the target it has found a shortest path in an unweighted graph. Note where weights break that argument and where Dijkstra’s idea picks up, without developing it formally.
Class Dates: 2026-11-18
Frame supervised learning as fitting a mapping from features to a label using examples where the label is known, and connect it to the tabular data students already handle: columns become features, one column becomes the target. Distinguish classification (a categorical label, scored by accuracy and error rates) from regression (a numeric label, scored by error magnitude). Make the case for held-out data carefully, because it is the central idea of the session: a model evaluated on the data it was trained on can memorize rather than generalize, so its score is not evidence that it works. Develop overfitting intuition through model flexibility — a model complex enough to fit noise will do so — and introduce the train/test split as the minimal defense. Close on interpretation: accuracy is meaningless without a baseline, class imbalance can make a high score worthless, and a model that predicts well still explains nothing about cause.
Class Dates: 2026-12-02
Revisit the course as a connected whole rather than a list of topics: values and types, control flow, functions, files, tabular data, and algorithmic ideas are the layers of a single toolkit, and a realistic task uses several at once. Treat debugging explicitly as a systematic method — reproduce the failure, read the error, form a hypothesis, shrink the input, test one change at a time — and contrast it with guess-and-rerun, which is where beginners lose the most time. Consolidate reproducibility as the through-line of the second half of the course: clear file organization, code that runs from the command line, fixed random seeds, and a commit history that records how the work developed. Close by connecting these skills to what comes next, including where each idea reappears in later statistics, analytics, and machine learning coursework.