Files, Modules, and Scripts

A file outlives the session that created it, which is what allows one program’s output to become another’s input. Safe file handling rests on a few habits: a with block that closes the file whatever happens, an explicit UTF-8 encoding, input and output locations that are named rather than assumed, and a read-back of anything written. Which format to write in depends on the shape of the data. CSV suits simple rectangular tables and moves easily between spreadsheets and programs, at the cost of careful handling of headers, quoting, embedded commas, missing fields, and the conversion of every value from text. JSON carries structured documents — nested objects and lists, numbers, booleans, and null — which makes it a fit for configuration, metadata, API payloads, and records that are not flat. JSONL puts one complete JSON record on each line, so large collections of independent records such as logs, events, and model outputs can be appended and processed one record at a time.

Code becomes reusable once it is divided along its responsibilities — reading, validating, transforming, analyzing, writing — and those functions are gathered into modules that other programs import. A script needs one explicit entry point: a main() function behind an if __name__ == "__main__": guard, so the same file can be imported without side effects or run directly from the terminal. What these conventions share is that they make a program’s inputs, assumptions, transformations, outputs, and execution steps visible rather than implicit, which is what allows an analysis to be repeated and checked by someone other than its author.

Figure: Infographic about Files and Modules

Listen

Overview
Deep Dive

Read

Hands-on

Notebooks in 05-Files-Modules-Scripts