Functions and Decomposition
Decomposition is how a problem too large to hold in your head becomes a handful of steps you can name, write, and check one at a time. In Python those steps are functions: each one a contract with a single responsibility and clear expectations about its inputs, its return value, and its edge cases.
The same discipline is what lets the pieces fit back together. Returning a value instead of printing keeps a function usable by a test or by the next stage of a pipeline; local names and no shared global state keep one piece from quietly breaking another. Docstrings, type hints, and a little input validation record what each unit promises — so a program assembled from small, isolated parts stays something you can read, test, and change.

Listen
| Speaker | Text |
|---|---|
| Alex | This is the brief on Python function design. We’re looking at class notes for IFI 8410 about why Python functions are built the way they are. Basically, how they keep your data science projects from turning into an unpredictable, untrustworthy mess. Now, I’m just covering the high-level concepts right now, so you absolutely need to listen to the full podcast and review the Jupiter notebooks for this session to get the complete coding examples. First, functions aren’t just handy tricks to avoid copy pasting code, they are strict contracts. They set explicit boundaries for what data goes in, what comes out, and all your assumptions. I mean, if you hire someone to calculate a discount, you’d be pretty mad if they also secretly reorganize your filing cabinet, right? Well, your function needs to do exactly what its name promises and literally nothing else. Second, let’s talk return versus print. Print just shows data to us humans, but return actually hands data back to the program for testing. You might be wondering, why do hidden global variables matter if my code works right now? Look, hidden global states create massive race conditions during parallel execution. Your results will literally depend on random timing instead of actual data. Finally, decomposition. You’ve got to break tasks down during the design phase, not after your code is already a tangled mess. Think of it like building blocks. How do you know your pipeline won’t crash on a missing value? Because you built small, testable blocks where raising a value error or type error is an intentional, predictable choice. When your function’s name, rules and tests all tell the exact same story, your code becomes incredibly easy to reuse, parallellyze, and absolutely trust. |
| Speaker | Text |
|---|---|
| Alex | Why do so many massive, like, really critical projects just end up completely stalling out? Oh, it happens all the time, right? I mean, whether it is a multimillion dollar corporate initiative or, you know, just that garage organization project you’ve been avoiding for six months, |
| Sam | yes, the garage, we all have that garage project, |
| Alex | exactly. But the failure point is almost never a lack of talent. Or a lack of resources. The issue is that the human brain is literally incapable of processing hundreds of interdependent decisions at the exact same time. It just can’t do it. You stare at the mountain of work, your working memory hits its absolute limit, and you just freeze. |
| Sam | Yeah, I mean, it is a very real cognitive overload. |
| Alex | Welcome to today’s deep dive, everyone. I’m your host, and today we are looking at the ultimate antidote to that exact paralysis. |
| Sam | I’m so glad we are digging into this today. I’m your resident expert for this deep dive, and the concept we are unpacking is called decomposition. |
| Alex | Decomposition. And just a quick note for you listening, this entire conversation is based strictly on the. Materials we’ve gathered specifically focusing on this core document about decomposition, |
| Sam | which is such a fascinating document. |
| Alex | It really is. OK, let’s unpack this because we aren’t just talking about making a giant to do list, right? What is the actual mechanism at play here? |
| Sam | No, it’s definitely not just a to do list. At its core, decomposition is, well, it’s a deliberate process of taking a massive, imposing goal. And dividing it into smaller understandable tasks and steps, right? |
| Alex | Understandable being the key word, |
| Sam | exactly. But the crucial shift in perspective, and the sources are very clear on this, is that decomposition is really about shifting the cognitive load, shifting the load, meaning what exactly? Meaning it doesn’t make the ultimate goal any smaller or, you know, any less important. It just changes how your brain interacts with it. Ah, |
| Alex | I see. So it changes where the heavy lifting is happening. |
| Sam | Yes, precisely. You are moving the burden of organizing all that chaos out of your fragile working memory and putting it into a reliable external framework. |
| Alex | I love that it makes the invisible work completely visible. |
| Sam | That is the perfect way to frame it. When you effectively decompose a problem, everyone involved can suddenly see the architecture of the project. They see the matrix, basically, yeah, they see what needs to happen in what sequence, who owns which piece, and like. How a single completed step unlocks the very next one, |
| Alex | which allows you to maintain singular focus on just one meaningful action at a time without losing sight of the final prize. |
| Sam | Exactly, it cures the paralysis. |
| Alex | OK, I understand the theory, but I have to push back on the reality of this for a second. If I am already completely overwhelmed by a massive project and your advice is to, you know, break it down, my immediate fear is that you are just asking me to create. A terrifyingly long list, |
| Sam | right? The endless scroll of doom. |
| Alex | Exactly. Like if I take a complex 5000 piece puzzle and just dump it on the table without looking at the picture on the box, I haven’t made it easier to |
| Sam | understand. You’ve just made a huge mess, |
| Alex | right? I’ve just created a pile of cardboard on my dining table. Isn’t breaking a project into 1000 tiny pieces just generating massive amounts of busy work? |
| Sam | That is the exact trap most people fall into. And honestly, It is why their projects still fail even after they try to plan them out, |
| Alex | because they just dumped the puzzle pieces. |
| Sam | Yes, the source material specifically addresses this danger, actually. True decomposition is not about generating a longer list of disjointed chores. So what is it then? It’s about designing a schematic to ensure you are building a useful plan and not just a chaotic list. The text outlines 4 fundamental pillars. 4 pillars, OK, yeah, these are essentially the filters your breakdown has to pass through. If your breakdown doesn’t meet these criteria, you were just making busy work. |
| Alex | So if a massive to do list is the wrong approach, these 4 killers act as our guardrails. What is the first filter we should be using? |
| Sam | The first pillar is clarity. Clarity |
| Alex | makes |
| Sam | sense, |
| Alex | yeah, because large goals are inherently vague, right? Like if your goal is improve customer service or organize the annual summit, those are aspirations, not actions. You can’t just do an annual summit on a Tuesday afternoon. |
| Sam | Exactly. By breaking them down, you are forcing those vague goals to become concrete realities. Clarity means every resulting task has a recognizable purpose. That demonstrably moves the work forward. |
| Alex | It forces you to distinguish what is actually completed from what is just theoretical, |
| Sam | right? And it also reveals the hidden traps. Oh, |
| Alex | you mean the dependencies. Like you can’t design the event invitations until you’ve actually booked the venue |
| Sam | because you physically need the address to print on the card. You literally cannot do step B until step A is finished, right, |
| Alex | which sounds obvious, but people mess that up constantly, |
| Sam | all the time, which brings us directly to the second pillar, which is |
| Alex | reuse, reuse like recycling tasks, |
| Sam | sort of. When you take on a large effort, you are going to encounter workflows that you will inevitably need again. A good decomposition preserves the cognitive effort you spent figuring out the problem the first time. Oh, |
| Alex | I see, because if you just spent like a full week figuring out the absolute optimal sequence for onboarding a new client, you really don’t want to start from absolute zero the next time a client signs a contract. |
| Sam | No, that would be a huge waste of time. But reuse doesn’t mean mindless robotic repetition either. |
| Alex | What does it mean then? |
| Sam | It means preserving a proven structural approach so you can adapt it to the nuances of the new situation. It saves massive amounts of time and mental energy. That is |
| Alex | huge for preventing burnout, |
| Sam | honestly. Oh, absolutely. So moving on, the third pillar is checking. |
| Alex | Wait, let me challenge that one for a second. Checking progress along the way sounds great in theory, but, um, if I’m building a bridge. I can’t really check if it holds the weight of a truck until the bridge is actually finished, right? The final test, yeah. So how does checking work practically when the project itself isn’t complete yet? |
| Sam | Well, you might not be able to drive a truck on a half-built bridge, but you absolutely can and honestly must test the tensile strength of the steel before you assemble it, and you check the concrete mixture before you pour the foundation. That is what this pillar is about. So checking the components, exactly. Breaking work into distinct manageable tasks makes it entirely possible to verify quality before the final deliverable. |
| Alex | Instead of getting to the very end of a massive project and discovering a fatal flaw when it is like. Astronomically expensive to fix. |
| Sam | Precisely. You review each critical result along the way. |
| Alex | That makes total sense. You are stopping the domino effect of a bad decision. You are essentially asking, does this next task actually have what it needs to begin successfully? |
| Sam | And if it doesn’t, you catch it early, which |
| Alex | saves everyone. OK. What is the 4th pillar? |
| Sam | The 4th and final pillar is collaboration. When you are dealing with large goals, you are almost always dealing with multiple people, right? Yeah, |
| Alex | very rarely is a massive project a solo mission. |
| Sam | Decomposing the work allows a team to share the low without every single member needing to manage the entire behemoth in their head at once, |
| Alex | which eliminates that awful ambiguity where two people are doing the exact same work. Oh, |
| Sam | the worst. Or even worse than that, nobody is doing the work because they both just assumed the other person was handling it. |
| Alex | The classic miscommunication, |
| Sam | right? Good decomposition creates clean handoffs. So those are the pillars clarity, reuse, checking, and collaboration. |
| Alex | OK. But to my earlier point about the pile of smashed puzzle pieces on the table. The source provides a very specific methodology for how to actually execute this breakdown, doesn’t it? |
| Sam | It does. It gives a very practical step by step process. |
| Alex | So practically speaking, how do I build this schematic without just getting completely lost in the weeds? |
| Sam | You have to start at the absolute end. |
| Alex | The end. |
| Sam | You define the final outcome in plain descriptive language. You. Not start by listing activities or tools or software. You |
| Alex | just describe what success looks and feels like. |
| Sam | Exactly. From there you identify the major high level tasks required to reach that outcome. So |
| Alex | working backwards from the finished product to identify like the main load bearing walls of the project, |
| Sam | yes, the load bearing walls, and then you divide those major tasks into smaller steps. But here is the critical limitation, what we can call the Goldilocks rule of detail. |
| Alex | The Goldilocks rule. I like that. What is it? |
| Sam | You only divide a task until the next clear action is obvious. |
| Alex | Ah, so you don’t just keep zooming in infinitely until you are listing out like open laptop, click mouse, type letter H. |
| Sam | No, because that is when a plan. Stops being a helpful map and starts becoming a terrible micromanaging distraction, right? It just becomes annoying. A task is too broad if your team looks at it and still doesn’t know where to start, like plan event, right? But it is too narrow when it devolves into microscopic actions that obscure the actual meaning of the work. So you have to find that sweet spot. Once you hit that sweet spot of clarity, then you order. Tasks you map the dependencies, you assign the owners, and you establish your checking points. |
| Alex | I really want to ground this in reality because looking at the sources, they provide three distinct examples of how this scales up from small to huge. |
| Sam | They do, and the progression is really helpful to see. |
| Alex | Let’s look at the first one, which is incredibly mundane but honestly proves the point perfectly. Hosting a Saturday dinner party for friends. |
| Sam | The dinner party. Everyone’s been there. |
| Alex | We don’t need to list out every time you chop a carrot, but intuitively we naturally use decomposition here, don’t we? |
| Sam | We do, but usually we do it poorly, which is why dinner parties can be so stressful. |
| Alex | Oh, absolutely. |
| Sam | If you just keep host dinner as a single block in your mind, you hit that working memory limit we talked about, and then chaos ensues. You buy the fish too late. You realize you need the oven at 2 different temperatures simultaneously or Uh, you find out a guest is gluten-free right as you are serving the pasta. |
| Alex | The absolute ultimate nightmare scenario, right? |
| Sam | But if we apply the methodology, we define the outcome first, a relaxed, delicious meal for our friends. |
| Alex | OK, so then the major tasks intuitively break down into planning the meal, preparing the environment, and preparing the food. |
| Sam | And notice how those four pillars map directly onto those major tasks. Clarity forces you to separate the planning phase from the cooking phase. You check dietary needs before you choose the menu, right? |
| Alex | And checking happens when you cross-reference your shopping list against the recipes before you leave for the store, |
| Sam | preventing that mid-cooking panic where you realize you forgot the garlic, which |
| Alex | I do every time. And collaboration is seamless with this model. One person can run to the grocery store while another person sets the dining room table. |
| Sam | Because the tasks are cleanly separated, they don’t step on each other’s toes at all. Plus, if the menu is a massive hit, you’ve preserved the cognitive effort. Oh |
| Alex | right, the reuse pillar. You have a template ready to deploy for the next holiday. |
| Sam | It seems basic, I know. But it is the exact same underlying mechanics used in complex project management, |
| Alex | which brings us to the second medium complexity example from the text, organizing a one day workshop for students. |
| Sam | And the stakes are definitely higher. |
| Alex | Yeah, we aren’t just in our kitchen anymore. We are dealing with physical venues, external attendees, learning materials, and really tight communication schedules. |
| Sam | The source outlines 5 major tasks here defining the workshop, arranging logistics, communicating with participants, preparing the session itself, and finally running and reviewing the event. OK, 5 distinct buckets. But what is absolutely critical to understand here is a causal mechanism between these tasks. The text describes it as a flow of completed results, |
| Alex | a flow of completed results, like a cascade effect. Yes, |
| Sam | the output of one task becomes the mandatory input for the next. OK, give me an example. Let’s say you are defining the workshop. You have to lock in the topic and the target audience. The completed result of that task dictates the logistics, right? |
| Alex | You can’t book an appropriate room until you know if you are expecting 20 people or 200 people. |
| Sam | Exactly. And you literally cannot execute the communication task like sending out the invitations. Until the logistics task of booking the room and confirming the date is fully finalized, because |
| Alex | what are you going to put on the invitation? Come somewhere at some time. |
| Sam | Exactly. If you tried to jump ahead and send save the dates without a venue, you are introducing massive risk and confusion. |
| Alex | The flow of results forces discipline. |
| Sam | It really does. |
| Alex | I want you, the listener, to think about a project that is currently stalled on your desk at work right now. Are you procrastinating because you are secretly trying to execute step 4 when the completed result of step 2 hasn’t actually been locked down yet? |
| Sam | That is such a good question to ask yourself. That ambiguity is almost always where paralysis lives. |
| Alex | It’s so true. |
| Sam | The checking pillar becomes your absolute best friend here too. In the workshop example, a key checking point is reviewing the actual presentation materials against the learning goals you set way back in step 10 wow, |
| Alex | yeah. If you skip that check, you might deliver a beautifully designed presentation that entirely fails to teach the students what they actually came to learn. |
| Sam | It looks pretty, but it’s totally useless. |
| Alex | OK, so the dinner party and the workshop are great, but they are still ultimately managed by a small centralized group of people. What happens when the conditions are just completely chaotic? The stakes are literally life and death, and there is no single authoritative manager. |
| Sam | That is where the source material takes us for the 3rd and final example. A community improving its ability to prepare for, respond to, and recover from flooding. |
| Alex | This is a staggering leap in complexity. It is massive. We are talking about integrating municipal funding, emergency responders, vulnerable citizens, and massive physical infrastructure all at the same time. If a mayor just stands at a podium and says we need to be better prepared for floods, I mean, that is a completely useless goal. It’s way too vast. |
| Sam | It provides zero actionable clarity. So how does the text decompose this massive multi-agent chaos? It outlines 5 major tasks again. |
| Alex | OK, |
| Sam | what are they? First, understand risks and needs. Second, set priorities and prepare the plan. Third, improve physical readiness. Fourth, improve communication and response readiness. And 5th. Support recovery and long-term improvement. OK, let’s |
| Alex | dig into the why of this specific order, because this isn’t just an arbitrary list of good ideas. Not at all. For instance, why must understanding risks and needs be entirely isolated from improving physical readiness? Like, why not just do them both? |
| Sam | Because of causality. If you merge them or, you know, skip the first step, you get catastrophic failure. How so? Well, understanding risk means mapping the specific topography of the flood zones. It means identifying which neighborhoods have the highest population of elderly residents who can’t easily evacuate. OK, |
| Alex | so gathering critical data, right? |
| Sam | If you jump straight into physical readiness, like just blindly buying sandbags and fortifying levees without that risk data, you might spend millions of dollars protecting an industrial park while the vulnerable residential zone completely floods. |
| Alex | Oh wow. So the clarity of that first step literally dictates the survival rate in the later steps. |
| Sam | Exactly. And the second step, setting priorities. Acts as the bridge. You take the risk data, you look at your available municipal funding, and you make the really hard choices about what gets funded first. |
| Alex | And notice how the collaboration pillar is the absolute glue here. |
| Sam | Oh, it has to be. |
| Alex | In a flood response, the fire department, the hospital system, the local news stations, and the Department of Public Works all have to execute simultaneously, |
| Sam | right? Everyone is moving at once. |
| Alex | If the plan isn’t rigorously decomposed with crystal clear handoffs, you will have emergency shelter. I opened by the city, but no communication plan to tell the public where they are, |
| Sam | which defeats the entire purpose of the shelter. |
| Alex | Exactly. If we connect this to the bigger picture, it feels like this flood plan isn’t a straight line like the dinner party or the workshop |
| Sam | was. No, it’s not a straight line at all. A disaster preparedness plan is a continuous feedback loop, |
| Alex | a loop. OK, walk me through that. |
| Sam | It is entirely cyclical. Complex, large scale efforts operate across different time horizons. But they must remain connected. The risk data feeds the priorities. The priorities dictate the physical readiness. |
| Alex | And then inevitably a flood event actually happens or a full scale practice drill is run, |
| Sam | right? And the results of that event become the ultimate checking pillar. You document what failed. Maybe the warning sirens weren’t loud enough in the South District, |
| Alex | and that recovery data feeds directly back into understanding your new risks, starting the entire loop over again, |
| Sam | precisely. The RS pillar is also literally saving lives here. When the floodwaters are rapidly rising, you do not have the cognitive bandwidth to draft evacuation orders from scratch, |
| Alex | right? You reuse the communication templates you built during the planning phase months ago. Exactly. But you know, looking at this massive multi-agency municipal plan brings up a really concerning vulnerability, actually. What’s that? When you take a sprawling problem and you break it into hundreds of subtasks spread across dozens of disconnected teams, how do you prevent everyone from developing severe tunnel vision? |
| Sam | That is a very real threat, right? |
| Alex | Like if I’m the guy whose only job is to inspect drainage pipes. How do you stop me from losing sight of the fact that I’m part of a life saving disaster response? |
| Sam | It is the single greatest risk of decomposition, honestly. If you over fragment the work, people just become mindless cogs. Yeah, they just check a box. They execute their tiny task without understanding the broader context, which means if something unexpected happens, they can’t adapt. So how do we fix that? To prevent this fragmentation, the source introduces a vital safeguard. It’s a 4 question final review. Whenever you break a project down, you must run it through these 4 questions. |
| Alex | So these act as the final grounding mechanism to keep everyone tethered to the main objective. Yes, OK. What is the first check? Question |
| Sam | 1. Is every task demonstrably connected to the overall goal that |
| Alex | forces rigorous clarity. It really does, because if a team is spending 3 weeks redesigning the emergency department’s logo instead of updating the evacuation routes, you can’t draw a straight line from that task to the goal of flood survival. |
| Sam | Exactly. It’s vanity work and you just cut it. |
| Alex | Got it. OK. What about preserving our resources? |
| Sam | Question two asks, can a useful task or set of steps be used again? |
| Alex | OK, so this enforces the reuse pillar we discussed earlier. Yes, |
| Sam | it forces project managers to constantly look for workflows that can be templatized. If the public works department developed a really efficient way to inspect the levees, the city should capture that process so it doesn’t just leave when the current inspector retires. |
| Alex | That makes total sense. Institutional memory, |
| Sam | right. Now, question 3 focuses entirely on risk mitigation. Is there a clear point at which each important result can be reviewed? |
| Alex | This goes back to your bridge analogy. It ensures we aren’t just blindly executing tasks on a timeline, but actually building in mandatory stop signs to verify quality, |
| Sam | because you have to catch it before a minor miscalculation compounds into a disaster, |
| Alex | right? |
| Sam | And |
| Alex | the fourth question, |
| Sam | the final question ensures the human element doesn’t fail. Does everyone know who is responsible and critically what they need from others? |
| Alex | It’s the ultimate collaboration check. It completely eliminates the, I thought you were handling that excuse. |
| Sam | Yes, every single handoff is explicitly defined. As the text concludes, the true hallmark of successful decomposition is when a massive, intimidating problem becomes perfectly understandable without becoming disconnected, without losing the plot. Right? The individual tasks are small enough to overcome our cognitive paralysis and invite immediate action. But they remain woven together tightly enough to produce one unified result. |
| Alex | So what does this all mean for you listening to this right now? It means that whether you are trying to coordinate a meal for your in-laws, orchestrating a corporate seminar, or literally engineering a municipality’s defense. against a natural disaster, decomposition is the universal mechanism. It really scales |
| Sam | to anything. |
| Alex | It is the architectural tool for taking overwhelming anxiety, pulling it out of your working memory, and translating it into visible collaborative reality. |
| Sam | It is a profound shift in how we approach the external world, but you know, there’s a deeper implication here that is really worth reflecting on. Oh. |
| Alex | Where do you see this framework applying outside of traditional project management? |
| Sam | Well, we’ve spent this entire deep dive looking at external projects, right? Dinners, workshops, infrastructure. But the most paralyzing goals we face as humans are rarely external. They are internal. Think about the pursuit of deep personal growth or the ambition to master a complex new skill, or the terrifying prospect of navigating a complete career overhaul in your forties. |
| Alex | Those are incredibly vague, massive goals, and they trigger the exact same cognitive overload and. A paralysis we talked about at the very beginning of the show. They absolutely do. You sit on your couch, you think I need to change my life, and you are just frozen by the sheer weight of it. |
| Sam | Exactly. But what if we applied the rigorous architecture of a municipal flood plan to our own internal development? That is a wild thought. What if you define the precise outcome of that career change, identified the major structural pillars, and then divided those pillars down following the Goldilocks rule only until the very next action was crystal clear. |
| Alex | You would stop staring at the terrifying mountain of change my life, and you would just focus on the clarity of the single next step, |
| Sam | knowing that the overarching schematic was already holding the big picture together for you. |
| Alex | That is a truly powerful way to reframe our own potential, breaking down the complexity of our own lives, one visible action at a time. Thank you for joining us for this deep dive, everyone. Keep asking those questions, and we’ll see you next time. |
| Speaker | Text |
|---|---|
| Alex | Imagine you write a Python script, right? It, um, works flawlessly on Monday. Oh, |
| Sam | I know exactly where this is going. Yeah, |
| Alex | you feed it your data set. It cleans the columns, runs the math, and it just spits out this beautiful summary. The |
| Sam | perfect day for a data |
| Alex | analyst. Exactly. But then on Tuesday, you run the exact same script on a, you know, slightly updated file, and it silently deletes 10,000 rows of customer |
| Sam | data, |
| Alex | and |
| Sam | it doesn’t even. Throw an error. |
| Alex | Nothing. It just quietly ruins your analysis. And the reason why is because you didn’t write a contract, |
| Sam | which is, I mean, that’s the terrifying reality of moving from just learning Python syntax to actually doing data science in the real world, right, |
| Alex | because when you’re a beginner. Just getting the code to run without a syntax error feels like a massive victory. |
| Sam | Oh, a huge victory. But in a professional environment, you know, especially when you’re dealing with complex dependencies or inheriting code written by someone who left the company like 3 years ago, code that merely runs is dangerous. |
| Alex | Yeah, if you can’t absolutely trust the results your program computes, your entire analysis is basically worthless. Exactly, it’s a house of cards. So today in this deep dive, we’re tackling that exact transition. How do you build analytical code that you can actually trust? We’re looking at a set of class notes from a course, uh, IFI 8410. It’s titled Deep Dive Why Functions Are Designed This Way in Python. |
| Sam | And the core thesis here, it really flips how most people learn to code, yeah, |
| Alex | because usually, You’re taught that a function is just a shortcut, right? Like a way to avoid typing the same 10 lines of code over and over again, |
| Sam | a handy little macro, basically. But for a data scientist building reproducible, you know, auditible pipelines, functions are something else entirely. The foundational idea in these notes is that a function is a strict legally binding contract within your code base. |
| Alex | OK, let’s unpack this idea of a contract, because when I think of Python, I tend to think of a fairly loose, flexible language. |
| Sam | Sure, dynamic typing and all that. |
| Alex | Yeah. So what does a contract actually mean when we’re just talking about a handful of lines of code? |
| Sam | Well, a well-designed function contract answers four very explicit. Questions to guarantee predictability. |
| Alex | OK, what’s the first one? |
| Sam | First, what is its one specific job that says its name? Second, what raw materials does it demand? Those are its parameters. |
| Alex | Makes sense. |
| Sam | Third, what finished product does it guarantee to deliver? That is the return value. And finally, what happens if the rules are broken? |
| Alex | Ah, like if it gets bad data. |
| Sam | Exactly. Those are the preconditions and the edge cases. |
| Alex | So the notes actually use a function called calculated discounted total to make this concrete. |
| Sam | That’s a great example. |
| Alex | Yeah, if we look at its parameters, it doesn’t just ask for like some data, it explicitly demands a list of non-negative prices formatted as floats. And a discount fraction. |
| Sam | And if you try to hand it a negative price, it refuses to guess what you meant, right? |
| Alex | It immediately raises a value error. Or |
| Sam | if you hand it a string instead of a number, it raises a type error. It violently rejects bad materials, |
| Alex | violently rejects. I love that. Drawing a hard, unforgiving boundary. Yeah, |
| Sam | the function essentially says provide exactly what I asked for, and I guarantee a mathematically sound discounted total. Give me garbage, and I will crash your program loudly and immediately. |
| Alex | So you actually know there’s a problem. It’s like, um, it’s like hiring a contractor to build a 200 square foot deck in your backyard. OK, yeah, let’s run with that. The contract is hyper-specific. It states that you provide the building permit. And the pos those are your inputs. The contractor promises to deliver the wooden deck that is the return value. And the beauty is once that contract is signed, you don’t need to stand in the yard micromanaging how they swing their hammers, you know, |
| Sam | you really don’t care about their internal implementation |
| Alex | exactly. You just care about the inputs and the final output |
| Sam | and that separation between the caller, the homeowner, and the author, the contractor. That’s the entire point of modular design. |
| Alex | It frees up the function’s author. To go back in, say, 6 months later and completely rewrite the internal math |
| Sam | just to make it run 10 times faster and as long as they don’t change the terms of the contract. The inputs it takes and the outputs it gives, the rest of your data pipeline won’t even notice the internals changed. |
| Alex | So a strict contract guarantees what comes out. But I’ve definitely written code where the output was technically correct, but my pipeline still completely fell apart in the next step. And that brings us to a massive stumbling block. The source material highlights for beginners. The difference between showing an answer and actually delivering it, |
| Sam | right, the classic distinction between printing a value and returning a value, |
| Alex | yes, because when you’re doing quick exploratory data analysis, you know, in a Jupiter notebook, hitting print feels like you’re getting the job done because |
| Sam | you literally see the answer with your own eyes on the screen, right? |
| Alex | But structurally, it’s a huge trap. |
| Sam | Let’s frame this in a data science context. Say you write a function called clean data, and instead of returning your clean data frame, you just use print data frame.head. You run it, you see the 1st 5 rows of your beautiful clean data on the screen, and you think you’re good to go. |
| Alex | But then when you try to pass the result of that function into your next pipeline step, like to run a groupie operation. The whole thing throws a nun type error and dies |
| Sam | because I mean print is basically an illusion. |
| Alex | It really is. |
| Sam | It just sends pixels to your monitor for your human eyes to read, but it puts absolutely nothing into the computer’s memory registers for the next step to use |
| Alex | right in Python, if a function doesn’t execute an explicit return statement. It implicitly hands back a completely empty object called none. |
| Sam | So your downstream group B operation isn’t receiving a data frame at all. It’s receiving none. |
| Alex | Going back to our deck contractor, it’s like they build a gorgeous 3D model of your deck on an iPad. They show you the screen so you can see it, but they refuse to actually pour the concrete in your yard. Yeah, |
| Sam | you can look at it, but your patio furniture can’t stand on it. |
| Alex | That’s perfect. So the practical rule for analytical code is simple. You only print when you are at the absolute end of the line, |
| Sam | like communicating with an interactive human user. Otherwise, you return the data. Returning is the official interface for composition, |
| Alex | allowing step A to feed into step B. All right, so returning safely gives data out. But we also need to talk about how a function sneaks data in without telling anyone. Oh, |
| Sam | this is crucial, |
| Alex | because if I’m trying to trust my computations, hidden dependencies seem like an absolute nightmare. |
| Sam | You’re heading on scope and global variables, and in data science, where auditability is just everything, this is where messy code becomes dangerous code. |
| Alex | The notes contracts two very different approaches to this. The good example is a function called convert Celsius to Fahrenheit, a classic, yeah. Inside that function. The author creates a variable called multiplier and another called Fahrenheit, but the moment the function finishes its job and returns the final temperature, those variables are just wiped from the computer’s memory. They |
| Sam | are strictly local. It’s a totally sealed environment, |
| Alex | which means it’s safe to reuse 1000 times. It has no memory of the last phone it ran. |
| Sam | Exactly. Then the notes show the bad example, a function called total with tax. Right, |
| Alex | in that one, the functions contract only asks for a single input, the price. But internally to calculate the total, it secretly reaches outside of itself to |
| Sam | grab a global variable called tax rate. Let’s say |
| Alex | it’s set to 0.08, which is defined somewhere completely different in the script. |
| Sam | And on the surface, you call total with tax 100 and you get 108, it seems fine. |
| Alex | But if we connect this to the bigger picture of a real world data science project. You’ve just destroyed your ability to audit your own work |
| Sam | completely. Imagine you have a complex notebook in cell 3. Somebody tweaks that global tax rate variable to 0.10 because they were testing something. Oh, done that. We all have. And then down in cell 45, you make the exact same function call total with tax 100, and Suddenly it returns 110. The |
| Alex | exact same input produced a different output. And |
| Sam | looking purely at the function call, you have absolutely no idea why the math changed |
| Alex | because the dependency wasn’t in the contract. It was hidden, |
| Sam | right? In analytical code, preserving the provenance of a result is non-negotiable. If you look at a number in your final report, you must be able to trace exactly how it was calculated, |
| Alex | and you do that by forcing all dependencies to be explicit inputs. |
| Sam | That function shouldn’t be allowed to sneak out into the global scope to hunt for a tax rate. The contract must force you to pass tax rate through the front door as a parameter. |
| Alex | OK, so we’re building these sealed, isolated functions with strict contracts and explicit inputs. But like most data pipelines aren’t just one or two calculations. No, no, they’re massive. Yeah. So how do we apply this modular thinking to a massive messy project by |
| Sam | recognizing that decomposition isn’t something you do after your code becomes 1000 lines of unreadable spaghetti. Decomposition is the design phase. You architect your pipeline in discrete stages from |
| Alex | day |
| Sam | one. |
| Alex | The notes lay out a really clean pipeline architecture. Instead of one monstrous script, you break the analysis into explicit steps. First, load survey data. Then validated each, then cleanage column, then |
| Sam | summarized by region, and finally, export summary. |
| Alex | Exactly. Each block answers one very specific question. And because of the strict contracts, they become swappable components, which |
| Sam | is huge. If a year from now your data set outgrows pandas and you decide you need to migrate to polars for performance, you don’t have to rewrite the entire universe. |
| Alex | You just pull out cleanage column and slot in a new polar’s version, right? |
| Sam | And as long as the new function honors the same contract, taking in raw data and returning clean data. The rest of the pipeline functions normally. |
| Alex | Hang on though. I’m looking at this pipeline example and something scares me from a real world perspective. If I have a massive panda’s data frame with millions of rows and I pass that entire data frame into a function called leage column. Isn’t that just an open box? Oh, I see. Couldn’t that function accidentally overwrite or delete my original data set behind the scenes? |
| Sam | That’s a highly valid concern, and it exposes the danger of unannounced side effects, specifically mutation. We need to talk about why Python handles data the way it does. When you pass a simple integer into a function, Python passes a copy of that value. But when you pass a massive object like a list, a dictionary, or a data frame, Python doesn’t make a copy. Because |
| Alex | of memory, right? If I have a 5 gigabyte data frame, automatically duplicating it every time it passes into a new function would just crash my laptop’s RAM in seconds. Exactly. |
| Sam | To be memory efficient, Python passes a reference. Essentially a memory address pointing to your original data, but that efficiency comes with a terrifying risk. |
| Alex | The function now has direct live access to your original source data. |
| Sam | The notes show a dangerous example of this with a function called remove missing. It takes a list of values and uses slice assignment written as values bracket colon bracket to filter out. The missing items |
| Alex | and to the person calling the function, it looks like it just handed back a nice clean list, |
| Sam | but mechanically it reached into the computer’s memory and permanently mutated the original list in place. |
| Alex | Wow. So if the caller was relying on keeping that raw data untouched for an audit log, it’s just gone. |
| Sam | So how do you protect your data provenance when you’re working with something like pandas? The notes highlight the clean sales example. The very first thing that function does internally. Is execute data.copy. |
| Alex | It creates a completely independent quarantine duplicate of the data frame. Yes, |
| Sam | it drops the null values from the copy and returns the copy. |
| Alex | But calling dot copy on a massive data set carries a huge computational cost. It does. |
| Sam | It burns time and memory, but it guarantees your original data. Remains pristine. Now in the real world, you might hit a scenario where your data set is so enormous that you simply cannot afford the RAM to copy it. |
| Alex | You have to mutate it in place just to survive. |
| Sam | When that happens, the rule of the contract is that you must explicitly announce the mutation. |
| Alex | So you change the function’s name to something like remove missing place. |
| Sam | Exactly. The function’s name and its documentation must scream at the caller, warning. I am going to permanently alter the raw data you hand me. |
| Alex | There can be no surprise side effects. |
| Sam | The contract must reflect the reality of the operation. |
| Alex | All right, so we’ve got perfectly designed pipelines, explicit inputs. We’re protecting our memory, but let’s be honest, real world data is chaotic. |
| Sam | Very true. |
| Alex | It is messy and unpredictable. If our function is this strict legal contract. What mechanically happens when the raw data violates the terms? Well, |
| Sam | because Python is a dynamically typed language, it doesn’t stop you from passing the word apple into a math function until the script is already running. |
| Alex | Yeah, that flexibility is the double-edged sword of Python. A variable can. The integer right now and a string 5 milliseconds from now. |
| Sam | It’s fantastic for rapid prototyping, but it means the burden falls entirely on you, the function author, to decide what constitutes an invalid input and how the program should react. |
| Alex | Let’s look at the average function from the source material. It takes a total and a count. Sounds incredibly simple. Return total divided by count. |
| Sam | But what if the data is dirty and count is zero? Or what |
| Alex | if someone passes the word 25 as a string instead of the integer 25? |
| Sam | The author has to catch those edge cases explicitly, and the notes clarify an important distinction in how you structure your errors. If the pipeline hands you an inappropriate object. Entirely like the string instead of a number, you raise a type error, you’re signaling the fundamental nature of what you gave me is wrong. |
| Alex | But what if they give me the right type of object, but it breaks the mathematical rules of the contract, like a negative count of observations, |
| Sam | right? An integer is the correct data type, but a negative count is mathematically impossible in this context. Exactly. That’s when you raise a value error. You’re signaling the data type is fine. But the specific value violates our agreement. Of |
| Alex | course, if you’re processing a million rows of data, you probably don’t want the entire pipeline to abruptly crash because of one corrupted entry. |
| Sam | No. Depending on your organization’s data quality policies, you might choose to catch that bad value, return a missing indicator like Nan, and Log a warning to a file, |
| Alex | so the specific mechanism matters less than the fact that your handling of edge cases is deliberate, consistent, and explicitly stated in the contract. Exactly. So taking all of the, the contracts, the error handling, the avoidance of hidden global variables. What does this mean when a listener needs to scale up? Uh, scaling up. Let’s say they’re moving from processing a few 1000 rows on their laptop to running jobs on a distributed cluster using tools like desk or Spark UDFs. |
| Sam | When you scale up to parallel execution, every single bad habit we’ve discussed goes from being a minor annoying bug to a catastrophic system failure. |
| Alex | This is where the notes get into the concept of the race condition. And I want to really dig into the mechanics of why this happens. Sure, |
| Sam | let’s imagine you have a global variable called completed tasks that keeps a tally of finished chunks of data. If you have 50 different worker nodes running concurrently on a cluster, and they all try to read and update that single global counter at the exact same time, the final tally. Won’t depend on your data. |
| Alex | It will depend entirely on the microscopic nanosecond timing of the computer’s processors, |
| Sam | right? Because updating a variable isn’t a single magical action, right? It’s a multi-step physical process at the CPU level, |
| Alex | right? Worker A has to read the memory address, which says 0. But before worker A can add 1 and write the new total back to memory, worker B also reads that same memory address, and it still says 0. So worker A calculates 0 + 1 and saves 1. Worker B calculates 0 + 1 and saves 1. |
| Sam | You just processed two huge chunks of data, but your global counter only went up by 1. You’ve completely lost data, and worst of all, there was no error message. |
| Alex | That is a race |
| Sam | condition. It is a nightmare to debug because it is non-deterministic. It happens unpredictably. Based on processor load. This is why when you scale up to distributed tools like DSk or concurrent.futures, your contracts must take the form of what we call purer functions. |
| Alex | A pure function, meaning a function that exists in a total vacuum. It doesn’t look at anything outside itself. |
| Sam | Exactly. It’s the strictest possible contract, explicit inputs in, a return results out, absolutely zero global state, |
| Alex | no mutation of the caller’s objects, |
| Sam | and no reliance on the timing or order of execution. If you had a pure function the same input, it is mathematically guaranteed to give you the exact same output every single time, regardless of what is happening concurrently on 1000 other servers in your cluster. OK, |
| Alex | so we know what the perfect contract looks like, but to bring this all together. How do we actually enforce this over time? That’s the real challenge, |
| Sam | right? |
| Alex | If I write a flawless pure function today, how do I make sure a junior analyst doesn’t accidentally break the terms of the contract? Three months down the |
| Sam | line you enforce it through dock strings and automated tests. A dock string isn’t just a casual comment. It is the written formal record of the |
| Alex | contract. It’s since read at the top of the |
| Sam | function, and it tells the next reader, which, let’s be honest, is usually just you 3 months later having forgotten everything you wrote. Oh, completely. It tells them exactly what the function’s purpose is. What specific types it demands, what it promises to return, and what exceptions it will throw if the rules are broken. |
| Alex | But a doc string is just text on a screen. I mean, it can’t actually stop someone from feeding bad data into the function. |
| Sam | And that is where the testing framework comes in. Using a tool like Py Test, you convert your written doc string into executable evidence. Tests are how you prove the contract holds up under pressure. |
| Alex | The notes give a specific example of this, a test called Test of ridge rejects zero count. Love |
| Sam | that |
| Alex | test. Instead of just hoping the function works, the test aggressively attacks it. It explicitly passes a 0 into the count parameter of the average function and |
| Sam | then uses a feature called pitest.raises to intercept the |
| Alex | output. The test is literally asserting. If I feed this function as zero, it must throw a value error, |
| Sam | right? If it doesn’t throw the error, the test fails, and you know the contract is |
| Alex | broken. So in the name of your function, the parameter types, the dock string, and your suite of tests all align and tell the exact same story. |
| Sam | That is when you know you’ve built something real. That is when your analytical code graduates from just running. To being genuinely reproducible, scalable, and fully auditable. |
| Alex | We’ve covered a ton of ground today, from treating your functions like demanding construction contractors to avoiding the silent data loss of distributed race conditions. It’s been |
| Sam | a great deep |
| Alex | dive. Thank you for joining us on this exploration into the architecture of trustworthy code. But before you open up your IDE and go back to your pipelines, I wanna leave you with a final thought to chew on. Oh, let’s hear it. We talked entirely about how building a trusted data pipeline requires you to write incredibly strict defensive contracts to protect against bad inputs and hidden mutations. But think about your day to day workflow. Are you actually the one writing the contracts, or is the chaotic, messy, entirely unpredictable reality of the real world data dictating the terms you are forced to sign? Think about who is really in control the next time you type deaf. Catch you next time. |
Presentation
- Decomposition - Breaking a large problem into smaller pieces: tasks, steps, and the four payoffs of clarity, reuse, checking, and collaboration
- Functions as Decomposition - How the four payoffs of decomposition become Python function design
Read
- Decomposition: Breaking a Large Problem into Smaller Pieces (source document for podcast)
- Python Modules and Self-Standing Scripts (source document for podcast)
- Wes McKinney: Python for Data Analysis: Chapter 2
- Wes McKinney: Python for Data Analysis: Chapter 3
Hands-on
Notebooks in 04-Functions-Decomposition