Up to now, most of what I wrote ran as one long script. Functions are where code gets split into named pieces that can be called, reused and tested on their own.
I went in thinking functions were mostly about saving keystrokes. They are not a speed tool either. A call even adds a little overhead. What they buy is readability and one place to fix things: the main program becomes a short list of named steps, and a bug gets fixed once instead of in every copy. The syntax was the quick part. What took longer was the quieter behavior: what a function hands back, which name Python finds, and what a function shares with the code that called it.
Print is not return
A parameter is the name in the definition. An argument is the value passed in on a call. That vocabulary helped, but the bigger lesson was the difference between showing a value and handing it back:
def size_in_kb(num_bytes):
print(num_bytes / 1024, "KB")
total = size_in_kb(2048) + size_in_kb(1024)
It prints 2.0 KB and 1.0 KB, and then it fails:
TypeError: unsupported operand type(s) for +: 'NoneType' and 'NoneType'
A function with no return still returns something: None. print shows a value to a person. return gives it to the code that called the function, and only that can be added, compared or stored. Changing the body to return num_bytes / 1024 makes total equal 3.0.
What tripped me up: the printed output looked right, so I assumed the function was right. The error showed up on the addition, not where the mistake was. Now when I see NoneType in an error, I first check whether a function I called forgot to return.
Stub loudly, and mind the parentheses
I tended to write a lot of code before running any of it. Functions make it easier to write the outline first and fill in one piece at a time. A function that does not exist yet can be a stub, but pass quietly returns None to everything downstream. I prefer a stub that says it is unfinished:
def parse_timestamp(text):
raise NotImplementedError("parse_timestamp is not written yet")
A gap I can see beats a gap that looks like a real answer.
A related surprise: def binds a name to a function object, and the parentheses are what call it. print(report_title) shows something like <function report_title at 0x...> (the address varies) instead of the title, with no error. That same idea is why sorted(files, key=len) works: it passes len itself so sorted can call it on each item.
Copy, paste, and miss one name
Several of my function bugs came from copying a working function to make a similar one:
kb = 4
def kb_to_bytes(kb):
return kb * 1024
def mb_to_bytes(mb):
return kb * 1024 * 1024
kb_to_bytes(2) gives 2048. mb_to_bytes(2) gives 4194304 instead of 2097152. In fact it returns 4194304 for any argument, because the formula still multiplies the outer kb, which is 4, and never uses mb. I renamed the parameter and forgot the formula. Because a variable named kb happened to exist outside the function, Python used it instead of raising a NameError, and the answer just looked like a big number.
What tripped me up: the code ran, which is what made it hard to see. I now test each copied function on an input where I already know the answer before I trust it.
Which name does Python find?
A name assigned inside a function is local. Each call gets its own local names, and inside the function they hide any outer name with the same spelling. When Python looks up a name, it checks the local scope first, then the scopes of any enclosing functions (which only exist when one function is defined inside another), then the module's global names, and finally the built-ins such as len and list. That order is often summarized as LEGB, and the search stops at the first match.
That order is why reusing a built-in name causes trouble:
list = ["a.txt", "b.txt"]
names = list("xyz")
That fails with TypeError: 'list' object is not callable, because Python finds my list before it reaches the built-in. I stopped using list, str, sum and input as variable names.
Reading a global inside a function works. Assigning to it does not, unless I say so:
files_seen = 0
def note_file():
files_seen += 1
Calling that raises UnboundLocalError. Because the function assigns to files_seen, Python treats it as local for the whole function, so there is nothing to read yet. global files_seen fixes it, but I treat needing global as a hint that the value should be passed in and returned instead.
Two names, one list
This is the part I most needed to learn. Python does not copy an argument when it passes it. The parameter becomes a second name for the same object, and what happens next depends on whether the function changes that object or just points its own name somewhere else:
def add_tag(tags):
tags.append("reviewed")
def clear_tags(tags):
tags = []
my_tags = ["unread"]
add_tag(my_tags)
clear_tags(my_tags)
print(my_tags) # ['unread', 'reviewed']
append changed the one list both names point to, so the caller sees it. tags = [] only moved the function's local name to a new list. The caller's list was never touched.
When I do not want a function to change my data, I hand it a copy. Using the same add_tag:
original = ["unread"]
working = original[:]
add_tag(working)
print(original) # ['unread']
print(working) # ['unread', 'reviewed']
This felt familiar. Digital forensics has a similar habit: examine a working copy and preserve the original. The same habit protects a list from a function I did not write.
Defaults, and the list that remembers
Defaults let a caller skip arguments that usually stay the same, and keyword arguments make a call say what each value means:
def format_size(num_bytes, unit="KB", digits=2):
divisor = {"KB": 1024, "MB": 1024 ** 2}[unit]
return f"{num_bytes / divisor:.{digits}f} {unit}"
print(format_size(1536)) # 1.50 KB
print(format_size(3_500_000, unit="MB", digits=1)) # 3.3 MB
Then there was the trap I would not have guessed. A default value is evaluated once, when the def statement runs, not on each call. With a list, that matters:
def log_event(event, history=[]):
history.append(event)
return history
print(log_event("open")) # ['open']
print(log_event("save")) # ['open', 'save']
print(log_event("close")) # ['open', 'save', 'close']
Each call was supposed to start fresh. Instead, every call without its own list shares one default list, and it keeps growing.
The fix is to default to None and build the list inside:
def log_event(event, history=None):
if history is None:
history = []
history.append(event)
return history
Now the first two calls print ['open'] and then ['save'].
More than one answer, and logic a test can reach
For completeness: *args collects extra positional arguments into a tuple, and **kwargs collects extra keyword arguments into a dictionary. I rarely write them yet, but recognizing them makes other people's code easier to read.
A function can hand back more than one value as a tuple, and the last habit tied everything together: keep input and output under a guard so the logic can be imported and tested.
def pack(items, per_box):
boxes = int(items / per_box)
leftover = items - boxes * per_box
return boxes, leftover
if __name__ == "__main__":
count = int(input("How many items? "))
boxes, leftover = pack(count, 12)
print(f"{boxes} full boxes, {leftover} left over")
int() drops the fraction, so pack(50, 12) returns (4, 2): four full boxes, two left over. (int() truncates toward zero, while // rounds down, so the two differ for negative results that are not whole numbers.) Run directly, the file asks its question. Imported, it only defines pack, so a test can check pack(24, 12) == (2, 0) without anyone typing anything. The function's name, arguments and return shape become a promise to its callers. If I change what comes back, the tests should break.
The pattern behind my mistakes
None of my mistakes were about the idea of a function. They were about what crosses its boundary: a value printed instead of returned, a stale name from copy and paste, a shadowed built-in, a list I assumed was copied, and a default that was shared. None of them showed up on the line where the mistake was, and most produced something plausible first.
What I am applying
- Return values, print at the edges. Only the top-level code talks to the user.
- Stub loudly with
raise NotImplementedError, not a silentpass. - Test every copied function on an input where I already know the answer.
- Keep built-in names free, and pass values in and out instead of using
global. - Pass a copy when a function should not change my data, and use
Noneas the default for anything mutable. - Keep input and output under `if __name__ == "__main__":` so the logic can be tested directly.
What this insight does and does not support
- Observed. Learning Python functions, the concepts were manageable. My mistakes were at the boundary: a missing
return, a missed rename, a shadowed built-in, a shared list and a shared default. The examples are my own and were run on Python 3.13. - Provides context. Useful for anyone moving from single scripts to functions. It suggests paying the most attention to what goes into a function and what comes back out.
- Not established by this piece alone. These are one learner's notes, and I am still early in this. They are not a claim about any tool or product, not a statement about how any organization writes software, and not evidence that these habits alone make code correct.
Check your own functions the same way: call each one on an input where you know the answer, and look at what it hands back, not what it printed on the way.
Views are my own.