On this page
Iterators and generators in Python let you process data one item at a time instead of loading everything into memory. An iterator is any object that returns its next value on request; a generator is the easiest way to write one. Together they let you stream files larger than RAM without changing how your loops read.
The saving is dramatic. A list of a million squares measured 8,448,728 bytes with sys.getsizeof below; the equivalent generator measured 208. The catch is just as real: a generator runs once, and a second loop over it silently produces nothing. This guide is for Python users working with large files, logs or streams. It covers the iterator protocol, yield, generator pipelines, itertools, yield from, and when a plain list is still the better choice.
How does the iterator protocol work?
An iterator is any object with __iter__ and __next__ methods. The iterator types section of the Python documentation defines the protocol: __next__ returns the next item and raises StopIteration when there are no more. A for loop calls these methods for you:
nums = [1, 2, 3]
it = iter(nums)
print(next(it)) # 1
print(next(it)) # 2
print(next(it)) # 3
next(it) # raises StopIteration
A for loop is roughly this, with the StopIteration handled:
it = iter(nums)
while True:
try:
item = next(it)
except StopIteration:
break
print(item)
The distinction worth keeping straight is iterable versus iterator. A list is iterable: iter(nums) gives you a fresh iterator each time. An iterator is its own iterator, so iter(it) is it is True, and once it is used up it stays used up.
You rarely write iterator classes by hand. For comparison, here is a countdown written the long way:
class Countdown:
def __init__(self, n):
self.n = n
def __iter__(self):
return self
def __next__(self):
if self.n <= 0:
raise StopIteration
self.n -= 1
return self.n + 1
print(list(Countdown(3))) # [3, 2, 1]
A generator does the same job in a fraction of the code.
How do you write a generator with yield?
Put yield inside a function. Any function containing yield returns a generator:
def countdown(n):
while n > 0:
yield n
n -= 1
for x in countdown(3):
print(x) # prints 3, then 2, then 1
Calling countdown(3) runs none of the body. Instead, it returns a generator object. Execution starts on the first next() and pauses at each yield, preserving local variables. As the Python tutorial on generators puts it, the __iter__ and __next__ methods are created automatically, and StopIteration is raised when the function ends.
The memory difference is the point:
import sys
def squares_list(n):
return [i ** 2 for i in range(n)]
def squares_gen(n):
for i in range(n):
yield i ** 2
print(sys.getsizeof(squares_list(1_000_000))) # 8448728 (about 8 MB)
print(sys.getsizeof(squares_gen(1_000_000))) # 208
Those figures come from Python 3.14 on 64-bit Windows; other versions differ by a few bytes. Note also that getsizeof counts only the list's array of pointers, not the million integer objects it points to, so the real gap is larger still. The generator's size does not depend on n, because it stores a position rather than results.
Iterators and generators in Python for large files
Wrap the file loop in a generator that yields one cleaned line at a time. Reading big files is the canonical use, and the one you will hit first in data work:
def read_records(path):
with open(path, encoding="utf-8") as f:
for line in f: # file objects are already lazy
line = line.strip()
if line:
yield line
Memory stays flat regardless of file size. The with block stays open for as long as the generator is being consumed, which is what you want. However, it also means you should consume the generator promptly rather than storing it and using it much later.
A worked pipeline with real output
Generators compose into pipelines. Suppose data.csv contains a header, a blank line and some messy rows:
id,name,value
1,alpha,3.5
2,beta,-1
3,gamma,oops
4,delta,2.25
5,epsilon
Each stage below is a small generator:
def parse(lines):
for line in lines:
parts = line.split(",")
if len(parts) != 3:
continue # wrong number of fields
try:
value = float(parts[2])
except ValueError:
continue # header or bad number
yield {"id": parts[0], "name": parts[1], "value": value}
def filter_valid(records):
for r in records:
if r["value"] > 0:
yield r
valid = filter_valid(parse(read_records("data.csv")))
print(valid) # <generator object filter_valid at 0x...>
for r in valid: # nothing has been read until this loop runs
print(r)
Output:
{'id': '1', 'name': 'alpha', 'value': 3.5}
{'id': '4', 'name': 'delta', 'value': 2.25}
Nothing executes until the final loop, and only one record exists in memory at a time. As a result, each stage is independently testable: pass it a plain list of strings and check what comes out.
What is a generator expression?
A generator expression is a list comprehension written with parentheses, and it produces values lazily instead of building a list:
squares = (x ** 2 for x in range(1_000_000)) # lazy
squares_list = [x ** 2 for x in range(1_000_000)] # eager
They are especially good as an argument to an aggregating function, where the extra brackets are redundant:
total = sum(x ** 2 for x in range(1_000_000)) # no intermediate list
print(total) # 333332833333500000
names = ["Asha", "Ravi", "Meenakshi"]
print(max(len(name) for name in names)) # 9
print(any(n > 2 for n in [1, 5, 2])) # True
Because any and all short-circuit, combining them with a generator can stop reading early. That saves real time on large inputs. The Functional Programming HOWTO makes the same recommendation for infinite streams and very large data.
Why does a generator only work once?
Because a generator holds a position, not a collection. Once it reaches the end it stays exhausted:
gen = (x for x in range(3))
print(list(gen)) # [0, 1, 2]
print(list(gen)) # [] <- exhausted, and no error
This causes real confusion when a generator is passed to a function that iterates twice, for example one that first counts rows and then processes them. If you need multiple passes, materialise it:
data = list(read_records(path)) # now re-iterable, but uses memory
Alternatively, produce a fresh generator each time by calling the function again.
Which itertools functions are worth knowing?
The itertools module covers most streaming patterns:
from itertools import islice, chain, groupby, count, tee
# take the first n from any iterable, lazily
first_ten = list(islice(read_records(path), 10))
# concatenate iterables without building a list
print(list(chain([1, 2], (3, 4)))) # [1, 2, 3, 4]
# infinite sequences
for i, name in zip(count(1), ["Asha", "Ravi"]):
print(f"{i}. {name}") # 1. Asha / 2. Ravi
# group consecutive items (input must be sorted by the key)
fruit = sorted(["banana", "apple", "cherry", "avocado", "blueberry"])
for letter, group in groupby(fruit, key=lambda s: s[0]):
print(letter, list(group))
# a ['apple', 'avocado']
# b ['banana', 'blueberry']
# c ['cherry']
# two independent iterators from one
it_a, it_b = tee(read_records(path), 2)
groupby catches people out: it starts a new group every time the key changes, so unsorted input produces fragmented groups. Sort first.
Batching comes up constantly for API calls and database inserts:
from itertools import batched # Python 3.12+
print(list(batched(range(7), 3))) # [(0, 1, 2), (3, 4, 5), (6,)]
batched yields tuples, and the last one may be shorter. On Python 3.11 and earlier, the islice idiom does the same job:
def batched(iterable, n):
it = iter(iterable)
while chunk := list(islice(it, n)):
yield chunk
What does yield from do?
yield from delegates to another iterable:
def flatten(nested):
for item in nested:
if isinstance(item, list):
yield from flatten(item) # recurse
else:
yield item
print(list(flatten([1, [2, [3, 4]], 5]))) # [1, 2, 3, 4, 5]
For simple cases it is a shorter way to loop and yield each item. It also does more: according to the yield expressions reference, it forwards values sent with send() and exceptions passed with throw(), and it evaluates to the sub-generator's return value.
Iterator vs generator vs list: quick reference
| Iterator class | Generator function | Generator expression | List | |
|---|---|---|---|---|
| How you write it | __iter__ and __next__ |
def with yield |
(expr for x in data) |
[...] or list() |
| Memory | Current state only | Current state only | Current state only | Every element |
| Reusable | No, once exhausted | No, call the function again | No | Yes |
len() and indexing |
Only if you add them | No | No | Yes |
| Best for | Complex stateful iteration | Pipelines, files, streams | One-line transforms into sum, any, max |
Small data, many passes |
Common mistakes with generators
- Returning a generator from inside a
withblock. The file closes when the function returns, so the firstnext()raisesValueError: I/O operation on closed file.Useyieldinside thewithblock instead, as inread_recordsabove. - Calling
len()on a generator. It raisesTypeError: object of type 'generator' has no len(). Count while you iterate, or use a list. - Storing
groupbygroups for later. Each group is only valid untilgroupbyadvances. Collecting(key, group)pairs first and reading them afterwards gives empty groups, so calllist(group)inside the loop. - Using
teewhen one branch runs far ahead. The itertools documentation warns thatteemay need significant auxiliary storage; if one copy consumes everything before the other starts,list()is faster. - Iterating the same generator twice. The second pass silently yields nothing, as shown above.
When should you use a list instead?
Generators are not automatically better. Use a list when you need len(), indexing, multiple passes, or when the data is small.
# a generator adds nothing here
names = [row.name for row in small_table]
The rule of thumb: if the data might not fit in memory, or you only need one pass, generate. Otherwise use a list and keep the code simple.
The Sunday Growth Brief
One email a week: the best new comparisons, a fresh roadmap and the tech news worth your attention.
No spam. Unsubscribe in one click.
Related reading
Lambda functions pair naturally with generator expressions and key= arguments. User-defined functions covers function structure more broadly, and inheritance and composition in Python shows where iterator classes fit into larger designs. For how for loops consume iterators, see Python for loops and range().
Frequently asked questions
What is the difference between a list and a generator?
A list computes and stores every element up front. A generator computes each element on demand and keeps only its current state. For a million rows, the list holds a million objects in memory while the generator holds one position, which is why generators suit files and streams that may not fit in memory.
Why can I only loop over a generator once?
Because a generator holds a position, not a collection. Once it has yielded its final value it is exhausted, and iterating again yields nothing without raising an error. If you need multiple passes, either materialise it with list() when the data is small enough, or call the generator function again to get a fresh generator.
What does yield actually do?
yield suspends the function, hands a value back to the caller and preserves all local variables. When the caller asks for the next value with next() or a for loop, execution resumes on the line after the yield. When the function body finishes, Python raises StopIteration for you, which ends the loop.
When should I not use a generator?
Skip the generator when you need random access, len(), or more than one pass, or when the data is small enough that laziness buys nothing. A list is simpler to debug and print, and simpler is usually right. Reach for generators when the data is large, unbounded or arrives as a stream.
Written by
Althaf Ashraf
AI Systems Engineer, Tata Consultancy Services
AI systems engineer working on agentic decision systems and retrieval architectures at TCS, with a focus on getting AI into real workflows rather than demos.

