asyncio
Read the transcript
1. One loop, one coroutine at a time
Host: Let’s start with the thing people get wrong about asyncio before we even touch syntax: it does not make your code run in parallel. It’s a single-threaded scheduler, and today we’re going to build the mental model that explains basically every quirk, every gotcha, and every misuse you’ll ever hit with it.
Guest: Right, and the core of that model is really simple: one event loop runs one coroutine at a time, full stop. The only moment control can move to something else is at an await — that’s the suspension point, the place where a coroutine says ‘I’m waiting on something, go run whoever else is ready.’ Everything else asyncio offers is just bookkeeping around that one rule.
Host: So the speedup people feel with asyncio isn’t from doing more computation at once, it’s from not sitting idle while waiting on a network call — you overlap the waiting instead of the work itself. Which also means if your bottleneck is actually CPU-bound work, none of this helps at all, because there’s no waiting to overlap.
2. The toolkit: tasks, groups, timeouts, and backpressure
Host: Okay, so once you accept that it’s all one loop taking turns, what do you actually reach for to manage a bunch of those turns at once? Walk me through TaskGroup versus gather, because I see both in code and I’m never sure which one is the adult in the room.
Guest: TaskGroup, from 3.11, is the structured version — you open a block, spawn tasks into it, and when the block exits it’s waited everything and propagated the first failure while cancelling its siblings. Gather is older and looser: it runs things concurrently too, but unless you pass return_exceptions=True, one failure surfaces immediately while the others just keep running in the background, orphaned. On 3.10 or earlier you don’t have a choice, TaskGroup doesn’t exist, so you’re doing gather plus wait_for and being careful about cleanup yourself.
Host: And timeouts layer on top of that how? I’ve seen asyncio.timeout as a context manager and wait_for wrapping a single call, and I’ve never been sure they’re doing the same job.
Guest: wait_for wraps one awaitable and cancels it on expiry; asyncio.timeout, also 3.11+, is a context manager so it composes over a whole block, including a TaskGroup full of things. Underneath, both raise CancelledError at the await point, and since 3.8 that inherits from BaseException, not Exception, so a bare except Exception silently stops catching it. Then for actually limiting concurrency you want a Semaphore, used as async with so the slot releases on every exit path, and for producer-consumer handoff a Queue — bind maxsize or the default of zero means unbounded, and a slow consumer just quietly turns into an out-of-memory crash. And when the blocking thing genuinely can’t be awaited, run_in_executor or the newer to_thread hands it to a thread pool capped at the smaller of thirty-two or your CPU count plus four workers, so the loop keeps moving, but only up to that ceiling before calls start queuing.
3. Where it actually breaks
Host: So let’s do a rogues’ gallery. If someone drops a synchronous requests.get or time.sleep inside an async def, what actually happens versus what they think happens?
Guest: They think they’ve slowed down one request; they’ve actually frozen the loop for everyone, because nothing yields control back. Same story with create_task if you don’t hold the reference — the task can get garbage-collected mid-flight, silently, and you never even see the work finish. And the mirror-image bug is swallowing CancelledError instead of cleaning up in finally and re-raising it, which makes that task permanently uncancellable, so your shutdown or your timeout just hangs waiting on something that will never yield to it again.
Host: And the gather trap — bounding fan-out inside the loop, not outside it — plus return_exceptions=True quietly handing back failures nobody checks. Given all that, is there one thing you’d tell someone to turn on before they ship anything?
Guest: Set the environment variable Python asyncio debug to one, every time, it’s free. It flags coroutines that block the loop past a threshold and tasks that were created but never retrieved. Turn it on in dev, fix what it complains about, and you’ve caught those two classes of bug before production ever gets a chance to.
Not covered
The planner wanted these and found nothing in the source to support them:
- A live walkthrough of a full production gateway request path with rate limiting, retries, and circuit breaking
- A deep dive into distributed rate limiting or Redis-backed quota enforcement
- Interview-style coding exercises or grading rubrics for asyncio components
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
At a Glance
Section titled “At a Glance”A single-threaded cooperative scheduler. One event loop runs one coroutine at a time; every await
is a point where that coroutine may yield control back to the loop. Concurrency comes from
overlapping waiting, not from parallel execution — which is why asyncio wins decisively for
I/O-bound work and does nothing for CPU-bound work.
Reach for it when a process spends most of its time waiting on network calls and the libraries involved are async-native. Reach for processes instead when the work is computation.
Key Concepts
Section titled “Key Concepts”| Concept | What it actually is |
|---|---|
| Coroutine | What async def f() returns when called. Nothing runs until it is awaited or scheduled |
| Task | A coroutine handed to the loop to run independently, via asyncio.create_task() |
| Event loop | The scheduler. One per thread; asyncio.run() creates one, runs a coroutine, closes it |
await |
A suspension point. Control returns to the loop, which may run something else |
asyncio.TaskGroup |
Structured concurrency (3.11+). Scopes tasks to a block, propagates the first failure, cancels siblings |
asyncio.gather() |
Runs awaitables concurrently. Without return_exceptions=True, a failure surfaces immediately but siblings keep running |
asyncio.timeout() |
Context manager (3.11+) that cancels the block on expiry. Composes; wait_for wraps a single awaitable |
CancelledError |
Raised at the await point inside a cancelled task. Inherits from BaseException since 3.8 |
Semaphore |
Bounded concurrency. async with it, so the slot is released on every exit path |
Queue |
Producer/consumer handoff. Bound it — maxsize=0 means unbounded, which means unbounded memory |
run_in_executor |
Escape hatch for blocking calls: runs them on a thread pool so the loop keeps moving |
to_thread() |
The 3.9+ ergonomic form of the above, for a single blocking call |
Numbers That Matter
Section titled “Numbers That Matter”| Quantity | Value | Why it matters |
|---|---|---|
| Event loops per thread | 1 | Two loops in one thread is a bug, not a scaling strategy |
| Default thread-pool workers | min(32, cpu_count + 4) |
The ceiling on concurrent to_thread() calls before they queue |
Queue(maxsize=0) |
Unbounded | The default. A slow consumer becomes an OOM, silently |
asyncio.run() on an existing loop |
RuntimeError |
Why library code takes a coroutine rather than calling run() itself |
CancelledError base class |
BaseException (3.8+) |
except Exception no longer catches it — older code that relied on it is now subtly different |
TaskGroup |
Python 3.11+ | Along with asyncio.timeout(); on 3.10 you are back to gather and wait_for |
uvloop throughput gain |
Roughly 2× on I/O-heavy loops | A drop-in loop replacement, near-zero code change |
| GIL effect on threads | No CPU speedup for pure Python | Threads help blocking I/O only |
Common Gotchas
Section titled “Common Gotchas”- A coroutine that is never awaited does nothing and emits a
RuntimeWarning— often at interpreter shutdown, long after the call site scrolled away. time.sleep()insideasync deffreezes the whole loop, not just that request. Same forrequests, synchronous database drivers, and CPU-heavy parsing.create_task()without keeping a reference lets the task be garbage-collected mid-flight. Hold the reference, or use aTaskGroup.- Swallowing
CancelledErrormakes a task uncancellable: shutdown hangs and timeouts stop taking effect. Clean up infinally, then re-raise. gather()over caller-controlled input is unbounded fan-out. Put the bound inside the loop.- Exceptions in
gather(..., return_exceptions=True)are returned, not raised — code that does not inspect the results swallows every failure. - A
Semaphoreacquired outsidetryleaks a slot on any exception before thefinally.async withavoids the whole class of bug. - Async context managers need
__aenter__/__aexit__— a syncwithon an async resource fails at a confusing place. - Debug mode is free signal:
PYTHONASYNCIODEBUG=1reports coroutines that block the loop too long and tasks that were never retrieved.
Where to Go Deeper
Section titled “Where to Go Deeper”- Module 1: Production Python — the full treatment: layered timeouts, backpressure, cancellation as control flow, and graceful drain.
- Module 1 → Failure Modes — the production consequences of each gotcha above.
- Track: AI Systems Coding — the components built on these primitives that coding rounds ask for.
async-ai-gateway— bounded concurrency, retry with jitter, circuit breaking, and drain, as running code.