14 - Async Python¶
Previous: 13 - Pydantic | Index: All guides | Next: 15 - pytest
Quick reference for asynchronous programming in Python with async / await and asyncio: running many slow I/O tasks (API calls, LLM requests, database queries) at the same time.
Last verified: 2026-09-27. For newer changes, check the Official docs links in the Introduction.
Introduction¶
Before you start¶
You should know: Python functions and exceptions (10 - Python Basics), and that an API call is a request over the network that takes time to answer (09 - HTTP and APIs). Async is an intermediate topic: it is fine to skip it until you need to make many slow calls at once.
The problem it solves: you need 100 answers from an LLM, and each takes 5 seconds. Called one after another, that is over 8 minutes, and your program sits idle almost the whole time, just waiting for the network. A web server handling many users has the same problem: one slow request should not block everyone else.
Before async: the usual answer was threads (several lines of work running "at once"), which work but are harder to reason about and heavier when you need thousands. Python added async / await (Python 3.5, 2015) as a lighter way for one program to juggle many waiting operations. FastAPI and the LLM SDKs support it directly.
Think of it like: the chef in the mental model below, who puts the pasta on to boil and starts the bread while waiting, instead of staring at the pot.
What is async programming?¶
Most time in web and AI apps is spent waiting: for an LLM to answer, a database to return rows, a file to download. Normal (synchronous) Python waits idle for each call to finish before starting the next. Async Python lets one program start many operations and switch to other work while each one waits. It is not about using more CPU cores; it is about not wasting time waiting.
Mental model¶
Imagine a chef cooking 3 dishes:
Synchronous chef (one thing at a time):
boil pasta [==========wait==========] -> bake bread [=====wait=====] -> make soup [====wait====]
total time = sum of all waits
Async chef (switches while waiting):
start pasta ---wait--------------------------> done
start bread ---wait-------------> done
start soup ---wait--------> done
total time = roughly the longest single wait
The event loop is the chef: a single loop that runs tasks until each hits an await (a waiting point), then switches to another task that is ready. async def defines a task that can pause; await marks where it pauses.
| Kind of work | Async helps? | Use instead |
|---|---|---|
| I/O-bound (HTTP, LLM APIs, DB, files, network) | Yes, a lot | asyncio |
| CPU-bound (maths, pandas, model inference) | No (blocks the loop) | multiprocessing, ProcessPoolExecutor |
| Blocking library with no async version | Only via threads | asyncio.to_thread(...) |
Why learn it?¶
- Parallel LLM calls: summarise 100 documents in the time of ~10 instead of 100 sequential calls.
- FastAPI is async; knowing when to use
async defvsdefavoids freezing your server. - Agents and streaming (tokens arriving over time) are naturally async.
- Most AI SDKs have async clients (
AsyncAnthropic,AsyncOpenAI,httpx.AsyncClient).
Key terms¶
| Term | Meaning |
|---|---|
| Coroutine | What calling an async def function returns; it does nothing until awaited |
await |
Pause here until the awaited thing finishes; let other tasks run |
| Event loop | The scheduler that runs coroutines and switches between them |
| Task | A coroutine scheduled to run concurrently (asyncio.create_task) |
| Concurrency | Many tasks in progress at once (interleaved) |
| Parallelism | Many tasks running at the exact same time on several cores |
| Blocking call | A normal function that waits without letting others run (time.sleep, requests.get) |
| Semaphore | A limit on how many tasks may run a section at once |
Where it fits: used with 09 - HTTP and APIs (httpx), 27 - LLM APIs (async clients), 40 - FastAPI, 32 - AI Agents; tested with 15 - pytest (pytest-asyncio).
Official docs¶
Where to read the latest, authoritative documentation:
| Resource | Link |
|---|---|
| asyncio documentation | https://docs.python.org/3/library/asyncio.html |
| HTTPX async support | https://www.python-httpx.org/async/ |
Contents¶
- async def and await
- Running Async Code
- Running Tasks Concurrently (gather)
- Tasks and TaskGroup
- Limiting Concurrency (Semaphore)
- Timeouts
- Error Handling
- Processing Results as They Finish
- Async HTTP with httpx
- Async LLM Calls
- Async Iterators and Streaming
- Async Context Managers
- Calling Blocking Code from Async
- Queues: Producer / Consumer
- Async in Jupyter and FastAPI
- Threads vs Processes vs Async
- Troubleshooting
- Try It
1. async def and await¶
The two keywords that define and use asynchronous functions.
async defcreates a coroutine function; inside it,awaitanother coroutine to wait for its result without blocking others.Use it in any function that does I/O through an async library.
import asyncio
async def fetch_price(item: str) -> float:
"""Pretend to call an API that takes 1 second."""
await asyncio.sleep(1) # non-blocking wait (NOT time.sleep)
return 9.99
async def main():
price = await fetch_price("lamp") # waits 1 s, but other tasks could run meanwhile
print(price)
- You can only use
awaitinsideasync def. - Calling
fetch_price("lamp")withoutawaitreturns a coroutine object and does nothing.
2. Running Async Code¶
Starting the event loop from normal code.
asyncio.run(main())creates a loop, runsmainto completion, and closes the loop.Use it for the single entry point of an async script.
Call asyncio.run once at the top level, not inside other async functions.
3. Running Tasks Concurrently (gather)¶
Starting several coroutines at once and waiting for all of them.
asyncio.gather(*coroutines)runs them concurrently and returns results in the same order.Use this when many independent API / LLM calls.
import time
async def main():
start = time.perf_counter()
prices = await asyncio.gather(
fetch_price("lamp"),
fetch_price("chair"),
fetch_price("desk"),
)
print(prices, f"{time.perf_counter() - start:.1f}s") # [9.99, 9.99, 9.99] 1.0s (not 3.0s)
items = ["a", "b", "c", "d"]
results = await asyncio.gather(*(fetch_price(i) for i in items)) # from a list
4. Tasks and TaskGroup¶
Scheduling coroutines to run in the background while you do other things.
asyncio.create_taskstarts a coroutine now;TaskGroup(Python 3.11+) starts several and waits for all, cancelling the rest if one fails.Use it for background work, or structured concurrency with safe error handling.
async def main():
task = asyncio.create_task(fetch_price("lamp")) # starts running right away
await do_other_things()
price = await task # collect the result later
async with asyncio.TaskGroup() as tg: # all tasks finish before the block exits
t1 = tg.create_task(fetch_price("a"))
t2 = tg.create_task(fetch_price("b"))
print(t1.result(), t2.result())
Keep a reference to tasks you create; unreferenced tasks can be garbage-collected.
5. Limiting Concurrency (Semaphore)¶
Allowing at most N tasks to run a section at the same time.
async with semaphore:waits for a free slot before entering.Use it for respecting API rate limits when launching hundreds of calls.
MAX_CONCURRENT = 5
async def summarize_all(docs: list[str]) -> list[str]:
"""Summarize many documents, at most 5 LLM calls at a time."""
sem = asyncio.Semaphore(MAX_CONCURRENT)
async def one(doc: str) -> str:
async with sem: # only 5 inside this block at once
return await summarize(doc)
return await asyncio.gather(*(one(d) for d in docs))
6. Timeouts¶
Giving up on something that takes too long.
asyncio.timeout(seconds)(3.11+) orasyncio.wait_for(coro, timeout)raisesTimeoutError.Use it in any external call that might hang.
async def main():
try:
async with asyncio.timeout(5):
result = await slow_call()
except TimeoutError:
result = None
result = await asyncio.wait_for(slow_call(), timeout=5) # older style
7. Error Handling¶
What happens when one of many concurrent tasks fails.
gatherraises the first error by default;return_exceptions=Truereturns errors as values instead;TaskGroupraises anExceptionGroup.Use it for batch jobs where one failure should not lose all other results.
results = await asyncio.gather(*coros, return_exceptions=True)
for item, res in zip(items, results):
if isinstance(res, Exception):
logger.warning("failed %s: %s", item, res)
else:
save(res)
try:
async with asyncio.TaskGroup() as tg:
...
except* ValueError as eg: # handle a group of ValueErrors (3.11+)
for e in eg.exceptions:
logger.error(e)
8. Processing Results as They Finish¶
Handling each result as soon as it is ready, instead of waiting for the slowest.
asyncio.as_completedyields awaitables in completion order.Use it for progress bars, streaming results to a UI, saving results early.
for next_done in asyncio.as_completed([fetch_price(i) for i in items]):
price = await next_done
print("got", price)
9. Async HTTP with httpx¶
Making many HTTP requests concurrently. One shared
httpx.AsyncClient(reuses connections) +gather.Use it for scraping, calling several APIs, webhooks. Sync
requestswould block the loop.
import httpx
async def fetch_all(urls: list[str]) -> list[int]:
async with httpx.AsyncClient(timeout=10) as client:
responses = await asyncio.gather(*(client.get(u) for u in urls), return_exceptions=True)
return [r.status_code if isinstance(r, httpx.Response) else -1 for r in responses]
10. Async LLM Calls¶
Calling an LLM for many inputs concurrently. Use the SDK's async client with
await, plus a semaphore for rate limits.Use it for classifying / summarising / extracting over many documents.
import anthropic
client = anthropic.AsyncAnthropic()
sem = asyncio.Semaphore(8)
async def classify(text: str) -> str:
async with sem:
msg = await client.messages.create(
model="claude-opus-5",
max_tokens=256,
messages=[{"role": "user", "content": f"Label as positive/negative/neutral. Reply with the label only.\n\n{text}"}],
)
return next(b.text for b in msg.content if b.type == "text").strip()
async def main(texts: list[str]):
labels = await asyncio.gather(*(classify(t) for t in texts))
For very large offline jobs, the provider's batch API is cheaper than many parallel calls (see 27 - LLM APIs).
11. Async Iterators and Streaming¶
Looping over values that arrive over time.
async forover an async iterator; write your own withasync def+yield(async generator).Use it for LLM token streams, reading a stream of events, paginated APIs.
async def stream_answer(question: str):
async with client.messages.stream(
model="claude-opus-5",
max_tokens=16000,
messages=[{"role": "user", "content": question}],
) as stream:
async for text in stream.text_stream:
print(text, end="", flush=True)
async def numbers(n: int): # async generator
for i in range(n):
await asyncio.sleep(0.1)
yield i
async for x in numbers(3):
print(x)
12. Async Context Managers¶
Setup / cleanup that itself needs
await(open connections, sessions).async withcalls__aenter__/__aexit__; build your own with@asynccontextmanager.Use it for HTTP clients, DB connections, FastAPI lifespan.
from contextlib import asynccontextmanager
@asynccontextmanager
async def db_connection():
conn = await connect()
try:
yield conn
finally:
await conn.close()
async with db_connection() as conn:
rows = await conn.fetch("SELECT 1")
13. Calling Blocking Code from Async¶
Using normal (sync) functions without freezing the event loop.
asyncio.to_thread(func, *args)runs it in a thread; CPU-heavy work goes to a process pool.Use this when a library has no async version (pandas, some SDKs, file parsing).
import asyncio
from concurrent.futures import ProcessPoolExecutor
df = await asyncio.to_thread(pd.read_csv, "big.csv") # I/O-ish blocking call
loop = asyncio.get_running_loop()
with ProcessPoolExecutor() as pool:
result = await loop.run_in_executor(pool, heavy_cpu_function, data)
14. Queues: Producer / Consumer¶
Tasks passing work to each other through a queue. Producers
putitems; a fixed number of workersgetand process them.Use it for pipelines (download -> parse -> embed), controlled worker pools.
async def worker(name: str, queue: asyncio.Queue):
while True:
item = await queue.get()
try:
await process(item)
finally:
queue.task_done()
async def main(items):
queue: asyncio.Queue = asyncio.Queue()
for item in items:
queue.put_nowait(item)
workers = [asyncio.create_task(worker(f"w{i}", queue)) for i in range(4)]
await queue.join() # wait until every item is processed
for w in workers:
w.cancel()
15. Async in Jupyter and FastAPI¶
Environments that already run an event loop. Jupyter and FastAPI own the loop; you just
awaitdirectly.Use it for notebooks and web endpoints.
# Jupyter: top-level await works directly (no asyncio.run)
results = await asyncio.gather(*(classify(t) for t in texts))
# FastAPI: async endpoint with async libraries
@app.get("/prices")
async def prices():
return await asyncio.gather(fetch_price("a"), fetch_price("b"))
In FastAPI use plain def for blocking code (FastAPI runs it in a thread for you). See 40 - FastAPI.
16. Threads vs Processes vs Async¶
The three ways to do several things at once in Python. They differ in what they share and what they speed up.
Use it for pick by the type of work.
Async (asyncio) |
Threads (threading, ThreadPoolExecutor) |
Processes (multiprocessing, ProcessPoolExecutor) |
|
|---|---|---|---|
| Good for | Lots of I/O with async libraries | I/O with blocking libraries | CPU-heavy work |
| Parallel CPU | No | No (GIL, except free-threaded builds) | Yes |
| Scale | 1000s of tasks | 10s to 100s | Number of cores |
| Overhead | Very low | Low | High (separate memory) |
| Example | 500 LLM calls | 20 downloads with requests |
Parallel feature engineering |
17. Troubleshooting¶
| Problem | Fix |
|---|---|
RuntimeWarning: coroutine '...' was never awaited |
Add await before the call |
SyntaxError: 'await' outside async function |
Put the code in an async def |
RuntimeError: asyncio.run() cannot be called from a running event loop |
In Jupyter / FastAPI just await; do not call asyncio.run |
| Async code is not faster | Using blocking calls (requests, time.sleep); switch to httpx, asyncio.sleep, async SDK clients |
| Whole app freezes | CPU-heavy or blocking code inside async def; use to_thread / process pool |
| 429 rate limit errors | Add a Semaphore, reduce concurrency, respect retry-after |
| One failure cancels everything | gather(..., return_exceptions=True) |
Task was destroyed but it is pending |
Program exited before tasks finished; await or cancel them properly |
Event loop is closed (Windows) |
Close clients inside the loop (async with), avoid reusing clients across asyncio.run calls |
18. Try It¶
Short exercises to practise this guide. Try each task yourself first, then open the solution.
Use it right after reading the guide, or later as a quick self-test.
Exercise 1: Concurrency proof¶
Run three tasks that each sleep 1 second so the total takes about 1 second.
Solution
Exercise 2: At most 5 at once¶
Call an async function for 100 items with at most 5 running at the same time.
Solution
Exercise 3: Timeout¶
Give slow_call() at most 2 seconds and fall back to None.
Solution
Previous: 13 - Pydantic | Index: All guides | Next: 15 - pytest