Python Async Explained: Do You Really Understand It?

Python GIL
In CPython (the standard Python interpreter), The GIL is a mutex around Python object access to simplify memory management and make the interpreter implementation safe.
The Global Interpreter Lock (GIL) lets only one thread execute Python bytecode at a time in a single process (which would run a single CPU).
This means threads won’t speed up CPU-heavy code; they still help for I/O-bound tasks (because threads release the GIL while waiting on I/O).
For CPU-bound parallelism, we use processes.
Note: So, in python even if you spin up multiple threads it likely that its being executed concurrently only one at a time
Coroutines
Coroutines are specialized functions that can pause and resume execution. Unlike threads or processes, which are managed by the Operating System, coroutines are managed by the application itself, making them "cooperative" rather than "pre-emptive."
Concurrency
Concurrency is the ability of a system to handle multiple tasks at the same time, by overlapping their execution.
Concurrency is about dealing with many things at once, not necessarily doing them simultaneously. Whereas parallelism deals with executing multiple tasks at the exact same time on multiple CPUs/cores. So, concurrency can exist without parallelism.
CPU Core 1: Task A: |----work----| IO |----work----| IO
CPU Core 1: Task B: | IO |----work----| IO |--work--|
Time ------------------------------------------------>
What is AsyncIO?
AsyncIO is Python's built-in library for writing concurrent code using the async/await syntax. AsyncIO is single-threaded and uses cooperative multitasking, where tasks voluntarily give up control.
Asynchronous doesn't mean faster instead asynchronous programming means doing other useful work instead of idly waiting for I/O operations (like network requests, database queries).
I/O-Bound vs. CPU-Bound Tasks:
I/O-bound tasks involve waiting for external operations. AsyncIO excels at these.
CPU-bound tasks require heavy computation. For these, multiprocessing is more suitable.
Event Loop
The event loop is the "engine" that runs and manages asynchronous functions. It's a scheduler that keeps track of tasks, suspending them when they await an I/O operation and resuming them when the operation completes. asyncio.run() starts and manages the event loop
Awaitables
These are objects that can be "awaited," meaning their execution can be paused and resumed later. Synchronous functions or libraries like time.sleep are not awaitable because they don't have the underlying mechanism to yield control to the event loop. Instead, asyncio.sleep should be used.
- The
awaitkeyword can only be used within a function defined with theasynckeyword.
Types of Awaitable Objects
Coroutines: Created when an
asyncfunction is called. They are functions whose execution can be paused.Tasks: Wrappers around coroutines that are scheduled on the event loop.
Futures: Low-level objects representing eventual results, similar to promises in JavaScript. While tasks use futures under the hood, developers rarely interact with futures directly.
Running Coroutines and Tasks
Coroutine are created by calling async functions. To run it and get a result, you must await it. Awaiting a coroutine directly schedules and runs it to completion at the same time.
Tasks are how coroutines are truly run concurrently. asyncio.create_task() schedules a coroutine to run on the event loop, allowing it to execute independently and yielding control back to the event loop when it encounters an await statement.
Coroutine Concurrency vs Thread Concurrency
Thread Concurrency: -
When you don’t have any library that can make our blocking code asynchronous then we launch it in a new thread.
The thread voluntarily yield control to the event loop
Switching happens at well-defined suspension points
Each thread will have its own stack but will share the same memory and GIL lock.
Can be achieved with
Taskobjects (asyncio.create_task) withasyncio.to_threadfor concurrent threads in python.
import asyncio
import time
def blocking_io(secs):
time.sleep(secs)
return secs
async def main():
start_time = time.time()
# Start three thread blocking operations concurrently
tasks = [asyncio.create_task(asyncio.to_thread(blocking_io, i)) for i in range(1,4)]
results = await asyncio.gather(*tasks,return_exceptions=True)
end_time = time.time()
print(results)
print(f"Completed in {end_time - start_time} seconds")
asyncio.run(main())
[1, 2, 3]
Completed in 3.0386593341827393 seconds
Coroutine Concurrency:-
Coroutines voluntarily yield control to the event loop
Switching happens at well-defined suspension points
Many coroutines can run on a single thread
Can be achieved with
Taskobjects (asyncio.create_task) for concurrent coroutines in python.
import asyncio
import time
async def blocking_io(secs):
await asyncio.sleep(secs)
return secs
async def main():
start_time = time.time()
# Start three coroutine/task blocking operations concurrently
tasks = [asyncio.create_task(blocking_io(i)) for i in range(1,4)]
results = await asyncio.gather(*tasks,return_exceptions=True)
end_time = time.time()
print(results)
print(f"Completed in {end_time - start_time} seconds")
asyncio.run(main())
[1, 2, 3]
Completed in 3.012946605682373 seconds
When to use what?
If you use Multi-processing
Each task runs in its own entirely separate Python instance with its own memory(heap, program code).
True Parallelism: Can use multiple CPU cores simultaneously.
Heavyweight: High memory overhead because you are spawning new runtimes.
Best for: CPU-bound tasks (complex math, image processing).
import asyncio
from concurrent.futures import ProcessPoolExecutor
import time
import os
def cpu_intensive_task(n):
"""Simulate a CPU intensive task like calculating fibonacci"""
task_start = time.time()
print(f"Process {os.getpid()} starting task with n={n}", flush=True)
def fibonacci(num):
if num <= 1:
return num
return fibonacci(num - 1) + fibonacci(num - 2)
result = fibonacci(n)
task_end = time.time()
task_duration = task_end - task_start
print(f"Process {os.getpid()} completed: fib({n}) = {result} in {task_duration:.2f} seconds", flush=True)
return (result, task_duration)
async def main():
start_time = time.time()
loop = asyncio.get_running_loop()
# Run multiple CPU-intensive tasks in parallel using ProcessPoolExecutor
with ProcessPoolExecutor(max_workers=4) as executor:
tasks = [
loop.run_in_executor(executor, cpu_intensive_task, n)
for n in [35, 36, 37, 38]
]
results = await asyncio.gather(*tasks)
end_time = time.time()
print(f"\nAll processes completed in {end_time - start_time:.2f} seconds")
print(f"Results: {results}")
return results
if __name__ == "__main__":
results = asyncio.run(main())
Process 30068 starting task with n=35
Process 51640 starting task with n=37
Process 7536 starting task with n=36
Process 16472 starting task with n=38
Process 30068 completed: fib(35) = 9227465 in 8.88 seconds
Process 7536 completed: fib(36) = 14930352 in 13.27 seconds
Process 51640 completed: fib(37) = 24157817 in 19.41 seconds
Process 16472 completed: fib(38) = 39088169 in 28.40 seconds
All processes completed in 30.38 seconds
Results: [(9227465, 8.884115219116211), (14930352, 13.268191814422607), (24157817, 19.40844225883484), (39088169, 28.399179935455322)]
The code above tackles a classic CPU intensive problem calculating Fibonacci numbers. The recursive implementation is intentionally inefficient (more on that later), which makes it perfect for demonstrating parallel processing:
The Gotcha: This recursive approach has exponential time complexity 𝑂 ( 2 𝑛 ), meaning fib(38) takes roughly twice as long as fib(37). This makes our CPU work really hard, which is exactly what we want for this demonstration.
Here's where the magic happens. Python's asyncio is fantastic for I/O-bound tasks (like web requests), but it can't help with CPU-bound work because of the GIL. The solution? Use ProcessPoolExecutor to spawn separate Python processes
Best practice is to set max_workers = os.cpu_count() to the number of cpu core.
Breaking it all down, the code creates a pool of 4 separate python processes, each with its own GIL loop.run_in_executor() bridges synchronous functions to asyncio's event loop and asyncio.gather() waits for all tasks concurrently and collects results in order as seen by the result of it running in 30.38 seconds even though the he highest time for the most cpu intensive was 28.40 seconds.
If you use Multi-threading
Tasks share the same memory space within a single process.
The GIL (Global Interpreter Lock): In Python, only one thread can execute bytecode at a time. You don't get true parallelism for CPU tasks.
Pre-emptive Switching: The OS decides when to switch threads. This can happen at the "wrong time," such as interrupting a thread right before it starts a network request, causing suboptimal delays.
Best for: Basic I/O tasks where the thread can "sleep" while waiting.
import asyncio
from concurrent.futures import ThreadPoolExecutor
import time
import os
def cpu_intensive_task(n):
"""Simulate a CPU-intensive task like calculating fibonacci"""
task_start = time.time()
print(f"Process {os.getpid()} starting task with n={n}", flush=True)
def fibonacci(num):
if num <= 1:
return num
return fibonacci(num - 1) + fibonacci(num - 2)
result = fibonacci(n)
task_end = time.time()
task_duration = task_end - task_start
print(f"Process {os.getpid()} completed: fib({n}) = {result} in {task_duration:.2f} seconds", flush=True)
return (result, task_duration)
async def main():
start_time = time.time()
loop = asyncio.get_running_loop()
# Run multiple CPU-intensive tasks using ThreadPoolExecutor (BAD for CPU tasks!)
with ThreadPoolExecutor(max_workers=4) as executor:
tasks = [
loop.run_in_executor(executor, cpu_intensive_task, n)
for n in [35, 36, 37, 38]
]
results = await asyncio.gather(*tasks)
end_time = time.time()
print(f"\n All threads completed in {end_time - start_time:.2f} seconds")
print(f"Results: {[r[0] for r in results]}")
print("Notice: Same process ID but different thread IDs, and MUCH slower!")
return results
if __name__ == "__main__":
results = asyncio.run(main())
Process 28256 starting task with n=35
Process 28256 starting task with n=36
Process 28256 starting task with n=37
Process 28256 starting task with n=38
Process 28256 completed: fib(35) = 9227465 in 15.40 seconds
Process 28256 completed: fib(36) = 14930352 in 34.90 seconds
Process 28256 completed: fib(37) = 24157817 in 53.01 seconds
Process 28256 completed: fib(38) = 39088169 in 55.10 seconds
All threads completed in 55.28 seconds
Results: [9227465, 14930352, 24157817, 39088169]
Notice: Same process ID but different thread IDs, and MUCH slower!
We see here that threading fails spectacularly for CPU-bound tasks notice how:
Same process ID (28256): All tasks run in the same process, confirming they are sharing memory space.
Sequential Execution: Despite using 4 threads, tasks complete one after another due to the GIL.
Performance Penalty: 55.28 seconds vs 30.38 seconds - threading is actually 82% slower than multiprocessing!
The culprit is python's Global Interpreter Lock(GIL). Think of it as a single key to the python interpreter, only one thread can hold it at a time. When your thread is doing CPU work (like calculating Fibonacci), it holds this key, forcing other threads to wait.
What happens under the hood with threading:
Thread 1: Calculate fib(35) [holds GIL]
Thread 2: Wait for GIL...
Thread 3: Wait for GIL...
Thread 4: Wait for GIL...
If you use Coroutines (Async/Await)
Tasks run within a single thread and a single process, but they cooperate.
No GIL issues: Since it’s only one thread, there’s no locking overhead.
Cooperative Switching: The code decides when to yield control (using
await). This ensures the CPU is never interrupted in the middle of critical logic, only when it's genuinely waiting for something else.Lightweight: You can run thousands of coroutines with the memory cost of just a few threads.
import asyncio
import aiohttp
import time
async def fetch_url(session, url, delay):
"""Simulate I/O-bound task - network request with delay"""
print(f"Starting request to {url}")
await asyncio.sleep(delay) # Simulate network latency
# In real world, this would be:
# async with session.get(url) as response:
# return await response.text()
return f"Response from {url} after {delay}s"
async def main():
start_time = time.time()
# Simulate multiple API calls
urls_and_delays = [
("https://api1.example.com", 2),
("https://api2.example.com", 1),
("https://api3.example.com", 3),
("https://api4.example.com", 1.5)
]
async with aiohttp.ClientSession() as session:
tasks = [
fetch_url(session, url, delay)
for url, delay in urls_and_delays
]
results = await asyncio.gather(*tasks)
end_time = time.time()
print(f"\nAll requests completed in {end_time - start_time:.2f} seconds")
print("Results:", results)
if __name__ == "__main__":
results = asyncio.run(main())
Starting request to https://api1.example.com
Starting request to https://api2.example.com
Starting request to https://api3.example.com
Starting request to https://api4.example.com
All requests completed in 3.02 seconds
Results: [
'Response from https://api1.example.com after 2s',
'Response from https://api2.example.com after 1s',
'Response from https://api3.example.com after 3s',
'Response from https://api4.example.com after 1.5s']
Summary of Execution Strategies
| Feature | Multi-processing | Multi-threading | Coroutines (Async) |
| Managed By | Operating System | Operating System | Application (Event Loop) |
| Multitasking Type | Pre-emptive | Pre-emptive | Cooperative |
| Memory Usage | High (Separate per process) | Medium (Shared) | Very Low (Shared) |
| Python GIL Impact | Bypasses it (Parallel) | Limited (1 Thread at a time) | Single-threaded (No lock needed) |
| Context Switching | Expensive (Kernel level) | Medium | Cheap (User level or Application level) |
⭐ Conclusion
Processes are for when you need more CPU power.
Threads are a middle-ground but suffer from Python's GIL and unpredictable OS switching.
Coroutines are the gold standard for high-performance I/O. They provide the most control, the lowest memory footprint, and the most predictable execution flow.
If you are building a web scraper, a chat server, or an API wrapper, Coroutines are likely your best tool.



