๐ง From Processes to Promises: The Complete Guide to Concurrency, Blocking, and Why Your Code Sleeps
A journey from the kernel's process table to the event loop, written for the endlessly curious student.
- The Process โ The container of everything
- Infinite Loops โ Why servers never die (and why that's OK)
- The Fundamental Truth โ I/O always blocks
- Signals โ How the OS taps you on the shoulder (or shoots you)
- Threads โ The OS's original concurrency tool
- The Kernel's Waiting Room โ epoll, kqueue, IOCP
- Coroutines & Goroutines โ Cooperative multitasking
- Promises & Async/Await โ Sugar over state machines
- libuv โ The hidden loop inside Node.js
- WebSockets in Practice โ JS vs PHP vs Swoole
- Frameworks & Runtimes โ Octane, Swoole, OpenSwoole
- The Grand Comparison โ Every language, side by side
- The Unifying Truth โ Blocking is everywhere
- Conclusion โ What to take away
Every program that runs on your computer is a process. A process is not your code โ it's the container the OS creates to hold:
- ๐ Your code (in memory)
- ๐งฎ A stack (local variables, return addresses)
- ๐ฆ A heap (dynamically allocated memory)
- ๐ Open file descriptors (sockets, files, pipes)
- ๐ท๏ธ A PID (Process ID)
- ๐ A state (running, sleeping, stopped...)
A PID is a unique integer the kernel assigns to every process. Key facts:
- PID 1 is always
initorsystemdโ the ancestor of all processes. - PIDs are assigned sequentially and recycled after reaching
pid_max(typically 32,768 or 4,194,304). - Every process has a PPID (Parent PID). If a parent dies, children are reparented to PID 1.
- Your shell has a PID. Type
echo $$to see it.
A PID is not "the process" โ it's just a handle the kernel uses to refer to it.
At any instant, a process is in exactly one of these states:
| State | Symbol | Meaning | CPU Usage |
|---|---|---|---|
| Running | R | Executing or ready to execute | Actively using CPU |
| Sleeping | S | Waiting for I/O or event | 0% |
| Uninterruptible Sleep | D | Waiting for disk I/O | 0% |
| Zombie | Z | Finished, but parent hasn't reaped it | 0% |
| Stopped | T | Paused (e.g., Ctrl+Z) | 0% |
You can see these in ps aux or top. A healthy server spends 99.9% of its time in S โ sleeping, waiting, doing nothing. That's not a bug. That's the entire design.
Here's a truth that surprises many beginners: a process ends when its code ends. When your script reaches the last line, the OS frees its RAM, closes its file descriptors, and removes its PID. Done.
So how does a server stay alive? With an infinite loop:
while (true) {
$data = fread($socket, 1024);
process($data);
}But wait โ doesn't that burn CPU? This is where the crucial distinction lives:
while (true) {
$x = 1 + 1; // doing nothing useful
}This runs 100% CPU. The OS scheduler sees it as "runnable" and keeps giving it time slices. This does hurt the system.
while (true) {
$data = fread($socket, 1024); // BLOCKS until data arrives
process($data);
}When fread() finds no data, the kernel puts the process to sleep. The process enters state S and is removed from the CPU run queue entirely. It uses 0% CPU. When data arrives, the kernel wakes it up.
๐ก The key insight: A well-written infinite loop spends 99.9% of its time sleeping inside a blocking syscall, not spinning. The loop exists only to re-enter the blocking call.
Every long-running service on Earth uses this pattern:
| Service | Where it blocks |
|---|---|
| Nginx | epoll_wait() |
| Node.js | epoll_wait() (via libuv) |
| Swoole | epoll_wait() (Reactor threads) |
| sshd | accept() |
| Your PHP script | fread() on a socket |
| systemd | epoll_wait() |
The pattern is always the same: block โ wake on work โ process โ block again.
Here is the single most important fact in this entire article:
Blocking is not a language flaw. It is a hardware reality.
A CPU executes billions of instructions per second. But I/O devices operate on a completely different timescale:
| Operation | Latency | CPU cycles (at 3GHz) |
|---|---|---|
| L1 cache access | ~1 ns | 3 |
| RAM access | ~100 ns | 300 |
| SSD read | ~100 ยตs | 300,000 |
| Network round-trip (LAN) | ~500 ยตs | 1,500,000 |
| Network round-trip (internet) | ~50 ms | 150,000,000 |
| Disk seek (HDD) | ~10 ms | 30,000,000 |
When your code says "read from this socket," the data is not there yet. The CPU has exactly two choices:
- Spin โ execute a loop checking "is it here yet?" โ burns 100% CPU for millions of cycles doing nothing.
- Sleep โ tell the kernel "wake me when data arrives" โ 0% CPU, but the process is paused.
Blocking is the CPU's way of not wasting itself on waiting. It's the efficient answer to an unavoidable physical gap.
โ Busy-wait (spin):
CPU: is it ready? is it ready? is it ready? ... [1,000,000 times] ... yes!
CPU usage: 100% ๐ฑ
โ
Blocking (sleep):
CPU: kernel, wake me when ready. [SLEEPS] ... kernel: WAKE UP!
CPU usage: 0% ๐
There is no third option. The data physically cannot arrive faster than the network allows.
A signal is an asynchronous notification sent to a process by the kernel or another process. It's the OS's way of saying "hey, something happened."
When you press Ctrl+C in a terminal:
- The terminal driver sends
SIGINT(signal 2) to the foreground process group. - The process either:
- Dies immediately (default behavior), or
- Catches the signal and runs a handler (graceful shutdown).
This is why even an infinite loop dies when you press Ctrl+C โ the OS interrupts it. Your loop was sleeping in fread(), and the signal wakes it with a special error.
| Signal | Number | Default Action | Triggered By |
|---|---|---|---|
SIGINT |
2 | Terminate | Ctrl+C |
SIGTERM |
15 | Terminate (graceful) | kill PID (default) |
SIGKILL |
9 | Terminate immediately (cannot be caught) | kill -9 PID |
SIGHUP |
1 | Terminate | Terminal closes |
SIGSTOP |
19 | Pause (cannot be caught) | kill -STOP PID |
SIGCONT |
18 | Resume | kill -CONT PID |
SIGUSR1 |
10 | User-defined | kill -USR1 PID |
SIGKILL cannot be caught, blocked, or ignored. The kernel terminates the process immediately โ no cleanup, no file flushing, no graceful goodbyes. This is why you use it as a last resort.
ps aux # list all processes
top # live view, sorted by CPU
pgrep -a php # find processes named "php"
pstree -p # tree view with PIDs
kill 1234 # graceful terminate (SIGTERM)
kill -9 1234 # force kill (SIGKILL)
pkill php # kill all named "php"
killall php # same, by exact nameA thread is the smallest unit of execution the kernel can schedule. It consists of:
- A program counter (where in the code it is)
- A stack (local variables, return addresses) โ typically 1โ8 MB
- A set of CPU registers
- An entry in the kernel's scheduler
When the OS switches threads, it must:
- Save all registers of the current thread
- Save the stack pointer
- Load registers of the next thread
- Load its stack pointer
- Flush TLB entries
This costs 1โ10 microseconds. Sound fast? A CPU executes ~3,000โ30,000 instructions in that time. It's expensive.
Multiple threads inside one process, sharing the same memory space:
Process
โโโ Thread 1 โโโ
โโโ Thread 2 โโโค shared heap, shared globals
โโโ Thread 3 โโโค separate stacks
โโโ Thread 4 โโโ
Pros:
- โ True parallelism on multiple cores
- โ Shared memory = fast data exchange
Cons:
- โ Race conditions โ two threads writing the same variable
- โ Deadlocks โ two threads each waiting for the other's lock
- โ Memory cost โ 8 MB ร 1000 threads = 8 GB just for stacks
- โ Debugging is hard โ bugs are non-deterministic
- โ Limited scale โ a few thousand threads per machine before things degrade
Want 100,000 WebSocket connections with one thread each? That's 800 GB of stack memory. Impossible. This is why the industry moved to lighter-weight abstractions.
Imagine 10,000 open sockets. You want to know which one has data. The naive approach:
// BAD: old select() approach
for (int i = 0; i < 10000; i++) {
if (has_data(sockets[i])) { // check each one
read(sockets[i]);
}
}That's O(n) per check โ 10,000 syscalls just to find one ready socket. Terrible.
These are kernel mechanisms that let one thread watch many file descriptors at once:
int epfd = epoll_create1(0);
epoll_ctl(epfd, EPOLL_CTL_ADD, socket1, &event); // register once
epoll_ctl(epfd, EPOLL_CTL_ADD, socket2, &event);
// ... register all 10,000
struct epoll_event events[100];
int n = epoll_wait(epfd, events, 100, -1); // BLOCKS here
// returns ONLY the ready sockets, e.g. 3 of themThe key insight: epoll_wait blocks (thread sleeps at 0% CPU) until at least one registered fd has data. Then it returns only the ready ones. You registered 10,000 sockets, you get back 3.
| Platform | Mechanism | Since |
|---|---|---|
| Linux | epoll | 2002 |
| macOS / BSD | kqueue | 2000 |
| Windows | IOCP | 1994 |
| Solaris | event ports | 2000 |
| Cross-platform (older) | select / poll | 1980s |
This is a kernel feature. No language "has" epoll โ every language's runtime uses it on Linux.
A coroutine is a function that can pause itself and be resumed later. That's the whole idea.
| Thread | Coroutine | |
|---|---|---|
| Who pauses it | OS (preemptive) | Itself (cooperative) |
| When it pauses | Any instruction | Only at yield/await points |
| Stack size | 1โ8 MB | A few KB (or zero) |
| Managed by | Kernel | Language runtime |
| Context switch | ~1โ10 ยตs | ~10โ100 ns |
| How many | Thousands | Millions |
| Data races | Yes | Only if multi-threaded |
Cooperative scheduling is the key. A coroutine runs until it explicitly says "I'm waiting for something." No preemption means no data races between coroutines on the same thread.
- Stackful โ each coroutine has its own tiny stack. Go, Lua, Swoole.
- Stackless โ compiled into a state machine; no separate stack. JS async/await, Rust async, C++20.
A goroutine is Go's stackful coroutine with M:N scheduling:
Goroutines: G1 G2 G3 G4 G5 G6 G7 G8 ... G1000000
\ | / \ | /
OS threads: T1 T2 T3 T4
\ | /
CPU cores: Core 1 Core 2 Core 3 Core 4
When a goroutine blocks on I/O, the Go runtime parks it and runs another goroutine on the same OS thread. The OS thread doesn't block โ the runtime handles waiting via epoll/kqueue/IOCP.
The programming model is beautiful: you write code that looks blocking:
func handle(conn net.Conn) {
buf := make([]byte, 1024)
n, _ := conn.Read(buf) // looks blocking
process(buf[:n])
}
for {
conn, _ := listener.Accept()
go handle(conn) // spawn a goroutine per connection
}No async, no await, no callbacks. You can have millions of goroutines, each costing ~2 KB.
A Promise is not a mechanism. It's a value โ an object with three states and a list of callbacks:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Promise โ
โ state: "pending" โ
โ "fulfilled" (with a value) โ
โ "rejected" (with a reason) โ
โ callbacks: [fn1, fn2, fn3, ...] โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
It does no I/O. It doesn't know about sockets. It's just a container waiting for someone to call resolve or reject.
const promise = new Promise((resolve, reject) => {
const socket = connect(url);
socket.onData = (data) => resolve(data); // someone calls resolve
socket.onError = (err) => reject(err);
});The Promise sits in the gap while epoll_wait is blocking. When data arrives, the event loop calls resolve. The state changes to fulfilled, and all stored callbacks fire.
await is not a primitive. It's a keyword that the compiler turns into a state machine:
// You write:
async function getData() {
const res = await fetch(url);
const data = await res.json();
return data;
}The compiler transforms this into something like:
function getData() {
return new Promise((resolve) => {
let state = 0;
function step(value) {
switch (state) {
case 0: state = 1; fetch(url).then(step); return;
case 1: state = 2; value.json().then(step); return;
case 2: resolve(value); return;
}
}
step();
});
}Key facts about await:
- The function stops executing at
await. - Its locals are saved in a heap-allocated state machine.
- Control returns to the event loop, which runs other tasks.
- When the awaited Promise resolves, the runtime resumes the function.
| Promise | await | |
|---|---|---|
| What it is | A value (object) | A keyword (syntax) |
| When it exists | At runtime | At compile time |
| What it does | Holds a future result | Pauses the function |
| Can you have one without the other? | โ
Yes (Promises existed before await) |
โ No (await needs a Promise) |
await is sugar. Underneath, it's the same Promise mechanism.
- JS Promise: eager.
new Promise(() => doWork())runsdoWork()immediately. - Rust Future: lazy.
let f = async { do_work().await };does nothing until you.await.
This is why Rust's async feels so different โ the Future is a recipe, not a running task.
9. ๐ฉ libuv โ The Hidden Loop Inside Node.js
libuv is the C library that powers Node.js's event loop. It's the "hidden loop" we discussed earlier โ the thing that makes ws.onmessage feel effortless.
libuv picks the best I/O polling backend for each platform:
- Linux โ
epoll - macOS / BSD โ
kqueue - Windows โ
IOCP
When people say "Node.js is non-blocking," they mean something specific: the JavaScript thread never blocks. But something is still blocking:
Your JS code: ws.onmessage = handler
โ
Node's event loop: epoll_wait(...) โ BLOCKS HERE, 0% CPU
โ
Kernel: wakes when socket has data
โ
libuv: reads data, parses frame
โ
Your handler: runs on the JS thread
The epoll_wait() call blocks. The OS puts the Node process to sleep. "Non-blocking" is always a lie told at a higher level of abstraction.
There is no magic. To receive data continuously, something must repeatedly check "is there new data?" That something is always a loop.
๐จ JavaScript โ The Loop Is Hidden
const ws = new WebSocket('wss://stream.binance.com:9443/ws/btcusdt@trade');
ws.onmessage = (event) => console.log(event.data);You never see a loop. But underneath, V8 + libuv is doing something like:
while (true) {
events = epoll_wait(...); // blocks
for (event : events) {
data = read(event.socket);
frames = parse_websocket_frames(data);
for (frame : frames) {
dispatch_to_js_callback(frame); // calls your onmessage
}
}
}while (!feof($socket)) {
$header = fread($socket, 2); // blocks
$payload = fread($socket, $len);
echo decodeFrame($header . $payload)['payload'] . "\n";
}This is the same loop that JS hides. You just wrote it manually.
Co\run(function() {
$client = new Co\Http\Client('stream.binance.com', 9443, true);
$client->upgrade('/ws/btcusdt@trade');
while (true) {
$frame = $client->recv(); // coroutine-yields, doesn't block the process
echo $frame->data . "\n";
}
});Swoole's coroutine runtime uses epoll internally, pausing the coroutine while waiting and running others.
| Concept | JavaScript | PHP (raw) | PHP (Swoole) |
|---|---|---|---|
| The loop | Hidden in V8 | You write while |
Runtime handles it |
| Receiving data | ws.onmessage = fn |
fread() |
$client->recv() |
| Blocking? | Non-blocking | Blocking | Non-blocking |
| Who polls | libuv | Your loop | Swoole Reactor |
| Concurrent connections | Thousands | One per process | Thousands per process |
Octane boots your Laravel app once and keeps it in memory, then feeds it requests at high speed. It's a request-serving layer, not a non-blocking runtime.
The catch: Octane is only as non-blocking as the server you pair it with.
| Octane Server | Coroutine Support | Blocking I/O? |
|---|---|---|
| Swoole (standard) | Partial | |
| Swoole + Coroutine Hooks | โ Yes | โ True non-blocking |
| RoadRunner | โ No | |
| FrankenPHP | โ No |
Without coroutine hooks, Octane with Swoole is still blocking โ 8 workers with 1s I/O = 8 req/s. With coroutine hooks, those same 8 workers can handle thousands of concurrent requests.
Swoole is the original. OpenSwoole is a community fork. Both are C extensions providing:
- Coroutines (stackful, like Go)
- Multi-process + multi-threaded Reactor model
- Coroutine-aware I/O (
Co\MySQL,Co\Http\Client, etc.) - Runtime hooks that convert standard PHP blocking functions into coroutine-aware versions
The magic hooks:
OpenSwoole\Runtime::enableCoroutine(OPENSWOOLE_HOOK_ALL);
Co\run(function() {
sleep(5); // โ converted to non-blocking coroutine sleep!
// Process is NOT blocked. Other coroutines keep running.
});This works for sleep(), file_get_contents(), PDO, Redis, cURL, and more.
| Model | Processes | Workers | Total Time |
|---|---|---|---|
| Traditional PHP (blocking) | 3 | 3 | ~3 seconds |
| Swoole (coroutines) | 1 | 1 | ~1 second |
| PHP (FPM) | PHP (Swoole) | Go | JS (Node) | C (pthreads) | Java (threads) | Java 21 (virtual) | Rust (tokio) | |
|---|---|---|---|---|---|---|---|---|
| Unit | Process | Coroutine | Goroutine | Async task | Thread | Thread | Virtual thread | Async task |
| Scheduler | OS | Swoole | Go runtime | libuv | OS | OS | JVM | tokio |
| Preemptive? | โ | โ | โ | โ | โ | โ | โ | โ |
| Stackful? | N/A | โ | โ | โ | โ | โ | โ | โ |
| Stack size | ~8 MB | ~2 KB | ~2 KB | heap | ~8 MB | ~1 MB | ~few KB | heap |
| Max count | ~thousands | ~millions | ~millions | ~millions | ~thousands | ~thousands | ~millions | ~millions |
| True parallel? | โ | โ | โ | โ | โ | โ | โ | โ |
| Shared memory? | โ | โ | โ | โ | โ | โ | โ | โ * |
| Data races? | โ | Possible | Possible | โ | โ | โ | Possible | โ * |
| Runtime provided? | N/A | โ | โ | โ | โ | โ | โ | โ |
* Rust uses Send/Sync to guarantee no data races at compile time.
| Unit | Creation time | Memory |
|---|---|---|
| OS thread | ~10โ100 ยตs | 1โ8 MB |
| Goroutine | ~200 ns | 2 KB |
| JS async task | ~100 ns | ~few hundred bytes |
| Rust async task | ~50 ns | ~few hundred bytes |
| Swoole coroutine | ~1 ยตs | ~few KB |
| Switch type | Time |
|---|---|
| OS thread โ OS thread | 1โ10 ยตs |
| Goroutine โ Goroutine | 100โ200 ns |
| Async task โ Async task | 50โ100 ns |
The industry trend is clear: move away from OS threads toward runtime-managed coroutines.
Here is the mental model that ties everything together:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Your code โ
โ await fetch(url) yield value โ โ language syntax
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Language runtime โ
โ Promises, state machines, schedulers โ โ runtime
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ I/O library (libuv, tokio, Swoole, asyncio) โ โ library
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Kernel (epoll, kqueue, IOCP) โ โ kernel
โ epoll_wait() โ BLOCKS a thread โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโค
โ Hardware โ โ physics
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Every layer hides the blocking of the layer below it.
epoll_waitblocks the thread โ but only one thread.- The event loop doesn't block your code โ it runs other tasks while
epoll_waitsleeps. awaitdoesn't block the event loop โ it suspends your function and returns control.yielddoesn't block anything โ it hands control back to the caller or scheduler.
The blocking never disappears. It just moves down the stack. Every "non-blocking" language is blocking somewhere โ in a runtime, in a thread pool, or in the kernel.
Every language blocks. The only difference is who writes the
sleepand how many sleepers you can afford.
If you remember nothing else from this article, remember these ten truths:
- A process is a container โ code, memory, file descriptors, and a PID.
- Processes die when their code ends โ servers stay alive with infinite loops.
- Good loops block, bad loops spin โ one sleeps at 0% CPU, the other burns 100%.
- I/O blocking is physics, not a language flaw โ the CPU is orders of magnitude faster than I/O devices.
- Signals are how the OS interrupts โ Ctrl+C sends SIGINT;
kill -9sends SIGKILL. - Threads are preemptive and expensive โ a few thousand max, 1โ8 MB each.
- Coroutines are cooperative and cheap โ millions possible, a few KB each.
- epoll/kqueue/IOCP are kernel primitives โ one thread watches many fds, blocks at 0% CPU.
- Promises are values,
awaitis sugar โ both rely on the event loop and epoll underneath. - Non-blocking is always a lie at some layer โ blocking gets moved, never eliminated.
Imagine a restaurant:
- epoll = the kitchen's order bell system (one panel, lights up when any table's order is ready).
- Event loop = the waiter, standing by the panel, delivering orders as they light up.
- Promise = the ticket the kitchen hands back โ "your food will be ready."
- await = a customer saying "I'll order dessert after the main" โ the waiter serves others, comes back later.
- Coroutine = a chef who pauses plating one dish to start another, resuming seamlessly.
The bell is in the kitchen. The waiter is in the dining room. The customer's request is at the table. Three layers, all connected, none the same thing.
The world of concurrency is not about eliminating waiting. It's about waiting efficiently โ and every language, framework, and runtime is just a different answer to the same fundamental question:
"While one thing waits, what can the rest do?"
Written for the endlessly curious. Keep asking questions. ๐