Building Distributed Systems in Elixir: Part 6 — Named Processes
A PID is tied to one running incarnation of a process. When that process crashes, its PID does not magically transfer to the replacement. Therefore, using a PID as a public service address creates a hidden failure: callers may continue sending requests to an identifier that no longer represents a living worker. The article shows this directly. The old worker had PID<0.102.0 , while its replacement had PID<0.105.0 . A message sent to the old PID cannot reach the replacement. The supervisor repairs the worker, but it does not preserve the worker’s address. This matters whenever other processes need a stable way to find a service. PIDs remain useful for direct replies, monitors, and links, as the article notes. For long-lived services, Elixir programs commonly use registered names or registries instead. Those mechanisms let clients locate the current process after supervision restarts it, though messages sent during a restart still need careful handling.
Why does replacing a crashed worker make its PID an unreliable public address?
A PID is tied to one running incarnation of a process. When that process crashes, its PID does not magically transfer to the replacement. Therefore, using a PID as a public service address creates a hidden failure: callers may continue sending requests to an identifier that no longer represents a living worker.
The article shows this directly. The old worker had #PID<0.102.0>, while its replacement had #PID<0.105.0>. A message sent to the old PID cannot reach the replacement. The supervisor repairs the worker, but it does not preserve the worker’s address.
This matters whenever other processes need a stable way to find a service. PIDs remain useful for direct replies, monitors, and links, as the article notes. For long-lived services, Elixir programs commonly use registered names or registries instead. Those mechanisms let clients locate the current process after supervision restarts it, though messages sent during a restart still need careful handling.
What is a PID in Elixir, and what exactly does it identify?
In Elixir, a PID, or process identifier, is the runtime’s identifier for a process. It lets another process send messages, create a monitor, or create a link. The important detail is its lifetime: the PID identifies one particular running incarnation, not an abstract worker role or permanent service.
The article gives a concrete pair of PIDs. The original worker was #PID<0.102.0>. After it crashed, the supervisor started a replacement with #PID<0.105.0>. Even though the replacement performs the same job, it is a different process and therefore has a different PID.
This distinction shapes system design. Keep PIDs when you need a reply path or failure signal for that specific process. Do not assume a stored PID will find a future replacement. For stable lookup, standard Elixir provides names and registries, which are broader service-discovery tools than a single process identifier.
What is a named process, and how is its name different from its PID?
A named process is an Elixir process registered under a name, commonly an atom such as :counter. Clients can use that name to find or send to the current process. The name represents a lookup handle, while the PID represents one specific running incarnation.
For example, a supervisor can start a worker and register it as :counter. If that worker crashes, the replacement can register itself under :counter again. Clients that look up the name can then discover the new PID instead of retaining the obsolete PID shown in the article.
The name is not permanent by itself. Elixir removes a local registration when its process exits, and the replacement must register the name again. During that gap, lookup or sending may fail. Names also have scope and uniqueness rules, so larger systems may use a Registry or distributed naming mechanism rather than one simple local name.
What happens to messages sent to a service name when the process behind that name crashes and is restarted?
A registered name points to the process currently holding that registration. It is not a mailbox or durable message queue. When the process crashes, its registration disappears. A supervisor may start another process, but that new process becomes reachable through the name only after it registers the name again.
Suppose :worker names a service. A caller sends a message while the original worker is alive, and that message targets the original process. If the worker dies, messages already in its mailbox are lost with it. A send attempted while :worker is unregistered cannot be delivered. Once the replacement registers :worker, later messages target the replacement.
This behavior makes naming useful but not sufficient for reliable requests. Callers should handle lookup or send failures, and services may need retries, acknowledgments, or durable storage. The article’s PID example explains why the name must be resolved again after supervision creates a new incarnation.
How many processes can usually hold the same registered name at one time on a single Elixir node?
On a single Elixir node, a local registered name normally belongs to one process at a time. This rule makes lookup unambiguous: resolving :worker identifies one current process, not a competing set of processes. If another process tries to claim the same name, registration fails or is rejected.
For example, one supervisor may run a service registered as :cache. A second cache process cannot also own :cache on that node. After the first process exits, the name becomes available, and a replacement may claim it. The replacement still has a new PID, even though it uses the same name.
This one-name rule is useful for singleton services, but it is not a general scaling strategy. If several workers should share work, use distinct names, a Registry with richer keys, or a pool and routing layer. Across multiple nodes, distributed naming introduces additional coordination and failure concerns.
What alternatives exist for finding a service, such as a local name, a registry, or a name shared across distributed nodes?
A local registered name is the simplest option. It gives one process a name on one node, which suits a local singleton service. It is easy to use, but it does not automatically provide a cluster-wide view or support many processes sharing one key.
Elixir’s Registry supports richer local discovery. It can associate keys with processes and can support multiple processes under a key, depending on its configuration. For names shared across distributed nodes, Erlang’s :global provides cluster-wide registration. Other distributed systems may use a dedicated service registry or application-level routing instead.
These choices do not erase failure windows. A crashed process must disappear from discovery, and a replacement must become discoverable. Distributed naming also requires coordination when nodes disconnect or disagree. The article’s central lesson remains: choose a service-discovery handle instead of assuming one PID will survive replacement.
How do isolated processes, message passing, and supervision combine to make distributed systems tolerate failures?
Elixir processes are isolated. Each has its own state and mailbox, so one process’s crash does not directly corrupt another process’s state. Processes communicate by sending messages rather than sharing mutable memory. This separation limits damage and gives failures a clear boundary.
Supervision adds recovery. A supervisor watches workers and starts a replacement when one crashes, as the article’s tiny supervisor does. The replacement has a new PID, such as #PID<0.105.0> instead of #PID<0.102.0>. Message passing lets processes cooperate, while supervision restores failed components.
The combination supports fault-tolerant designs, but it is not automatic reliability. Callers must account for dropped messages, changing PIDs, restart gaps, and service discovery. Stable names or registries can locate replacements, while retries, acknowledgments, and careful state management handle unfinished work. Distributed nodes add network failures, so the same principles must be applied with explicit failure assumptions.
This brief was written by AI from the original reporting and checked by other models. Names, figures and quotes come from the source; read it for full context.
Read more in the JupiteX app
Pulse is free. New stories every 4 hours, each one broken into the questions that explain it.
Or read more news on the web