Connection Events
The last page already used three connection callbacks without naming them as a group: the disconnect callback logged the moment the link dropped, the reconnect callback logged the rejoin, and the reconnect-error callback recorded every failed attempt in between. Each one reports a transition of the same state machine.
This page collects that reporting into one surface. You'll see every event the connection can deliver, wire the one no earlier page covered — the closed event, the signal that the connection won't come back — and read the connection's state directly for a health check. Two new ideas carry the page: the connection event surface, and status polling.
The events a connection reports
A connection event is a notification the client delivers when the connection changes state or learns something about the cluster. Six events cover the lifecycle:
- Disconnect — the link is gone and the client has entered RECONNECTING. You watched this fire on Reconnection.
- Reconnect — the client is back on a server and the connection is CONNECTED again. Also from Reconnection.
- Reconnect error — one reconnect attempt failed; the client keeps trying. Reconnection wired this to make a long outage visible.
- Closed — the connection reached CLOSED and will not come back. No earlier page covers this event; this page does, below.
- Discovered servers — the client's server pool grew from cluster gossip, described in Connecting → Discovered servers.
- Lame duck — the server announced it's shutting down soon, so the client can move before the link is cut. It belongs to the next page, Drain & Shutdown.
Three of the first four mark edges of the state diagram from Reconnection; the reconnect error repeats on RECONNECTING, once per failed attempt, so the diagram shows it as an annotation on that state rather than an arrow:
CONNECTED ──disconnect──▶ RECONNECTING ──reconnect──▶ CONNECTED
│ │ reconnect error, per failed attempt
│ Close() by the app └──retries spent──▶ CLOSED
└─────────────────────────────────────────────▶ CLOSED
The closed event fires on the two edges into CLOSED, and it fires once: it's the last event the connection delivers. The other two events — discovered servers and lame duck — aren't part of the state machine at all. They arrive over a healthy connection, carrying information about the cluster.
How each client delivers the events
Every client reports this surface, but the delivery mechanism differs, so the same design reads differently in each language.
| Client | Mechanism | Where it's set |
|---|---|---|
| Go | one callback option per event | DisconnectErrHandler, ReconnectHandler, ReconnectErrHandler, ClosedHandler, DiscoveredServersHandler, LameDuckModeHandler |
| Python | one coroutine callback per event | connect arguments disconnected_cb, reconnected_cb, closed_cb, discovered_server_cb, lame_duck_mode_cb |
| Java | one listener, receiving an enum | connectionListener(...) on the options builder; the listener gets an Events value: DISCONNECTED, RECONNECTED, RESUBSCRIBED, DISCOVERED_SERVERS, LAME_DUCK, CONNECTED, CLOSED |
| JavaScript | one async iterator of status objects | nc.status(); each status has a type: disconnect, reconnect, reconnecting, update, ldm, close, and more |
| Rust | one async callback, receiving an enum | event_callback(...) on the connect options; the callback gets an Event value: Connected, Disconnected, LameDuckMode, Draining, Closed, and more |
| C# | one .NET event per kind | ConnectionDisconnected, ConnectionOpened, ReconnectFailed, LameDuckModeActivated, and more on the connection |
Three details are worth knowing before you rely on the table. Rust and
C# don't report a discovered-servers event; their pools still pick up
gossiped servers, they just don't announce it. C# has no closed event
at all — its terminal signal is the ConnectionState property, covered
below. And
JavaScript's update status carries the change itself: which servers
were added and which were deleted.
Wire the full surface on order-svc and log each event as it arrives.
The CLI wires all of these handlers internally, so its example watches
the events rather than wiring them:
- CLI
- C
#!/bin/bash
# Watch the full connection event surface from the CLI.
#
# The CLI wires every connection event handler itself, so there are no
# flags for wiring your own — that wiring is a client-library call. What the CLI
# does offer is the output of those handlers: with --trace it prints a
# ">>>" line per event. Restart a server in the pool and you see:
#
# >>> Connected to <url> on the first successful connect
# >>> Disconnected due to: <err>, will attempt reconnect
# >>> Reconnected to <url> after the client rejoins
# >>> Discovered new servers, known servers are now <urls>
# >>> Reconnect error: <err> per failed reconnect attempt
# >>> Connection is closed: <err> the closed event; the CLI exits
#
# The disconnect and closed lines print even without --trace; the rest
# are trace-only.
nats sub "orders.>" \
--server "nats://n1:4222,nats://n2:4222,nats://n3:4222" \
--connection-name order-svc \
--trace
// One callback per event. Disconnect, reconnect, discovered servers,
// and lame duck are logging; closed is the one that demands an action.
static void
onDisconnected(natsConnection *nc, void *closure)
{
printf("disconnected, reconnect loop running\n");
}
static void
onReconnected(natsConnection *nc, void *closure)
{
char url[256];
if (natsConnection_GetConnectedUrl(nc, url, sizeof(url)) == NATS_OK)
printf("reconnected to %s\n", url);
}
static void
onDiscoveredServers(natsConnection *nc, void *closure)
{
printf("server pool grew from cluster gossip\n");
}
static void
onLameDuck(natsConnection *nc, void *closure)
{
printf("server is shutting down soon (lame duck)\n");
}
static void
onClosed(natsConnection *nc, void *closure)
{
printf("connection closed, no more events after this\n");
}
const char *servers[] = {"nats://n1:4222", "nats://n2:4222",
"nats://n3:4222"};
if (s == NATS_OK)
s = natsOptions_SetServers(opts, servers, 3);
if (s == NATS_OK)
s = natsOptions_SetName(opts, "order-svc");
// Wire the full event surface before connecting. The C client has
// no per-attempt reconnect-error callback; a long outage shows as
// the gap between the disconnect and reconnect events.
if (s == NATS_OK)
s = natsOptions_SetDisconnectedCB(opts, onDisconnected, NULL);
if (s == NATS_OK)
s = natsOptions_SetReconnectedCB(opts, onReconnected, NULL);
if (s == NATS_OK)
s = natsOptions_SetDiscoveredServersCB(opts, onDiscoveredServers, NULL);
if (s == NATS_OK)
s = natsOptions_SetLameDuckModeCB(opts, onLameDuck, NULL);
if (s == NATS_OK)
s = natsOptions_SetClosedCB(opts, onClosed, NULL);
if (s == NATS_OK)
s = natsConnection_Connect(&conn, opts);
// Restart a server in the pool and watch the events print.
if (s == NATS_OK)
{
printf("connected, restart a server to see the events\n");
while (!natsConnection_IsClosed(conn))
nats_Sleep(100);
}
Disconnected versus closed
One branch in that event stream changes what your application does; the rest is logging. The branch is disconnected versus closed.
A disconnect is recoverable. The client is in RECONNECTING, cycling the pool with backoff, and everything on the last page exists to get it back to CONNECTED without your help. The right reaction is to log the event, count it in a metric, and let the reconnect loop work.
Closed is final. The client has stopped trying: no further events arrive, no publish succeeds, and nothing inside the client will change that. A service holding a CLOSED connection is broken until something outside the client acts — usually the process exits and a supervisor restarts it, or an alert reaches an operator.
The closed observer is whatever your client gives you to catch that
final transition: the ClosedHandler option in Go, closed_cb in
Python, the CLOSED event on Java's listener — its documentation calls
the state "permanently closed, either by manual action or failed
reconnects" — and Event::Closed in Rust. JavaScript adds a second
form: nc.closed() returns a promise that resolves when the client
closes, and resolves to the error if one caused the close (it never
rejects), so a main function can end by awaiting it. In C# you watch
the ConnectionState property instead.
Reconnection set MaxReconnect to -1, so for order-svc the
retries-spent edge never fires. CLOSED still happens: when the
application calls Close() or finishes a drain (next page), and on
errors where the client decides retrying can't help — in Go, for
example, receiving the same authentication error twice aborts
reconnecting regardless of the reconnect policy, unless you opt out with
IgnoreAuthErrorAbort. Credentials themselves are
TLS & Auth's topic; here it's
enough that CLOSED can arrive even when retries are unlimited.
So the routing for order-svc is: disconnect, reconnect, reconnect
error, discovered servers, lame duck — log them. Closed — exit or alert.
Polling the connection state
Events push transitions to you. Status polling is the pull side: asking the connection where it is right now. Every client exposes it, under different names:
- Go:
Status()returns one ofDISCONNECTED,CONNECTED,CLOSED,RECONNECTING,CONNECTING,DRAINING_SUBS,DRAINING_PUBS(the two draining states belong to the next page);IsConnected(),IsReconnecting(), andIsClosed()answer the common questions directly. - Python: the
is_connected,is_reconnecting,is_closed, andis_drainingproperties. - Java:
getStatus()returnsDISCONNECTED,CONNECTED,CLOSED,RECONNECTING, orCONNECTING. - Rust:
connection_state()returnsState::Pending,Connected, orDisconnected. - C#: the
ConnectionStateproperty returnsClosed,Open,Connecting,Reconnecting, orFailed. - JavaScript: there's no state getter beyond
isClosed()andisDraining(); you track state fromstatus()events, andclosed()covers the terminal case.
A poll is a snapshot, and a snapshot races the state machine: the answer
can be stale by the time you branch on it. A check that reads
CONNECTED and then publishes may publish into a connection that
dropped in between — which is fine, because the reconnect buffer from
the last page holds that publish. The mistake is using polls to react:
code that polls for RECONNECTING to trigger recovery logic duplicates
what the events already deliver, minus the ordering. Poll to report
state; react through events.
Go also offers the transitions as a channel: StatusChanged() returns
one that receives status changes (by default CONNECTED,
RECONNECTING, DISCONNECTED, CLOSED), for programs that prefer
select loops over callbacks.
A readiness check for order-svc
The two mechanisms combine into a health check. order-svc runs under a
platform that probes it — a container orchestrator, a load balancer, a
supervisor — and the connection state is exactly what the probe should
reflect:
- Ready while the connection is CONNECTED: the poll answers the probe.
- Not ready while it's RECONNECTING: the platform routes new work elsewhere, the process stays up, and the reconnect loop keeps working. RECONNECTING heals on its own; killing the process over it throws away the reconnect buffer.
- Dead on CLOSED: the closed observer exits the process, and the supervisor replaces it.
The CLI's version of a health check probes from the outside. nats server check connection opens its own connection and measures the
connect time, the server round-trip time, and a full publish/subscribe
round trip, each against a warning and a critical threshold, then exits
with the Nagios convention's code: 0 OK, 1 warning, 2 critical, 3
unknown. --format switches the output between nagios, json,
prometheus, and text. An external probe tells you the pool is
reachable and fast — it can't read your process's connection state; the
in-process check is a client-library call:
- CLI
- C
#!/bin/bash
# Probe the pool the way a monitoring system would.
#
# `nats server check connection` opens a fresh connection and measures
# three things against warn/critical thresholds: how long the connect
# took, the server round-trip time, and a full publish/subscribe round
# trip. The exit code follows the Nagios convention — 0 OK, 1 warning,
# 2 critical, 3 unknown — so the command slots into monitoring systems
# and container health probes directly.
#
# The thresholds below are the defaults, written out to show the knobs.
nats server check connection \
--server "nats://n1:4222,nats://n2:4222,nats://n3:4222" \
--connect-warn 500ms --connect-critical 1s \
--rtt-warn 500ms --rtt-critical 1s
# For scripting, `--format=json` renders the same result as JSON with a
# "status" field; `--format=prometheus` emits metrics and always exits 0.
# A quick latency read without thresholds: `nats rtt` computes the
# round-trip time to your servers, five round trips each by default.
nats rtt
// The readiness probe answers from a status poll: ready on
// CONNECTED, not ready while the reconnect loop works, dead on
// CLOSED. Poll to report state; react through the events.
switch (natsConnection_Status(conn))
{
case NATS_CONN_STATUS_CONNECTED:
{
// Connected: add the round-trip time as a latency read,
// one PING/PONG to the server, timed.
int64_t rtt = 0;
if (natsConnection_GetRTT(conn, &rtt) == NATS_OK)
printf("ready, rtt %dus\n", (int) (rtt / 1000));
break;
}
case NATS_CONN_STATUS_RECONNECTING:
// Heals on its own; report not ready so the platform
// routes new work elsewhere, and keep the process up.
printf("not ready: reconnecting\n");
break;
case NATS_CONN_STATUS_CLOSED:
// Terminal: the closed observer exits the process and
// the supervisor replaces it.
printf("dead: connection is closed\n");
break;
default:
printf("not ready\n");
break;
}
Force reconnect
One control on this surface points the other way: instead of observing a transition, you cause one. A force reconnect drops the current connection on purpose and runs the normal reconnect path — same pool, same backoff, same events.
The use case is rebalancing. A long-lived connection stays wherever it
landed. After a rolling restart of n1/n2/n3, every client that
failed over during the restart can end up on the one server that stayed
up the longest, and nothing ever moves them back. A forced reconnect
from some of the instances re-spreads the load.
Support and semantics vary by client, so check yours:
- Go:
ForceReconnect()starts the reconnect without waiting for it; if the client is already reconnecting, it forces an immediate attempt, skipping the current backoff delay. - Java:
forceReconnect()(plus an overload takingForceReconnectOptions) closes the connection immediately and runs the reconnect logic. - Python:
force_reconnect()moves a currently connected client to RECONNECTING and dials another server in the pool; it does nothing if the client is closed, already reconnecting, or not connected. - Rust:
force_reconnect()on the client. - C#:
ReconnectAsync()forces the reconnection and returns without waiting for the new connection. - JavaScript:
reconnect(), documented as experimental; it rejects if the client is closed or draining, and has no effect if the client is already disconnected. Its companionsetServers()replaces the server pool, and callingreconnect()afterwards moves the client onto the new pool immediately.
Treat it as a real disconnect, because it is one: the same risks apply as for any disconnect — JavaScript's documentation, for one, warns that in-flight requests are rejected — and the gap is covered by the same reconnect buffer with the same limits. And don't fire it across a whole fleet at once: a synchronized mass reconnect is the load spike that reconnect jitter exists to prevent. Stagger it.
Pitfalls
Three mistakes undo the observability this page adds. Each comes back to the two ideas: the event surface and status polling.
A status poll used as a trigger races the state machine. The poll
answers "where is the connection now", and the answer expires
immediately. Recovery logic gated on polling — a loop that checks for
RECONNECTING, a publish guarded by IsConnected() — acts on stale state
sooner or later, and re-implements what the event surface already
delivers in order. Poll for dashboards, metrics, and readiness probes;
react through the events.
No closed observer means permanent failure goes unnoticed. A
service that wires only the disconnect and reconnect callbacks logs a
disconnect, then nothing, forever: after CLOSED there are no more events
to log, and the service sits with a dead connection while its process
reports healthy. The closed event is the one that demands an action, and
it's the one most often left unwired. Wire it everywhere, even with
unlimited retries — a Close() on an error path or an authentication
failure can still get you to CLOSED:
- CLI
- C
#!/bin/bash
# React to CLOSED: exit so the supervisor restarts the process.
#
# The CLI already does what this page teaches. Its closed handler logs
# ">>> Connection is closed: <last error>" and terminates the process,
# while its disconnect handler only logs and leaves the recovery to the
# reconnect loop (the CLI sets its own MaxReconnects to -1, unlimited).
# You can't wire your own handler via flags; in a client library you wire the
# real pattern — a closed observer that exits or alerts.
#
# Run the subscriber, then close its connection from the server side to
# see the closed line and the exit.
nats sub "orders.>" \
--server "nats://n1:4222,nats://n2:4222,nats://n3:4222" \
--connection-name warehouse
// The closed observer: the last event the connection delivers. After
// this fires nothing inside the client recovers, so the reaction is to
// exit and let the supervisor restart the process, or raise an alert.
static void
onClosed(natsConnection *nc, void *closure)
{
const char *err = NULL;
natsConnection_GetLastError(nc, &err);
printf("connection closed: %s -- exiting for the supervisor\n",
(err == NULL) ? "no error" : err);
}
// Wire it even with unlimited reconnects: a Close() on an error
// path or a repeated authentication failure still reaches CLOSED.
if (s == NATS_OK)
s = natsOptions_SetClosedCB(opts, onClosed, NULL);
if (s == NATS_OK)
s = natsConnection_Connect(&conn, opts);
if (s == NATS_OK)
s = natsConnection_Subscribe(&sub, conn, "orders.>", onMsg, NULL);
// Block until the terminal transition, then exit non-zero.
if (s == NATS_OK)
{
while (!natsConnection_IsClosed(conn))
nats_Sleep(100);
}
Connection events and async errors are different signals. The events
on this page report state transitions of the connection. The async error
callback — Go's ErrorHandler, Python's error_cb, Java's
ErrorListener — reports problems that happen while the connection
stays CONNECTED, such as a subscription falling behind (that signal is
Slow Consumers' topic). In
JavaScript and Rust one stream carries both kinds, distinguished by
type. Folding them into one "connection problems" metric muddles
diagnosis: rising disconnects point at the network or the servers,
rising async errors point at your application. Count them separately.
Where you are
The connection state machine that Reconnection built is now observable from the outside. You have:
- the full event surface named and routed: disconnect, reconnect, reconnect error, discovered servers, and lame duck are logged; closed triggers an exit or an alert
- a closed observer on every long-lived service, so a permanent failure is acted on instead of silently absorbed
- status polling wired to the readiness probe — ready on CONNECTED, not ready on RECONNECTING, dead on CLOSED — and events, not polls, driving reactions
- force reconnect, where your client offers it, for deliberately rebalancing long-lived connections after cluster changes
Every transition so far happened to the connection; this page made them visible. The one transition you haven't driven deliberately is the exit: shutting the connection down cleanly.
What's next
The next mechanism is drain and shutdown: telling the connection to finish its in-flight work — deliver buffered messages, flush pending publishes — and only then close, so a deploy or a SIGTERM never loses work. That page also covers the lame-duck event from this page's list in working context.
Continue to Drain & Shutdown.
See also
- Reconnection — the disconnect/reconnect loop these events report on.
- Monitoring — the server's view of your connections, from the outside.
- Reference — the wire protocol and server configuration; the event and status APIs for each client live in its API reference.