All posts

Thursday, 14:32 — the Z2M cascade nobody was meant to hear

A fictional shift in the life of a solo integrator: three Z2M crashes after a silent add-on update, and the forty minutes between the first symptom and a quiet fix.

14:32 — A tile goes amber

Thursday, early afternoon. Sandra — a fictional solo integrator somewhere between Zürich and Konstanz, looking after 23 Home Assistant installations — is in her office working up a quote. The fleet dashboard runs on the second monitor. She isn’t looking at it. It’s always open and almost never noisy.

At 14:32:04 a tile flips from green to amber. Kanzlei Weber. Disk usage just past 70%. No drama. She takes a sip of coffee.

14:34 — Then the second. Then the third.

Two minutes later the Haus Zürichberg tile drops. Status: critical. Source: zigbee2mqtt. Adapter unreachable. Eight seconds after that, Büro Meier. Same adapter, same message.

Three installations, three customers, the same symptom, inside two minutes. That isn’t coincidence. That’s a pattern.

She closes the quote. With any luck none of the three customers has noticed yet that the lunchtime lighting automation didn’t fire.

14:38 — The common add-on

She filters the dashboard on zigbee2mqtt. Seven installations run it. Three are down. Four aren’t. She sorts by add-on version:

  • The three that are down all run 2.5.1. Auto-updated this morning.
  • The four still running are on 2.4.x. Auto-update off, because that’s how the customer wanted it.

There it is. Not the hub. The update.

14:44 — Three maintenance requests

Sandra opens the bulk action. Pick three customers, send one maintenance request with identical wording. Subject: “Restart Z2M adapter after update”. Reason: “Known bug in 2.5.1 — rolling back to 2.4.6”. Duration: 30 minutes.

She knows they won’t all accept at the same time. Frau Weber is in a client meeting. The Meier family is on holiday. That’s fine, the tunnel waits for each of them. And the direction matters: the connection dials out from the installation rather than the other way round, which is exactly why it still reaches boxes stuck behind CG-NAT with no open ports, the places a classic VPN never gets near.

14:51 — Three clicks at the customer end

At Haus Zürichberg the request lands as an actionable card right inside Home Assistant. The owner is at the kitchen table, sees the prompt, reads “Restart Z2M adapter”, taps Accept. 38 seconds later Sandra has a URL that opens straight into his HA frontend.

Frau Weber clears the request thirteen minutes later. The Meier family follows at 15:04, from a beach café. Three clicks, three sessions, three tunnels — each one with a 30-minute window, then automatically gone.

15:09 — Rollback, tunnel closed

The fix itself is trivial. Downgrade the add-on to 2.4.6, restart the adapter, status green. Under four minutes per customer. While she’s still working through the third tunnel she’s already typing the entry in her internal wiki:

“Z2M 2.5.1 — pause auto-update until patched.” Tag: incident, severity: medium, affected: 3/7.

At 15:24 everything is green again. None of the three customers called.


What would have happened, the old way

Before the fleet dashboard this story plays out differently.

Sandra would have spotted the Z2M crashes the next morning, when she checked her own home automation. Frau Weber would have called that evening because the office lighting wasn’t responding, appointment for Friday. The Meier family would have written in from holiday, annoyed that the heating wasn’t coming up. Trust slightly dented. Three separate drives, three separate VPN sessions, three separate explanations of what went wrong. This is the exact seam where a tinkering hobby either turns into a real business or doesn’t: anyone running Home Assistant as a business simply can’t afford three separate drives per incident.

What happened instead: forty minutes between the first symptom and a quiet fix. Three customers who didn’t notice a thing, and a Thursday afternoon that got rescued before it turned into an emergency. Sandra isn’t real. The pattern is, and anyone who recognises it in their own week can put their fleet under one shared dashboard in a few minutes and find out, next Thursday, whether the forty minutes hold.

This story is fictional. The people, companies and version numbers don’t exist as described. What’s real is the pattern.

DO
Denny Ovčar
Founder · ha-fleet-manager.com
Reply
Share