Skip to content

Fix fast polling when PollControl is not on first EP - #1817

Draft
TheJulianJES wants to merge 3 commits into
zigpy:devfrom
TheJulianJES:tjj/fix-poll-control-multi-endpoint
Draft

TheJulianJES wants to merge 3 commits into
zigpy:devfrom
TheJulianJES:tjj/fix-poll-control-multi-endpoint

Conversation

@TheJulianJES

@TheJulianJES TheJulianJES commented May 5, 2026 •

Copy link
Copy Markdown
Contributor

Proposed change

This fixes an issue where fast polling is not correctly activated during initial pairing of a device (or a re-interview) if the PollControl cluster is not on the first endpoint. This is an issue with all/many Frient end-devices.

Skim the AI description below for more information. Note: Also read the related issue at the bottom.

Impact on Frient devices with limited binding table

However, this breaks Frient devices – at least my Frient Intelligent Smoke Alarm that shipped with v4.0.8.

Currently, we automatically bind 4 clusters on endpoint 35 using ZHA (BinaryInput, PowerConfiguration, IasWd, IasZone). They all bind properly without the patch. With it, PollControl is also bound now, which then causes a TABLE_FULL error for IasZone...

  • Z2M explicitly unbinds the PollControl cluster and states that only 4 clusters can be bound: https://github.com/Koenkk/zigbee-herdsman-converters/blob/42d3b11e44fbdccf1bf75c0d63f8983e030af622/src/devices/develco.ts#L426-L432
  • Apparently this was an issue introduced with firmware v4.0.8 – v4.0.7 was fine but no newer version is available:
  • Note: We also bind TemperatureMeasurement but that's on EP 38, not 35. The 4 bind limit is apparently per EP.
  • Technically, we may not need the BinaryInput bind, as we don't use that cluster, but rely on IasZone instead.
    • Z2M seems to bind this to display the reliability attribute as a fault sensor
  • We may also not need the IasZone bind because the write to "IAS CIE Address" should deal with the change notifications instead?
    • TODO: Why does Z2M explicitly bind this cluster then?
    • TODO: Or why do we in general? Look at spec and compare CIE address vs binding
  • Note: A re-interview also does NOT clear bindings, only a factory reset does. This also seems to be the case for Z2M.
  • Z2M does not bind the PowerConfiguration cluster on this endpoint.
    • TODO: Investigate if we still get battery voltage reports anyway?
    • EDIT: Not true. Z2M also binds this. See my comment below here.

This needs to be checked further.

AI description

Bug

In Device._discover(), fast polling is supposed to be set up "as soon as we are aware of a PollControl cluster" — begin_fast_polling is called after each endpoint is initialised, until it succeeds. The retry-stop signal was implemented as the else branch of a try/except (TimeoutError, DeliveryError):

for ep in self.non_zdo_endpoints:
    await ep.initialize()
    if not initiated_fast_polling:
        try:
            await self.begin_fast_polling()
        except (TimeoutError, DeliveryError):
            pass
        else:
            initiated_fast_polling = True

But begin_fast_polling() also returns silently (no exception) when find_cluster(PollControl.cluster_id) raises ValueError — i.e. when no endpoint has a PollControl cluster yet. Endpoint clusters aren't populated until ep.initialize() runs on that specific endpoint, so on a fresh multi-endpoint device whose PollControl lives on something other than the first endpoint in non_zdo_endpoints, the very first call (after ep1 init) returns silently, the else branch fires, initiated_fast_polling is set to True, and the retry on subsequent endpoints is suppressed. No Bind_req for cluster 0x0020 is ever sent, no fast_poll_timeout write happens, the device never starts emitting PollControl check-ins, and zigpy never reaches it again.

Real-world reproduction

A frient WISZB-131 with endpoints [1, 35, 38] and PollControl on EP35. After OTA + re-interview, the only Bind_reqs on the wire are ZHA's configure-phase ones (BinaryInput, PowerConfiguration, TempMeas, IAS Zone) — no 0x0020. The Device does not support fast polling debug line fires once after EP1's Simple_Desc_rsp and zigpy never re-checks. Devices that don't persist their binding table across reboot/firmware change then silently stop sending check-ins.

Fix

Use self._fast_polling (only set on a real bind inside begin_fast_polling()) as the retry-stop signal instead of the absence-of-exception:

if not initiated_fast_polling:
    with contextlib.suppress(TimeoutError, DeliveryError):
        await self.begin_fast_polling()
    initiated_fast_polling = self._fast_polling

Entry-state initiated_fast_polling = self._fast_polling is preserved at the top of the block, so devices that are already in fast polling mode (e.g. mid-OTA re-init) still skip the loop entirely. The symmetric if self.all_endpoints_init: ... branch is unchanged — that path only calls begin_fast_polling() once and is unaffected.

Test coverage

Three new tests in tests/test_device.py:

  • test_initialize_fast_polling_pollcontrol_on_later_endpoint — multi-endpoint device with PollControl on ep2. Asserts dev._fast_polling is True, that bind() was called, and that fast_poll_timeout was written. Verified to fail against pre-fix code (the regression target).
  • test_initialize_fast_polling_pollcontrol_on_first_endpoint — same shape but PollControl on ep1, ensuring the previously-working path still works after the change.
  • test_initialize_fast_polling_no_pollcontrol — device with no PollControl anywhere; asserts no Bind_req is sent and _fast_polling stays False.

A small helper _make_fast_polling_init_mocks factors the shared mockepinit / mock_ep_get_model_info setup.

Full tests/test_device.py (66 tests) passes.

Out of scope

A separate, orthogonal issue exists in begin_fast_polling() itself: Cluster.bind() returns the Bind_req status tuple rather than raising on non-SUCCESS, so a TABLE_FULL (or other) bind failure is silently ignored and self._fast_polling is set to True regardless. Not addressed here — should be tackled separately.

`_discover` initialised endpoints sequentially and called
`begin_fast_polling` after each one, using a `try/except…else` to mark
fast polling as initiated. But `begin_fast_polling` also returns
silently when no `PollControl` cluster is found yet, so the `else`
branch fired even when nothing was bound. On a fresh multi-endpoint
device whose `PollControl` lives on a later endpoint, we'd skip every
subsequent retry once that first (futile) call returned, leaving fast
polling unbound for the rest of the interview.

Use `_fast_polling` (only set on a real bind) as the retry signal so we
keep trying after each endpoint until `PollControl` actually shows up.
Covers the case where `PollControl` lives on a later non-ZDO endpoint:
the first `begin_fast_polling` call (after ep1 init) returns silently
because no endpoint has `PollControl` yet, and the retry must still
fire once ep2 has been initialised.

Fails against the prior `try/except…else` logic and passes against the
`_fast_polling`-based retry signal.
…trol cases

The earlier regression test only covered PollControl on a later endpoint.
Add two more cases to guard the working paths against future regressions:

- PollControl on the first endpoint: bind + fast_poll_timeout write must
  still happen during interview.
- No PollControl on any endpoint: no Bind_req should be sent at all.

Also strengthen the later-endpoint test to assert the bind and
fast_poll_timeout write, and factor the shared mock setup into a helper.
@codecov

codecov Bot commented May 5, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 99.55%. Comparing base (2087790) to head (6f25fb9).

Additional details and impacted files
@@            Coverage Diff             @@
##              dev    #1817      +/-   ##
==========================================
- Coverage   99.55%   99.55%   -0.01%     
==========================================
  Files          64       64              
  Lines       13212    13210       -2     
==========================================
- Hits        13153    13151       -2     
  Misses         59       59              

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@TheJulianJES

Copy link
Copy Markdown
Contributor Author

Looking at the technical manual for this device, it seems like there's an "auto bind" for any PollControl clusters on the coordinator? If there is, we may not want to bind the PollControl cluster ourselves? Doing so seems to take another slot..
https://www.smarthome-store.de/media/documents/smoke-alarm-technical-manual.pdf

The "Power Configuration" and "Temperature Measurement" may also be bound automatically in "EZ Mode":

The following clusters are support in EZ-mode finding and binding:

  • Temperature cluster
  • Power configuration cluster

This might explain why Z2M doesn't bind them. Though the temperature cluster is on another endpoint anyways, which seems to not share the 4 binding limit with the main EP.

@TheJulianJES

Copy link
Copy Markdown
Contributor Author

In general, we should also investigate whether binding IasZone does something at all. We don't set up attribute reporting or anything, and the change notifications (should) use the IAS CIE address of the coordinator, not whatever the cluster was bound to.

Z2M also binds the cluster on a lot of devices, but I'm not sure why.

@TheJulianJES

Copy link
Copy Markdown
Contributor Author

Ah, Z2M does set up attribute reporting for PowerConfiguration, like we do. It's just generically done here: https://github.com/Koenkk/zigbee-herdsman-converters/blob/master/src/lib/modernExtend.ts (by Frient devices using m.battery(), then calling setupConfigureForReporting and setupAttributes above.)

@TheJulianJES

Copy link
Copy Markdown
Contributor Author

With the 4 binding slots on EP 35 filled, excluding PollControl, the auto-bind mentioned in the technical documentation does not happen. But if, for example, the IasZone cluster is not bound by ZHA, the smoke sensor does automatically bind the cluster and send check-ins ~40 seconds after the initial join. If you then try to bind the IasZone cluster, you'll get TABLE_FULL.

@TheJulianJES

TheJulianJES commented May 6, 2026 •

Copy link
Copy Markdown
Contributor Author

These are the Frient devices in our diagnostics that have more than 4 clusters on an endpoint that ZHA likely binds:

  • AQSZB-110
  • EMIZB-141
  • EMIZB-151
  • FLSZB-110
  • MOSZB-153
  • SCAZB-141
  • SMSZB-120
  • SPLZB-141 (no PollControl)
  • WISZB-131

In particular, these are the Telink TI 2530 based devices from the list (that may be more likely to have the 4 binds per EP limit):

  • AQSZB-110 (Air Quality Sensor)
  • FLSZB-110 (Water Leak Detector)
  • SMSZB-120 (Intelligent Smoke Alarm: known)

With this PR, we should ideally remove and reset the AQSZB-110 + FLSZB-110 sensors to see if further binds by ZHA error out with TABLE_FULL.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant