Monitoring external integrations

Hello,

Would it be possible to implement a kind of supervision with an alert on the communication method of your choice when an external integration changes its status from « Connected » to « Disconnected » (it must have been connected at least once, otherwise we will receive alerts as soon as the external integration is installed)?

Indeed, the goal is to be alerted, for example, for Daikin if the connection token expires to restart the connection or any other issues that would change the status to « Disconnected ».

Perhaps the same alert could also be set for the « In Progress » status.

Thank you

For TV integration, this isn’t possible. When the TV is on, we’ll have the info that we’re connected. But once turned off, we risk having false notifications.

I’m talking about this:

When the TV is off, does « Connection » switch to disconnected on your end?

Yes :slight_smile: because the websocket is closed

Let’s see if we can add some exclusions :wink:

But why isn’t it the service that manages its token renewal?
In my case, when I turn the TV back on, I have a WOL to physically turn on the TV and a socket reconnection.

For you, when the token expires, shouldn’t it automatically renew it?
Or do you have to do a manual action, of course? :thinking:

I haven’t encountered this case with Daikin as I believe the token expires in 1 year
So I’ll see in 1 year if it is automatically renewed or not :sweat_smile:

But I was just giving this as an example, it may be in a disconnected state because Daikin is experiencing a problem, for example, or the integration encountered an unhandled error and it crashed and ended up in a disconnected state

What are the connection statuses? Connected, disconnected? Maybe we should have a degraded state?

No idea :sweat_smile:
Yeah maybe a degraded state would be nice

Hi @prohand, hi @spenceur!

Thanks for bringing this up, the need is very real: an integration that fails silently (typically an expired token) is the worst possible failure in home automation. Nothing visibly breaks, and you only notice three weeks later :sweat_smile:

Good news: for external integrations, most of the monitoring already exists on the server side. Each integration sends a heartbeat to Gladys, a health check runs continuously, and the integration goes into a degraded state when it loses its connection to the third-party service or stops responding, with automatic restart and backoff behind the scenes.

The « degraded » state that you proposed @spenceur already exists internally :+1: What’s missing is the user-facing part: being notified.

On the « Connected → Disconnected » alert: the TV example shows the trap well. « Disconnected » can mean two completely different things:

  • a normal case: device turned off, WebSocket closed, expected behavior;
  • an abnormal case: expired token, revoked access, API error.

The distinction already exists partly in the model (the integration is running but its connection to the cloud has dropped, so degraded), we just need to ensure that each integration reports the correct signal, rather than blindly alerting on any disconnection.

On exclusions: I prefer to avoid making it the main mechanism. It forces everyone to empirically discover which integrations are « noisy », after receiving false alerts. If we need it from the start, it’s because the reported signal is not the right one.

Concretely, I would see:

  1. Make the state visible: the state of each integration + the date of the last successful exchange, directly in the interface. Zero notifications, but already a big part of the value.
  2. A global setting in Gladys’ parameters: « Alert me when an integration is in error ». A simple switch, no configuration per integration or scene to create. And it naturally covers your request @prohand of « only if it has already been connected »: degraded means by construction « it was working, it’s not working anymore ».
  3. Anti-flapping built into the setting: a minimum delay in degraded state before alerting (to let the automatic restart do its job), no duplicates, and a notification of return to normal. Without that, this kind of feature ends up deactivated by everyone in a week.

On automatic token renewal: 100% agreement in principle, but it has to be handled integration by integration, and even with perfect refresh there will still be cases where access drops (revocation by the manufacturer, API change…). Monitoring must therefore still exist.

In short: the foundations are already there, the bulk of the work is to expose all this to the user :slightly_smiling_face: