Hi @prohand, hi @spenceur!
Thanks for bringing this up, the need is very real: an integration that fails silently (typically an expired token) is the worst possible failure in home automation. Nothing visibly breaks, and you only notice three weeks later 
Good news: for external integrations, most of the monitoring already exists on the server side. Each integration sends a heartbeat to Gladys, a health check runs continuously, and the integration goes into a degraded state when it loses its connection to the third-party service or stops responding, with automatic restart and backoff behind the scenes.
The « degraded » state that you proposed @spenceur already exists internally
What’s missing is the user-facing part: being notified.
On the « Connected → Disconnected » alert: the TV example shows the trap well. « Disconnected » can mean two completely different things:
- a normal case: device turned off, WebSocket closed, expected behavior;
- an abnormal case: expired token, revoked access, API error.
The distinction already exists partly in the model (the integration is running but its connection to the cloud has dropped, so degraded), we just need to ensure that each integration reports the correct signal, rather than blindly alerting on any disconnection.
On exclusions: I prefer to avoid making it the main mechanism. It forces everyone to empirically discover which integrations are « noisy », after receiving false alerts. If we need it from the start, it’s because the reported signal is not the right one.
Concretely, I would see:
- Make the state visible: the state of each integration + the date of the last successful exchange, directly in the interface. Zero notifications, but already a big part of the value.
- A global setting in Gladys’ parameters: « Alert me when an integration is in error ». A simple switch, no configuration per integration or scene to create. And it naturally covers your request @prohand of « only if it has already been connected »: degraded means by construction « it was working, it’s not working anymore ».
- Anti-flapping built into the setting: a minimum delay in degraded state before alerting (to let the automatic restart do its job), no duplicates, and a notification of return to normal. Without that, this kind of feature ends up deactivated by everyone in a week.
On automatic token renewal: 100% agreement in principle, but it has to be handled integration by integration, and even with perfect refresh there will still be cases where access drops (revocation by the manufacturer, API change…). Monitoring must therefore still exist.
In short: the foundations are already there, the bulk of the work is to expose all this to the user 