#eu: downtime started, error: Simulated telematics device timed out waiting for recently added command. Usually this indicates the problem with flespi Telematics Gateway.
Update: our uplink provider is mitigating a large-scale DDoS attack (~50–60 Gbps) that is also impacting several upstream providers.
Traffic is being rerouted through a scrubbing center for filtering. Short intermittent disruptions remain possible.
Sorry for any inconvenience.
Traffic is being rerouted through a scrubbing center for filtering. Short intermittent disruptions remain possible.
Sorry for any inconvenience.
👍6
Now the networking problem is gone and we are experiencing the CPU overload in MQTT Broker and messages delivery delays within the MQTT bus. We are working on it.
👎7👍1
We are still working on the system. Unfortunately we had to shut down the telematics gateway due to high load on the MQTT Broker and now restoring the broker back, giving it more resources. Once settled we will slightly open window for devices for passing the traffic through, initially narrow. It will take at least dozens of minutes more.
👍6👎5
We are gradually starting channels and probing the load, now you may expect some messages to pass through the system. REST system works stable. MQTT Broker load we are evaluating. If everything will be OK we will gradually increasing the window for devices.
👍12
We are very close to 50% of uptime checking nodes to report their UP status and stop the global downtime. Gradually increasing the load and passing telematics data through...
👍11👎1
#eu: downtime started, error: Simulated telematics device connected to the channel, sent the packet with message, but channel didn't replied to it. Usually this indicates the problem with flespi Telematics Gateway.
#eu: downtime started, error: Simulated telematics device timed out waiting for recently added command. Usually this indicates the problem with flespi Telematics Gateway.
👎2
#eu: downtime started, error: Simulated telematics device timed out waiting for recently added command. Usually this indicates the problem with flespi Telematics Gateway.
#eu: downtime started, error: Simulated telematics device timed out waiting for recently added command. Usually this indicates the problem with flespi Telematics Gateway.
Mostly the load is now passed and full window for telematic data ingestion is opened already for dozens of minutes. Still we have a huge load and still periodically some remote nodes report delays. So we may expect some more short downtime reports.
We are now checking other systems down the data flow such as analytics and streams, checking the databases and services consistency and all these minor things that should be checked after such a major outage.
During this downtime we found and fixed some narrow places that triggered only on the high load. Later we will analyze and trace the outage in details and address whatever find.
You can track live flespi status and load here: https://flespi.com/status
We are now checking other systems down the data flow such as analytics and streams, checking the databases and services consistency and all these minor things that should be checked after such a major outage.
During this downtime we found and fixed some narrow places that triggered only on the high load. Later we will analyze and trace the outage in details and address whatever find.
You can track live flespi status and load here: https://flespi.com/status
Flespi
flespi platform status page - uptime, updates, counters
flespi live status page aggregates monthly uptime values, latest module updates, counters, change logs, and more.
We diagnosed that analytics system has intervals data gaps due to the downtime. We are working to restore the data.
👍8
All systems are stable and the load is fully under control. We are still restoring analytics intervals — currently at ~50% progress, with an estimated 2 more hours until fully restored.
If you experience any issues, please contact us in the flespi chat.
It's 1 AM on Saturday our local time, and almost the entire engineering team is still in the office (since 8 AM Friday), running the restoration process and keeping an eye on anything that may pop up.
If you experience any issues, please contact us in the flespi chat.
It's 1 AM on Saturday our local time, and almost the entire engineering team is still in the office (since 8 AM Friday), running the restoration process and keeping an eye on anything that may pop up.
👍13
We have restored all intervals data in analytics. All systems are stable. If you experience any issues, please report them in the flespi chat — we will take care of it.
A more detailed incident report will be published later in July.
Sorry for any inconvenience and for the ruined Friday - we did our best to make it as short as possible.
A more detailed incident report will be published later in July.
Sorry for any inconvenience and for the ruined Friday - we did our best to make it as short as possible.
👍17