An ESP32 cloud project should not stop doing its local job just because Wi-Fi, the router, DNS, or the cloud service disappears. The reliable architecture is local-first: sensing and control continue locally, network state is treated as temporary, data is buffered deliberately, and synchronization resumes after connectivity returns.
Espressif’s Wi-Fi documentation makes an important point: after a station disconnect event, the application is responsible for reconnection, and existing sockets may no longer be usable. Design for that failure instead of treating connection loss as exceptional.
1. Separate the device function from the cloud function
List which features must work offline and which genuinely require the cloud.
| Function | Offline behavior |
|---|---|
| Read sensors | Continue locally. |
| Safety/control loop | Must not depend on Internet availability. |
| Local display/relay logic | Continue using last valid local configuration. |
| Cloud telemetry | Queue or summarize while offline. |
| Remote commands | Unavailable offline; define safe default behavior. |
| OTA update | Pause cleanly and retry only through the update mechanism’s safe rules. |
If a loss of Internet disables a local control loop, the architecture is too tightly coupled unless the product truly cannot operate without a remote service.
2. Use explicit connection states
Do not scatter if (WiFi.status()) checks through the application. Model connectivity as a small state machine:
OFFLINE
CONNECTING_WIFI
WAITING_FOR_IP
CONNECTING_SERVICE
ONLINE
BACKOFF
Network events update the state. Application tasks publish data to a local queue and do not block waiting for the network.
3. Recreate sockets after Wi-Fi loss
In ESP-IDF, a station disconnection can invalidate existing TCP/UDP socket state. The application should close/re-create network connections as needed after Wi-Fi is restored. A robust sequence is:
- receive the disconnect event;
- mark cloud transport offline;
- stop using the old socket/client session;
- schedule Wi-Fi reconnection;
- wait for a valid IP event;
- create a fresh MQTT/HTTP/WebSocket connection;
- resume synchronization from the local queue.
Do not assume that “Wi-Fi connected again” means an old cloud session is still valid.
4. Reconnect with backoff, not a tight loop
A device that retries continuously can waste power and overload the network stack. Use bounded exponential backoff with jitter or a similar retry policy.
1 s → 2 s → 4 s → 8 s → 16 s → 30 s max
Reset the delay after a stable successful connection. Distinguish intentional disconnects from failures so the device does not immediately reconnect when your application deliberately turned Wi-Fi off.
5. Buffer data according to its value
Not every sample deserves permanent storage. Decide what must happen when the cloud is unavailable:
- Latest-state data: keep only the newest value.
- Events/alarms: persist every important event until acknowledged.
- High-rate telemetry: aggregate or downsample rather than filling flash.
- Billing/compliance data: use stronger persistence and sequence tracking appropriate to the application.
A ring buffer with sequence numbers and timestamps is usually easier to reason about than an unbounded queue.
6. Do not write every sensor sample to NVS
ESP-IDF NVS is designed for relatively small key-value data such as configuration, calibration values, and state flags. It includes wear-management mechanisms, but it is not a reason to write high-rate telemetry to flash one sample at a time.
Use NVS for things like:
- Wi-Fi/application configuration;
- calibration constants;
- last acknowledged sequence number;
- small state flags needed after reset.
For larger offline logs, use a storage strategy suitable for sequential records and the expected write rate, such as a filesystem/partition or external storage where appropriate.
7. Give every queued record an identity
Reconnect logic can create duplicates. Add a monotonic sequence number or unique event ID to records sent to the backend.
{
"device_id": "sensor-07",
"seq": 18452,
"timestamp": 1790672120,
"temperature": 28.4
}
The server can then detect retries instead of counting the same event twice. This is especially important when the ESP32 sends a record but loses the connection before it receives the application’s acknowledgement.
8. Separate transport acknowledgement from business acknowledgement
MQTT QoS, TCP delivery, and HTTP status codes help with transport reliability, but they do not automatically prove that your application processed a command or permanently stored an event.
For critical workflows, define application-level acknowledgement. Example:
- device sends event 18452;
- server stores/processes event;
- server acknowledges 18452;
- device advances its persisted acknowledged pointer.
This makes recovery after reset or reconnect deterministic.
9. Remote commands need expiry and idempotency
An old cloud command should not execute unexpectedly after a long outage. Commands should carry an ID and, where appropriate, an expiry time or version.
For actions that can be retried, design handlers to be idempotent when possible. “Set relay to ON” is easier to retry safely than “toggle relay.” The second command changes meaning if it is delivered twice.
10. Keep local time semantics clear
If records need timestamps, define what happens before Internet time synchronization. Options include:
- use a real-time clock;
- store monotonic uptime plus a sync reference;
- mark records as unsynchronized until wall-clock time becomes trustworthy.
Never silently label an estimated timestamp as precise cloud-synchronized time.
11. Add watchdog and brownout observability
Network problems often get blamed for failures caused by power or blocking code. Record reset reasons, reconnect counts, heap high-water/low-water information, queue depth, failed DNS/TCP/MQTT attempts, and the age of the oldest unsent record.
If Wi-Fi radio current peaks pull down a weak supply, reconnection loops can look like a software bug. Measure the power rail during transmit/reconnect activity when resets correlate with networking.
12. Test failure on purpose
A cloud project is not validated by one successful connection. Run controlled tests:
- boot with router unavailable;
- disable the access point for 1 minute and restore it;
- leave Wi-Fi connected but block Internet/cloud access;
- restart the MQTT/HTTP backend while the device runs;
- change the device IP address/DHCP lease;
- fill the offline queue to its defined limit;
- power-cycle the ESP32 while records are pending;
- restore connectivity and verify ordering, duplicates, losses, and recovery time.
Example acceptance criteria
| Test | Pass condition |
|---|---|
| Wi-Fi unavailable at boot | Local sensing/control starts normally; reconnect attempts do not block it. |
| AP disappears for 5 min | No crash; data policy works; device reconnects automatically after AP returns. |
| Cloud service unavailable | Local function continues and queue remains bounded. |
| ESP32 resets while offline | Required configuration/state survives and recovery logic starts cleanly. |
| Connectivity returns | New data resumes and retained events synchronize without silent loss. |
A reliable architecture in one sentence
Sensors and control write to local state; a separate communications layer synchronizes that state with the cloud whenever the network is healthy.
This separation makes the firmware easier to test and prevents a cloud outage from becoming a device outage.
Related EET guides: MQTT vs HTTP vs WebSockets on ESP32 and why embedded projects fail after deployment.
Primary references
- Espressif ESP-IDF — Wi-Fi station scenarios and reconnect behavior
- Espressif ESP-IDF — Non-Volatile Storage library