An ESP32 that connects to Wi-Fi once is not necessarily a reliable IoT device. Real networks change: access points restart, signal strength drops, DHCP leases renew, authentication can fail, and the ESP32 may temporarily lose connectivity.
The correct design goal is not “never disconnect.” It is detect the disconnect, identify why it happened, recover automatically, and keep the application responsive while reconnecting.
First, separate a Wi-Fi disconnect from an ESP32 reset
If the whole ESP32 restarts when Wi-Fi begins or reconnects, investigate power first. Wireless transmission can expose a weak 3.3 V rail and trigger a brownout. Our ESP32 random reboot guide shows how to distinguish brownouts, watchdogs, panic resets, and boot problems.
If the ESP32 keeps running but loses its IP connection or reports a station disconnect event, continue with the network checks below.
1. Log the disconnect reason instead of guessing
ESP-IDF provides Wi-Fi disconnection reason codes. Those codes can distinguish authentication problems, inactivity, handshake timeouts, access-point behavior, and other causes. A reliable system should record the reason whenever WIFI_EVENT_STA_DISCONNECTED occurs.
This matters because different failures need different fixes. Repeated authentication timeouts are not solved the same way as weak signal or an access point that disappeared.
2. Do not assume one call to connect means permanent connectivity
Espressif documents that esp_wifi_connect() attempts to connect to an access point, but applications that need reconnection should implement reconnect logic or use the available retry configuration. In a deployed IoT product, reconnection belongs in the application design.
Your application should therefore maintain a small Wi-Fi state machine rather than blocking inside a loop until the network returns.
A robust reconnection pattern
A good reconnection strategy has four parts:
- Receive or detect a disconnect event.
- Record the reason and timestamp.
- Schedule a reconnect attempt without blocking the rest of the application.
- Back off if repeated attempts fail.
For Arduino-style projects, the exact event names vary by ESP32 Arduino core version, but the architecture should look like this:
unsigned long lastReconnectAttempt = 0;
const unsigned long reconnectInterval = 10000;
void maintainWiFi() {
if (WiFi.status() == WL_CONNECTED) {
return;
}
unsigned long now = millis();
if (now - lastReconnectAttempt >= reconnectInterval) {
lastReconnectAttempt = now;
WiFi.reconnect();
}
}
void loop() {
maintainWiFi();
// Sensor reading, control logic, display updates,
// logging, and other tasks continue running here.
}
The important part is that the code does not sit in a long while loop waiting for Wi-Fi. Your sensors, control logic, watchdog servicing, and local user interface should continue to operate.
3. Add backoff instead of reconnecting continuously
A device that calls reconnect as fast as possible can waste power, overload logs, and create unnecessary traffic when the access point is genuinely unavailable.
A better sequence is:
- First retry after a few seconds.
- If that fails repeatedly, increase the delay.
- Cap the delay at a reasonable maximum.
- Reset the retry delay after a successful connection.
For example: 5 s → 10 s → 20 s → 40 s → 60 s maximum.
Battery-powered systems benefit especially from bounded retry behavior because continuous connection attempts can dominate energy consumption.
4. Measure signal strength where the device will actually run
A board that works perfectly beside your router may fail after installation behind a wall, inside a metal enclosure, or near motors and switching power supplies.
Log RSSI while the product operates in its final location. More negative RSSI values generally indicate a weaker received signal. The exact threshold for reliable operation depends on the environment, data rate, interference, antenna design, and access point.
Do not optimize around a single RSSI reading. Watch the range over time, especially when doors close, equipment switches on, or people move through the area.
5. Check the access point, not only the ESP32
Many “ESP32 Wi-Fi bugs” originate in router configuration or RF conditions.
Check these items:
- The SSID is available on 2.4 GHz for classic ESP32 devices.
- The access point is not automatically changing configuration in a way the device cannot follow.
- Security credentials are correct and stable.
- DHCP has enough address capacity.
- The router is not aggressively removing idle clients.
- There is no MAC filtering or client isolation causing unexpected behavior.
- Channel congestion is not extreme.
- Mesh or multi-AP environments are not causing undesirable roaming behavior.
6. Authentication and handshake failures need different treatment
ESP-IDF exposes reason codes for authentication expiry, handshake timeout, association problems, and other Wi-Fi events. When the same reason repeats, investigate that specific stage.
Examples:
| Observed pattern | Likely area | What to inspect |
|---|---|---|
| Handshake timeout | Security/authentication | Password, AP security mode, RF loss during handshake |
| Repeated inactivity disconnect | AP policy or link quality | Router settings, signal, power-save behavior |
| AP not found | Coverage or SSID availability | 2.4 GHz network, channel, antenna, location |
| Connects but no internet/service | IP/DNS/server layer | DHCP, gateway, DNS, remote service |
7. Wi-Fi connected does not mean the application service is healthy
Your ESP32 can remain associated with the access point while MQTT, HTTP, DNS, or the remote cloud service is unavailable.
Track connectivity in layers:
- Wi-Fi link: is the station associated?
- IP layer: does it have a valid IP address and gateway?
- Service layer: can it reach the MQTT broker, API, or server?
- Application layer: is useful data actually being acknowledged?
Do not reboot the entire ESP32 just because the cloud server is down. Recover the failed layer first.
8. Avoid blocking network code
Even when Wi-Fi itself is healthy, blocking HTTP requests, long DNS timeouts, or a client waiting forever for a server can make the project appear disconnected.
Use explicit timeouts and bounded retries. For long-running systems, every external operation should eventually return control to the main application.
A useful rule is: network failure must degrade the network feature, not freeze the device.
9. Consider Wi-Fi power saving when latency matters
Power-saving modes can affect response timing. That does not automatically mean power saving should be disabled; it means you should choose the power/performance trade-off deliberately.
For battery products, saving energy may be more important than instant response. For real-time control or low-latency interactive systems, test whether the selected sleep behavior meets the timing requirements.
10. Reconnect dependent services in the correct order
After Wi-Fi returns, application services often need to be rebuilt.
A clean recovery sequence is:
- Wi-Fi reconnects.
- IP address becomes valid.
- Time/DNS dependencies recover if required.
- MQTT/HTTP/WebSocket client reconnects.
- Subscriptions or sessions are restored.
- Buffered sensor data is sent if your application supports it.
Do not assume an old TCP or MQTT connection is still usable after the network link has disappeared.
11. Design the device to work locally while offline
A resilient IoT product should define what happens when the internet is unavailable for minutes or hours.
Depending on the application, the ESP32 may continue to:
- Read sensors.
- Control local outputs.
- Maintain safety limits.
- Store recent measurements.
- Update a local display.
- Accept local buttons or BLE commands.
The cloud connection should enhance the system, not become a single point of failure unless the product truly requires it.
12. Stress-test recovery before deployment
Do not validate reconnection by unplugging the router once. Test the failure modes deliberately.
- Restart the access point.
- Move the ESP32 to weak-signal areas.
- Disable the SSID for several minutes.
- Change the network password and observe the failure reason.
- Disconnect the internet while keeping Wi-Fi active.
- Stop the MQTT broker or API server.
- Cycle power repeatedly.
- Run the device for days, not minutes.
Record reconnect time, disconnect reason, retry count, free heap, and application state. A fault that cannot be reproduced and measured is much harder to fix permanently.
Production checklist
- Disconnect reasons are logged.
- Reconnect attempts are non-blocking.
- Retry rate is bounded with backoff.
- RSSI is tested in the final installation location.
- 2.4 GHz compatibility is confirmed for classic ESP32 hardware.
- DHCP, DNS, and remote-service failures are distinguished.
- MQTT/HTTP sessions are rebuilt after link recovery.
- The device retains safe local behavior while offline.
- Brownouts have been ruled out during Wi-Fi transmit.
- Long-duration reconnect testing has been completed.
Related ESP32 reliability guides
- ESP32 random reboots: brownouts, pins and power fixes
- ESP32 power optimization for wireless projects