YURUNOX / Embedded system reliability

What Is a Watchdog Timer? How to Select, Configure and Test One

A watchdog timer is a deadline monitor for embedded software and systems. The system must provide a valid service event before—or, for a window watchdog, within—the permitted interval. Missing or mistimed service triggers a defined action such as a reset. It helps recover from certain stalled-execution faults only when the feed represents real progress and the reset path restores the required function.

But a regular heartbeat does not prove that the application is healthy. A useful watchdog design connects three things: meaningful progress, the right timing and a recovery path that restores the product’s function.

By YURUNOX · For embedded engineers and component teams
Updated and source-checked:

STM32F4 Discovery development board with microcontroller, connectors and reset controls
A development board makes supervision testable. The important result is not just a reset—it is the return of the required function. Board shown for hardware context, not as a reported test. Photo: Teardown Central, Wikimedia Commons, CC BY-SA 2.0. Uncropped.

What Does a Watchdog Timer Detect—and What Does It Miss?

It checks whether an accepted service event arrives on time. The hardware does not usually know whether a sensor reading is correct, a control calculation is sensible or a network task is making progress. Software must make those health conditions part of the permission to feed.

Swipe to compare the decision conditions →

Condition → recommendation → evidence required → stop boundary
System conditionStarting recommendationEvidence requiredStop boundary
MCU watchdog covers the required faults and reset domainUse the internal hardware watchdog if its clock and operating-mode behavior are suitable.MCU reference manual, clock source, timeout limits, reset cause, sleep/debug behavior and configuration protection.Do not assume “internal” means independent of the main clock, supply or software configuration.
One critical task can stall while interrupts or other tasks continueGate the hardware feed on fresh task-level progress.Named tasks, milestones, maximum report age, supervisor timing and hardware fallback behavior.Stop if an unrelated interrupt or heartbeat can feed after required work has stopped.
Runaway code must be detected as well as missing serviceConsider a window watchdog.Earliest and latest legitimate feed times, scheduling jitter, clock tolerance and startup/re-enable rules.Do not use a tight window until every permitted operating mode fits its guaranteed limits.
MCU supervision shares an unacceptable failure dependencyAdd a suitable external watchdog or supervisor.Supply dependencies, WDI qualification, output type, pull-up, reset wiring, disable paths and timing limits.Stop if the fault output cannot reliably assert the intended reset or safe-state input.
Failure can move, heat, pressurize or energize a loadUse the watchdog only within the application’s risk-control architecture.Hazard analysis, required safe state, independent protections, diagnostic coverage and validated restart behavior.A general watchdog is not evidence of functional-safety compliance or a complete protective function.

The basic sequence: configure, service, expire, respond

After supervision starts, software periodically performs the required service operation. For an internal microcontroller watchdog, that may be a register sequence or instruction. For an external IC, it may be a qualifying transition on the watchdog input, often called WDI. A valid event restarts or satisfies the timer according to the device’s rules.

If acceptable service stops, the watchdog reaches its timeout and activates the configured response. Some implementations provide an interrupt or fault signal before a reset. An interrupt-only response still depends on the processor being able to execute the handler.

Basic watchdog block diagram linking processor service signal, timer clock and watchdog output to processor reset
Trace the service path toward the timer, then the reset path back to the processor. This particular diagram uses a shared clock; an independently clocked watchdog has a different clock dependency. Diagram: Lambtron, Wikimedia Commons, CC BY-SA 3.0. Unmodified.

A timeout is an event; a reset is a response. Also distinguish clearing the watchdog counter from resetting the processor. A function named “watchdog reset” may mean servicing the timer, while a boot log with the same phrase may identify why the processor restarted.

Why use one in unattended equipment?

A remote gateway, instrument or controller can otherwise remain unresponsive until someone intervenes. A watchdog puts a bound on certain lost-progress failures, provided those failures stop valid feeding. It can help restore availability, but repeated restarts are evidence to investigate—not proof that the underlying defect has been solved.

02 / Choose the supervision architecture

When Should You Use an Internal, External or Software Watchdog?

Location and timing behavior are separate properties. Internal or external describes where a hardware watchdog resides. Standard or windowed describes when it accepts service. An internal watchdog can be windowed; an external watchdog can use a simple timeout.

Swipe the table sideways to compare all columns →

What each implementation adds—and what still needs checking
ImplementationUseful capabilityDependency to examine
Internal hardwareHardware supervision without another IC.Clock, supply, reset domain and configuration protection.
External hardwareA separate supervisory device and physical fault output.Shared power faults, disable paths and reset wiring.
Software task watchdogTask-specific deadlines and richer health context.Scheduler, processor and software-timer availability.
Task checks + hardware fallbackTask-level detection with hardware escalation.Coordination and the full worst-case response time.

Choose by the failures the design must detect, not by assuming that “external” always means better or “hardware” means independent of everything.

An internal watchdog can have its own oscillator

ST’s STM32 watchdog tutorial contrasts an independent watchdog using a separate low-speed oscillator with a window watchdog derived from the main clock in its example. That distinction matters if the main clock stops.

Clock independence is only one layer. An internal watchdog can still share silicon, supply and reset logic with the processor. Review the exact MCU reference manual, including supported sleep modes and debug-freeze controls; the STM32 example is not a universal rule for all microcontrollers.

A software heartbeat may still feed hardware

A Linux daemon can service physical watchdog hardware through a driver; the daemon itself is not the hardware timer. Stopping or closing the service interface can behave differently depending on the driver and configuration. Review the Linux watchdog API documentation and test the deployed system’s behavior.

For task-level supervision, Zephyr’s task watchdog provides software channels and an optional hardware fallback. The fallback is useful when the software monitor or scheduler itself stops working. Its configuration and timing still need validation for the actual release.

03 / A deadline versus an allowed interval

When Is a Window Watchdog Better Than a Standard Timeout?

Yes, with a window watchdog. A standard timeout watchdog generally accepts a valid service event before the upper deadline. A window watchdog also has an initial closed interval: servicing too soon can trigger a fault, just as servicing too late can.

That lower limit can reveal a runaway loop that repeatedly executes the feed operation. It does not prove that the correct work happened between feeds. Even perfectly timed pulses can hide a broken application if they come from the wrong source.

Interactive timing example / not device specifications

Try the same event against two watchdog rules

Assume the last accepted feed was at 0 ms. The standard deadline is 900 ms. The windowed example accepts service strictly between 250 and 900 ms. Select the next attempted feed.

Swipe the chart sideways to see the full timeline →

Standard and window watchdog service intervals A service event at 500 milliseconds is before the standard deadline and inside the valid window from 250 to 900 milliseconds. Standard Before deadlineLate Windowed Too earlyValid intervalLate 02509001,100 ms Attempted service: 500 ms
Standard watchdog
Accepted

500 ms is before the 900 ms timeout.

Window watchdog
Accepted

500 ms is inside the allowed interval.

Each choice is a separate hypothetical trial, not a sequence of feeds. Startup rules, pulse qualification, timing tolerance and exact boundary behavior are device-specific. A late event does not undo a timeout that already occurred.

Published device example

What the TPS3430 adds beyond a nominal window

The TI TPS3430 datasheet defines the guaranteed service region using the latest lower boundary and earliest upper boundary. In its factory-programmed option with CWD unconnected and SET0 = SET1 = 0, those limits are 25.9 ms and 46.8 ms.

A hypothetical 32–38 ms service range fits between those limits, leaving 6.1 ms before its earliest event and 8.8 ms after its latest event. That is a timing comparison, not a validated design. This device uses falling WDI edges and has a separate first-pulse rule, so startup must also be checked.

Use a window only when legitimate operation can fit it reliably. If scheduling is highly variable, investigate the schedule and health policy before choosing tighter supervision.

04 / Budget for normal work and fault response

How Should You Set the Watchdog Timeout and Window?

Start with the longest legitimate gap between valid service events. Include peak workload, scheduling delay, communications retries, flash operations, startup and every permitted mode transition. An average loop time measured on an idle prototype is not that bound.

Then check two different limits. The shortest actual timeout controls false-reset risk; the longest actual timeout controls how long recovery may be delayed after valid feeding stops.

Worst valid service gap + margin < shortest actual timeout

Use limits across the specified supply and temperature range. Include clock, prescaler, programming-component and synchronization effects where applicable. Margin is an engineering choice, not a universal percentage.

Worked example: a “one-second” watchdog is not always one second

Illustrative counter model: suppose expiration takes 1,000 effective ticks, with a clock allowed to vary from 900 to 1,100 Hz. Timeout equals tick count divided by clock frequency.

Fast clock → earliest timeout
0.909 s

1,000 ticks ÷ 1,100 Hz

Slow clock → latest timeout
1.111 s

1,000 ticks ÷ 900 Hz

If valid service can be 0.600 s apart and the chosen margin is 0.100 s, the 0.700 s total remains below 0.909 s. That passes this one timing check under the assumptions. It does not establish acceptable total recovery time.

These are hypothetical values, not TPS3430 or TPS3431 specifications. Effective tick count is not necessarily the raw reload-register value. Use the selected device’s actual counter equation or specified capacitor-programmed timing limits.

Allow time for the whole recovery chain

The end-to-end response can include software health-detection delay, the maximum watchdog timeout, output propagation, reset assertion and restart or peripheral reinitialization. A short hardware timer cannot compensate for a software monitor that keeps accepting stale progress.

Fault response budget = health-detection delay + watchdog wait + reset-path delay + required recovery time

For a window watchdog, leave room on both sides: the earliest real feed must be after the latest possible opening, and the latest real feed must be before the earliest possible closing. Include scheduling jitter and clock uncertainty once each; do not overlook them or count them twice.

05 / A pulse should represent meaningful work

How Should Software Feed the Watchdog Without Hiding a Failure?

Feed only after the required work has shown fresh, acceptable progress. An unconditional timer interrupt can keep running while the main program is stuck. A hardware-generated pulse can likewise continue after the CPU has stopped doing useful work.

Microchip AN2747 specifically warns about interrupt-based clearing and widely scattered clear instructions. The practical lesson is to control the feed path and define what evidence authorizes it.

Documented software example / not a YURUNOX test

Zephyr’s sample: one thread is alive while another is stuck

In the official Zephyr task-watchdog sample, the control thread deliberately stops progressing while the main thread continues reporting activity. The documented output then identifies the control thread’s watchdog channel and resets the device.

1. Local failureThe control thread becomes stuck.
2. Other activity continuesThe main thread still runs, but cannot stand in for the failed task.
3. Task-specific responseThe expired channel is identified and the sample requests a reset.

The lesson is the coverage boundary: “some code is running” is weaker evidence than “the required tasks are progressing.” This is a published demonstration, not a field-reliability result. Hardware fallback depends on the sample’s board configuration.

A practical progress-checking pattern

  1. Give each critical task a meaningful milestone and its own allowed age or deadline.
  2. Report a fresh completion counter or timestamp after reaching that milestone—not simply on entering an interrupt.
  3. Let a supervisor verify the required reports for the current operating mode before authorizing a hardware feed.
  4. Reject stale evidence and execute the defined escalation policy when required progress disappears.

A Boolean “healthy” flag set once at startup is not enough. Reports need freshness rules, with counter wraparound and concurrent access handled correctly. Equally, a task waiting legitimately for a user or packet must not be treated as failed just because it has no new payload to process.

Illustrative gateway scenario

Network unavailable is not the same as network task deadlocked

A responsive communications task may correctly retry during a remote server outage. Resetting the gateway on every failed connection can make availability worse. Supervise whether the task can run, obey its retry policy and respond to local commands—not whether the outside network always succeeds.

By contrast, a task locked forever on a mutex should eventually fail its health check even while LEDs, timer interrupts and unrelated tasks continue.

06 / Reset is the start of recovery

What Must Recover After the Watchdog Triggers?

First determine what the reset signal actually reaches. A processor reset may leave an external modem, sensor or communication interface powered and stuck. Changing the watchdog timeout will not repair that reset-domain mismatch.

  1. 01 / PRESERVERead the reset causeCapture available cause flags before startup code clears them.
  2. 02 / REINITIALIZERestore the required stateBring the processor, affected peripherals and outputs through the intended sequence.
  3. 03 / VERIFYCheck the product functionConfirm valid inputs, working communications and permitted output behavior.

Where supported, retain a small diagnostic record with boot count, task state and firmware version. Do not assume ordinary RAM survives every reset, or that a last-minute flash write will complete during a fault.

How Do You Prevent Repeated Failures From Becoming an Endless Boot Loop?

If the same failure recurs during each startup, define an escalation path: a controlled recovery mode, a validated backup image or an inhibited state that reports the problem. The right choice depends on the product. Simply reaching the main loop does not prove useful service has returned.

Bootloaders, firmware updates and sleep transitions also need a named owner for watchdog service. Any permitted pause needs a defined duration and re-enable condition; otherwise a development workaround can silently remove supervision from production.

Which Faults and Safety Functions Still Need Separate Protection?

A watchdog monitors service activity. A brownout detector or voltage supervisor monitors supply voltage. The functions may share one IC, but they address different conditions. Check both when the system needs both, and do not assume the watchdog operates correctly outside its own specified supply range.

For equipment that can move, heat, pressurize or energize a load: define output behavior throughout the fault and restart. Reset, power-off and a safe state are not interchangeable. Use application-specific risk assessment and qualified engineering review; a general watchdog does not replace required interlocks or emergency protection.

07 / Prove the complete chain

How Do You Test the Complete Watchdog and Reset Path?

On an isolated bench, with hazardous loads disconnected or secured under an approved procedure, define the expected action and recovery deadline before injecting a fault. Test more than “stop all software and see a reboot.”

  1. Suppress valid service.Confirm that the configured hardware response occurs within its specified timing limits.
  2. Block one critical task.Leave unrelated tasks and interrupts running. Verify that they cannot mask the missing progress.
  3. Exercise the timing boundaries.For a window watchdog, test early, valid and late service events. Include the documented startup and re-enable behavior.
  4. Test operating modes and dependencies.Cover peak workloads, updates, bootloader handover, sleep and wake-up. Review clock-loss tests before performing device-specific fault injection.
  5. Verify recovery, not only reset.Check cause capture, repeated-fault handling, peripheral state and the function the user actually relies on.
Agilent 16902A logic analyzer showing its display and digital measurement equipment
A logic analyzer can correlate digital service and reset events. This is an equipment illustration, not a captured watchdog result. Photo: Daichinger, Wikimedia Commons. Public domain; unmodified.

Measure at the real pins

For an external watchdog, correlate WDI, the fault output and the processor’s reset input. A log message saying the watchdog expired does not prove the reset pulse reached the MCU.

Use an oscilloscope when voltage levels, edge shape or pulse integrity matter. Check the output’s polarity, pull-up and minimum reset-pulse requirement.

Keep evidence tied to hardware revision, firmware build, clock settings and operating conditions. Retest after changes to scheduling, boot code or reset wiring. Debugger freeze settings can hide failures: confirm behavior with the intended release configuration.

Which Checks Match Common Watchdog Symptoms?

Swipe the table sideways to compare all columns →

Start with the observed symptom
SymptomPossible explanationFirst check
Frozen application, no resetUnconditional feed or stale health report.Block critical work and trace the actual feed source.
Resets during valid workloadService gap underestimated or timing tolerance ignored.Measure worst-case gaps against the shortest timeout.
Window fault despite frequent feedsThe next event is too early.Compare actual edges with guaranteed window limits.
Fault output asserts, MCU stays runningWiring, logic level, pulse or reset-domain mismatch.Measure at the output and MCU reset input.
MCU reboots, product stays offlineExternal peripheral or recovery sequence remains faulty.Check peripheral state and end-to-end function.

These are diagnostic starting points, not confirmed causes. Extending the timeout without understanding the symptom may only postpone a real fault.

08 / Turn the requirement into a specification

Which Watchdog IC Specifications Must Match the Design?

Start with the fault coverage and reset connection. If the MCU’s internal watchdog meets those requirements, an external IC is not automatically necessary. If it does not, specify the missing capability rather than asking for the shortest available timeout.

Swipe the table sideways to compare all columns →

Two TI examples—not interchangeable parts
DeviceDocumented behaviorQuestion for the design review
TPS3431Standard programmable watchdog with enable controls and an active-low, open-drain fault output.What can disable supervision, and is the output compatible with the reset circuit?
TPS3430Window watchdog with programmable reset delay.Do normal service, startup and re-enable timing satisfy the device’s rules?

A similar package or nominal timeout does not establish pin compatibility, equivalent window behavior or the same recovery response. Engineering should approve substitutions against the complete datasheets.

  • Timing: minimum/maximum timeout, window limits, oscillator or capacitor tolerance and reset-pulse duration.
  • Electrical interface: supply range, service-edge requirements, output polarity, open-drain versus push-pull and pull-up limits.
  • Operating modes: startup delay, disable protection, bootloader ownership, low-power behavior and current consumption.

Also confirm how the watchdog is enabled, disabled and re-enabled. A programmable timeout or matching package does not prove that the service edge, pull-up network, reset pulse, low-power behavior or configuration safeguards are equivalent.

09 / Define the evidence before requesting a part or alternate

What Evidence Should Be Included in a Watchdog IC RFQ?

Send the failure-coverage and interface conditions that determine compatibility, not only a nominal timeout or package. These inputs let a sourcing team compare the requested ordering code with a proposed alternate without treating functional similarity as engineering approval.

  1. Exact device identity.Manufacturer, full ordering code, datasheet revision, approved alternates and any lifecycle or qualification requirement.
  2. Processor and operating range.MCU or processor, watchdog supply, logic and pull-up rails, temperature range, low-power modes and current limit.
  3. Supervision behavior.Internal or external architecture, standard or windowed service, required edge or pulse, enable/disable rules and fault coverage.
  4. Guaranteed timing.Minimum and maximum timeout or window boundaries, worst legitimate service gap, startup/re-enable timing and required reset-pulse duration.
  5. Reset and recovery path.Output type and polarity, pull-up, destination pin, affected reset domain, safe-state requirement and recovery deadline.
  6. Procurement evidence.Package, quantity, required date, date/lot traceability and inspection or documentation requirements.

Treat any proposed alternate as a design change until the complete supervision and recovery path has been reviewed and tested. Mark unknown operating conditions as open evidence gaps instead of filling them with assumed values.

For sourcing context, review YURUNOX’s Texas Instruments component page and quality-assurance information. Availability, suitability and any proposed alternative need confirmation for the actual enquiry.

Need a watchdog device for a specific design?

Share the exact desired part or candidate, MCU or processor, supply and temperature range, standard or windowed behavior, guaranteed timing limits, reset connection, operating modes and required quantity. Include the required date and traceability documents so the enquiry can be compared against the actual design.

Discuss your watchdog requirements →

YURUNOX is an electronic-component sourcing partner. Device suitability, fault coverage and substitutions require application-specific engineering approval.

10 / Verify behavior against the exact document revision

Which Primary Sources Support These Watchdog Decisions?

Manufacturer documentation supports device behavior; the Zephyr sample is a published software demonstration. Timing exercises and gateway scenarios are illustrative, not customer incidents or qualification results. Use the exact datasheet revision and deployed software release for implementation.

  1. Microchip AN2747 — Watchdog TimerNormal/window operation and why interrupt-based or dispersed clearing can hide failures.
  2. STMicroelectronics — Getting started with WDGSTM32 clock-source examples, debug behavior and device-specific configuration.
  3. Texas Instruments — TPS3430 datasheetSBVS366A, revised October 2021. Window timing, WDI behavior and factory-programmed limits.
  4. Texas Instruments — TPS3431 datasheetSNVSB66A, revised October 2021. Standard watchdog behavior, enable controls and open-drain output.
  5. Zephyr Project — Task WatchdogTask channels, kernel-timer dependence and optional hardware fallback.
  6. Zephyr Project — Task watchdog sampleDocumented stuck-control-thread demonstration used in the worked example.
  7. Linux Kernel — Watchdog driver APIUserspace servicing, driver behavior and configuration-dependent stopping rules.
Cart (0 items)