Texas Instruments’ TI TinyEngine NPU now sits inside a sub-dollar Arm Cortex-M0+ microcontroller, the MSPM0G5187, and a Cortex-M33 motor-control family, the AM13Ex. Announced on 10 March 2026, the move pushes neural-network acceleration into the cheapest, lowest-power corner of embedded design, where TinyML used to mean running everything on the CPU.
What TI announced and why it matters
According to TI’s press release, both new MCU families integrate the TinyEngine NPU, a dedicated accelerator that runs neural-network inference in parallel with the main CPU. TI says that, compared with similar MCUs without an accelerator, it lowers latency by up to 90 times and energy by more than 120 times per inference, and keeps the flash footprint small. TI also says it is integrating the NPU across its whole MCU portfolio.
Two parts headline the launch:
- MSPM0G5187: an 80 MHz Cortex-M0+ priced under US$1 in 1,000-unit quantities, in production now.
- AM13E23019 (first AM13Ex part): a Cortex-M33 real-time control MCU with the NPU and a trigonometric math accelerator, in preproduction, with more package and memory variants promised by the end of 2026.
The real news is not peak throughput. It is that a part in the cost and power class of a basic sensor MCU now has hardware for convolutional networks.
Inside the TI TinyEngine NPU: a small, fixed accelerator
The TinyEngine is not a programmable vector processor. It is a fixed-function engine tuned for convolutional neural networks (CNNs). TI’s product overview lists the layers it accelerates: generic, depthwise, pointwise and transposed convolutions, fully connected layers, and average and max pooling with batch normalization. Weights can be 8-, 4- or 2-bit, with mixed-precision configurations.

On the MSPM0G5187 product page, TI rates the NPU at 2.56 GOPS at 80 MHz, with 8- and 4-bit data paths. The surrounding chip matters just as much:
- 128 KB flash (dual bank for OTA updates) and 32 KB SRAM, both with ECC or parity
- 12-bit 1.6 Msps ADC with up to 26 external channels, a comparator and a DAC
- USB 2.0 full speed and a digital audio interface supporting I2S and TDM
- 12-channel DMA, an AES accelerator and secure key storage
- RUN 103 µA/MHz, STANDBY 1.5 µA with RTC and full SRAM retention, SHUTDOWN 88 nA
The design pattern is simple. DMA moves ADC or audio samples into SRAM, the CPU does light feature work, and the NPU runs the network while the core sleeps or services other tasks. Because the NPU and the application share the same 32 KB of SRAM, an independent analysis by IoT Digital Twin PLM points out that the peak activation tensor, not flash, is usually the first limit you hit.
Use cases: motor control, arc-fault detection and sensing
TI’s examples point to one clear sweet spot: time-series signals such as current, voltage, vibration and audio.
Arc-fault detection
TI lists an arc fault (AFCI) model among the MSPM0G5187 examples and offers a reference design, TIDA-010971, for AC arc-fault detection with edge AI in circuit breakers. A small CNN that classifies current waveforms can reduce nuisance trips that fixed thresholds struggle with.
Motor control and predictive maintenance
The AM13E23019 combines a Cortex-M33 (up to 250 MHz on TI’s current product page), 512 KB flash, 128 KB SRAM, three 12-bit ADCs, 30 PWM channels, eQEP encoder inputs and CAN FD. TI says it can keep precise real-time control loops running for up to four motors while the NPU handles adaptive control for load sensing and energy optimization. TI also claims up to 30% lower BOM cost than multi-chip designs, and a trig accelerator that is 10 times faster than CORDIC implementations.
Low-power sensing
Other TI examples include generic time-series classification, motor fault and ECG. Add the I2S/TDM interface and microamp standby, and battery-powered sensing nodes become a natural target.
Toolchain: CCStudio and TI Edge AI Studio
TI pairs the silicon with CCStudio Edge AI Studio, a free tool for selecting, training and deploying models across its embedded portfolio. TI says it ships more than 60 models and application examples. The workflow covers data collection and labelling, feature extraction, model selection and tuning, then compilation and deployment, and it supports PyTorch and ONNX. Underneath, TI’s neural network compiler turns a trained model into a library you link into a CCStudio project. TI says it can target the NPU or run the model in software on the CPU. The CCStudio IDE itself now has generative AI features for code, configuration and debugging.
TI TinyEngine NPU vs STM32N6 vs software-only TinyML
These options target different tiers. The STM32N6 is a vision-class MCU, while the TinyEngine is built for cheap time-series inference.
| Option | CPU | Acceleration | Memory | Best fit |
|---|---|---|---|---|
| TI MSPM0G5187 | Cortex-M0+, 80 MHz | TinyEngine NPU, 2.56 GOPS | 128 KB flash / 32 KB SRAM | Sub-dollar sensing, arc fault, audio triggers |
| TI AM13E23019 | Cortex-M33, up to 250 MHz | TinyEngine NPU + trig math unit | 512 KB flash / 128 KB SRAM | Multi-motor control with local fault detection |
| ST STM32N6 | Cortex-M55, 800 MHz, Helium | Neural-ART NPU, up to 600 GOPS at 1 GHz | 4.2 MB embedded RAM | Camera vision and heavier audio |
| Software-only TinyML | Any Cortex-M | None (CPU kernels) | Whatever the MCU has | Very small models, low duty cycle, existing hardware |
On peak numbers, the STM32N6’s NPU is more than 200 times faster. But it adds a camera pipeline, an ISP and megabytes of RAM, which a vibration classifier does not need. The more useful comparison for the MSPM0G5187 is with software-only inference on the same class of core. A Cortex-M0+ (Armv6-M) has no DSP or SIMD extension, so every multiply-accumulate takes several instructions and the CPU stays awake throughout. TI’s 90x and 120x figures are measured against exactly that baseline, so treat them as best-case vendor ratios rather than an independent benchmark.
What this means for embedded engineers
If your product classifies a 1D signal and must cost about a dollar and sleep at microamps, the TinyEngine NPU is worth prototyping on. A practical path:
- Start from a TI example. Open Edge AI Studio, pick the arc-fault, motor-fault or generic time-series example, and run it on an LP-MSPM0G5187 LaunchPad before collecting your own data.
- Budget SRAM first. Write down the peak activation size, input window, DMA buffers and stack, and compare them with 32 KB. Our guide to writing embedded C that uses less RAM applies directly here.
- Design for the operator list. Use depthwise-separable convolutions, small kernels, pooling and a small dense head. Any layer the NPU can’t run falls back to the CPU and eats into the speed-up.
- Train with quantization in mind. 4- and 2-bit weights save flash, but small models lose accuracy fast. Use quantization-aware training and validate on data from other machines and other conditions. Be wary of headline accuracy; see why high accuracy can hide a bad machine-learning model.
- Keep control deterministic. On the AM13Ex, run the FOC loop at top interrupt priority. Let the NPU raise flags for a supervisor to act on rather than drive the power stage directly.
- Measure, don’t trust ratios. Compile the same model for NPU and CPU, then compare latency and current over a full duty cycle on your own board.
For the wider skills picture, read what embedded engineers should learn next for edge AI on microcontrollers.
Key takeaways
- The TI TinyEngine NPU brings 2.56 GOPS of CNN acceleration to an under-US$1 Cortex-M0+ MCU.
- It targets time-series workloads: arc faults, motor faults, ECG, audio triggers. Vision is out of scope.
- The AM13Ex pairs the NPU with real-time control for up to four motors, but it is still in preproduction.
- The STM32N6 belongs to a different, vision-class tier. The TinyEngine’s real competitor is software-only TinyML on small cores.
- SRAM, operator support and quantization-aware training decide whether you actually get TI’s claimed speed-ups.
FAQ
What is the TI TinyEngine NPU?
It is TI’s own neural processing unit, built into C2000 F28P55x, AM13E230x and MSPM0G5187 microcontrollers. It runs quantized convolutional, fully connected and pooling layers in parallel with the CPU, at up to 2.56 GOPS on the MSPM0G5187.
How much does the MSPM0G5187 cost?
TI prices the MSPM0G5187 under US$1 in 1,000-unit quantities, and production quantities have been available on TI.com since the March 2026 launch.
Can the TinyEngine NPU run camera or vision models?
Not in any practical sense on the MSPM0G5187. With 32 KB of SRAM it suits small 1D and spectrogram models. For camera vision, an STM32N6-class MCU with megabytes of RAM is the better choice.
Which tools do I need to deploy a model?
Use TI’s free CCStudio Edge AI Studio to train and compile the model, then CCStudio IDE with the MSPM0 or AM13E2x SDK to integrate the generated library into your firmware. Models can come from PyTorch or ONNX.
Ready to build real edge AI and embedded skills? Explore the Educational Engineering Team courses at https://eduengteam.com/wp-content/uploads/2026/10/meta-muse-gadgets-esp32-ai-agent-diagram.jpg.
