Edge AI runs inference close to the sensor or user. It can reduce network transfers and dependence on a cloud round trip. It does not guarantee a particular latency, battery life, privacy outcome or prediction accuracy: these depend on the model, hardware and complete application.
Start with an acceptance test
For vibration monitoring, define faults, operating regimes and the cost of missed detections. For vision inspection, define defects, lighting, production speed and acceptable false rejects. For medical or safety-related use, define the intended purpose and applicable validation requirements before selecting a model.
Separate training data from validation and test data by machine, session or relevant time period. Report precision/recall, false-alarm rate and performance on new operating conditions. A result from one dataset cannot establish a universal 96% accuracy or a guaranteed maintenance warning several weeks in advance.
Hardware figures: read the units and conditions
| Platform | Published capability | Interpretation |
|---|---|---|
| ESP32-P4 | High-performance CPU cores up to 400 MHz | MHz is clock frequency, not a 400 GOPS AI throughput claim |
| STM32N6 | Neural-ART NPU up to 600 GOPS | Vendor accelerator peak; supported operators and memory traffic affect achieved throughput |
| Jetson Orin Nano | Up to 67 TOPS, configurable module modes in the 7–25 W range | Sparse INT8 peak in the appropriate configuration; not full-system power or model latency |
| NXP i.MX 8M Plus | NPU up to 2.3 TOPS | Vendor NPU peak; measure the complete model and board |
The linked vendor documentation defines these figures. TOPS/GOPS comparisons require matching numeric precision, sparsity assumptions and counting conventions. A faster peak accelerator may be slower for an unsupported operator or a pipeline dominated by camera acquisition and CPU preprocessing.
TinyML on a MCU suits suitably small models and constrained sensing tasks. Application processors and NPUs may support larger vision workloads. FPGA pipelines offer configurable parallel processing but require mapping, timing and memory design. Select by measured requirements, not a generic platform ranking.
Measure the complete pipeline
Measure acquisition, preprocessing, inference and output handling separately, then record end-to-end latency, including p95/p99 where deadlines matter. Specify model version, input dimensions, batch size, precision, clocks and temperature. Measure energy per inference and idle energy on the complete board, including sensors and communications.
Quantization can reduce model size and computation, but accuracy changes depend on calibration data and the task. There is no universal less-than-2% accuracy loss. Language-model deployment also needs RAM for runtime buffers and the KV cache, beyond stored weights; the label “small model” does not guarantee it fits an MCU.
Privacy, safety and maintenance
Local processing can help minimise data transfer, but personal-data processing remains subject to the GDPR. Logs, telemetry, retained images and remote maintenance can still disclose data. Decide what is stored, transmitted and deleted.
An anomaly detector is not automatically a validated emergency-stop function. Where safety is involved, establish the required safety architecture, independent safeguards and evidence for the intended use. Sign and authenticate model/firmware updates, keep versioned test evidence and monitor distribution drift after deployment.
AI Act scope
Risk classification follows the intended use and the operator’s role, not the location of inference. A local model does not automatically meet the AI Act or become high-risk. According to the Commission’s current timeline, general application began on 2 August 2026, with high-risk Annex III rules from 2 December 2027 and regulated-product high-risk rules from 2 August 2028. Apply the relevant category and transition provisions to the product.
Edge AI engineering · MCU selection · EU hardware legislation · Discuss a validation plan
Frequently asked questions
Does Edge AI guarantee inference below 10 ms?
No. Latency depends on the model, hardware, memory, input acquisition and preprocessing. Measure end-to-end latency on the actual product and state the test conditions.
Does ESP32-P4 deliver 400 GOPS?
The cited Espressif specification is a CPU clock of up to 400 MHz. It is not a published 400 GOPS NPU performance specification.
Can I compare all vendors’ TOPS directly?
Only after checking precision, sparsity, operation-counting conventions and configuration. Peak throughput is not a measurement of model latency or complete-system energy.
Is local inference automatically compliant with GDPR and the AI Act?
No. GDPR obligations depend on personal-data processing; AI Act obligations depend on intended use, risk classification and role. Local execution alone proves neither.
Primary sources
Technical and regulatory references checked on 1 October 2026.