Hello,
We're trying to understand how ReLU6 behaves on the i.MX95 Neutron NPU.
As far as we understand, ReLU6 should always keep its output between 0 and 6. We noticed
lower detection accuracy on the NPU compared to the CPU for an SSD MobileNet V1 model, so we
started checking intermediate layer outputs. For a layer using ReLU6, the value on the board
went above 6 (up to about 25), while the same layer on the CPU (same model, same input) always
stayed within 0-6, as expected.
We didn't see any error or warning about this at runtime on the board. During model conversion,
neutron-compiler does show a couple of general warnings about quantization, but none of them
seem related to this specific point.
Could you help us understand:
1. Is it expected that a ReLU6 output can go above 6 on this NPU?
2. If not, shouldn't the compiler or the runtime show an error or a warning in that case?
Environment:
- Board: i.MX95 EVK
- BSP: LF_6.18.20_2.0.0
- eIQ Neutron SDK: 3.2.3
- Model: SSD MobileNet V1 (uint8, Arm ML-Zoo)
Thank you!
Hi,
Thank you for your interest in NXP Semiconductor products,
ReLU6 is listed as a supported operator in Neutron Supported Operators markdown, I would like to confirm if such layer wasn't changed to ReLU, could you please share the steps to reproduce and the binaries you get?
You could try eIQ Model Zoo mobilenetv1 and convert it to Neutron.
Regards
Hi,
Thanks for the quick reply.
1) Confirming it's still ReLU6, not ReLU
No, this layer was not changed to ReLU. neutron-compiler's own NeutronIR (--dump-neutron-ir-final-file)
still shows FusedActivation = "Relu6" for it - not our interpretation of the file. See
evidence/01_confirm_still_relu6/ for the exact model, command, and output.
2) Steps to reproduce, and the binaries we get
See evidence/02_reproduction_npu_exceeds_bound/ for the exact commands, the compiled model, input,
script, console log, and raw output tensor. On the i.MX95 EVK board:
min=-15 max=96
ReLU6 upper bound (raw): 12 elements above the bound: 4650 of 720000 (0.65%)
The same layer on the CPU stays within 0-6 as expected (real max=5.999, vs. NPU's real max=24.664).
3) Re: trying eIQ Model Zoo mobilenetv1
We tried it (mobilenet_v1_0.25_128_quant.tflite, from your recipe.sh). None of its 28
CONV_2D/DEPTHWISE_CONV_2D operators actually have ReLU6 as a fused activation, so this model doesn't
reproduce (or rule out) the issue we're reporting.
For reference, this all uses eIQ Neutron SDK 3.2.3 throughout (compiler and on-board runtime). We
suspect the problem is in neutron-compiler itself: compiling this layer with ReLU6 changed to plain
ReLU produced byte-for-byte identical microcode, so the upper bound doesn't seem to make it into the
generated code at all.
Could you confirm this on your end, and let us know if there's anything else you need from us?
Thank you!