Which quantization precisions are officially supported in Ara240 16GB M.2 Module ?(INT4,8,16)
According to Ara240 Discrete Neural Processing Unit Data Sheet
|
Precision |
Officially documented support |
|
INT4 |
No documented support found |
|
INT8 |
Supported |
|
INT16 |
Supported |
|
INT32 |
Supported, though not in your INT4/8/16 list |