Transformed and Quantized your model with NPU of the MX8Mplus
Introduction
In this article, we will explain in this article which steps you have to take to transform and quantize your model with different TensorFlow versions. We are only looking into post training quantization.
We are using the MX8MPlus EVK run our model that have a dedicated neural network accelerator IP from VeriSilicon (Vivante VIP8000).
Press enter or click to view image in full size
i.mx8MPlus block diagram from NXP Is the neural processing unit (NPU) from NXP need a fully int8 quantized model we have to look into full int8 quantization of a TensorFlow lite or PyTorch model. Both libraries are supported with the eIQ library from NXP. Here we will only look into the TensorFlow variant.
The general overview on how to do post training quantization can be found on the TensorFlow website.
The operations for floating point are more complex than for integer (arithmetic’s, avoiding overflow). This results in the ability to use only the much simpler and smaller arithmetic units instead of the larger floating-point units.
The physical space needed for float32 operation much larger than for int8. This results in:
Lower power consumption,
Less heat development,
The ability to join more calculation units decreasing inference time.
Prerequisite
To create the code on your PC first, we recommend using Anaconda with a virtual environment running Python 3.6, TensorFlow 2.x, numpy, opencv-python and pandas.
Post training quantization with TensorFlow Version 2.x
If you created and trained a model via tf.keras there are three similar ways of quantizing the model.
First Method — Quantizing a Trained Model Directly
The trained TensorFlow model has to be converted into a TFlite model and can be directly quantize as described in the following code block. For the trained model we exemplary use the updated tf,keras_vggface model based on the work of rcmalli . The transformation starts at line 28.
from keras_vggface_TF.utils import preprocess_input
from keras_vggface_TF.vggfaceTF import VGGFace
converter.inference_output_type = tf.int8
import tensorflow as tf
import numpy as np
tfVersion=tf.version.VERSION.replace(".", "") # can be used as savename
print(tf.version.VERSION)
# TRAIN A MODEL
pretrained_model = VGGFace(model='resnet50', include_top=False, input_shape=(224, 224, 3), pooling='avg') # pooling: None, avg or max
#Create a generator for representative Dataset
folderpath='./All_croped_images/'
def prepare(img):
img = np.expand_dims(img,0).astype(np.float32)
img = preprocess_input(img, version=2)
return img
repDatagen=tf.keras.preprocessing.image.ImageDataGenerator(preprocessing_function=prepare)
datagen=repDatagen.flow_from_directory(folderpath,target_size=(224,224),batch_size=1)
def representative_dataset_gen():
for _ in range(10):
Img = datagen.next()
yield [img[0]]
#CONVERT YOUR MODEL TO TFLITE AND QUANTIZE TO INT8
converter = tf.lite.TFLiteConverter.from_keras_model(pretrained_model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset_gen
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.experimental_new_converter = True
converter.target_spec.supported_types = [tf.int8]
converter.inference_input_type = tf.int8
quantized_tflite_model = converter.convert()
#SAVE THE QUANTIZED MODEL AS TFLITE FILE
open('quant_model.tflite' , "wb").write(quantized_tflite_model)
After loading/training your model you first have to create a representative data set. The representative data set is used by the converter to get the max and min values to be able to estimate the scaling factor. This limits the error introduced by the quantization from float32 to intX. The error comes from the different number-space of float and int. Converting from float to int8 limits the number-space to integer values between -128 to 127. Calibrating the model on the dynamic range of the input limits this error.
Here you can just loop through your images or create a generator as in our example. We used the tf.keras.preprocessing.image.ImageDataGenerator() to yield images and do the necessary prepossessing on the images. As a generator you can of course also use the tf.data.Dataset.from_tensors() or … from.tensor_slices(). Just keep in mind to do the same pre-processing on your data here as you did on the data you trained your network with (normalization, resizing, de-noising, …). This can all be packed into the preprocessing_function call of the generator (line 19).
The conversion starts at line 28. A simple TensorFlow lite conversion would look like this:
converter = tf.lite.TFLiteConverter.from_keras_model(pretrained_model)
tflite_model = converter.convert()
open('model.tflite' , "wb").write(tflite_model)
The quantization part is in-between:
converter = tf.lite.TFLiteConverter.from_keras_model(pretrained_model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
converter.representative_dataset = representative_dataset_gen
converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
converter.experimental_new_converter = True
converter.target_spec.supported_types = [tf.int8]
converter.inference_input_type = tf.int8
converter.inteference_output_type=tf.int8
quantized_tflite_model = converter.convert()
open('quant_model.tflite' , "wb").write(quantized_tflite_model)
Line 3 : optimizations other than default are deprecated. No other options are available at the moment (Year 2020).
Line 4 : Here we set the representative data set.
Line 5 : Here we make sure that we have a full conversion to int8. Without this option, only the weights and biases would be converted but not the activations. This is used when we only want to reduce the model size. However, our NPU needs full int8 quantization. Having the activations still in floating point would result in overall floating-point and could not run on the NPU.
Line 6: Enables MLIR-based conversion instead of TOCO conversion, which enables RNN support, easier error tracking.
Line 7: Sets the internal constant value to int8. The target_spec corresponds with the TFLITE_BUILTINS from line 5.
Line 9 and 10: Set also the input to int8. This is fully available from TF 2.3
Now if we convert the model using TF2.3 with
experimental_new_converter=True
inference_input_type=tf.int8
inference_output_type=tf.int8
we receive the following model:
However if we do not set the inference_input_type and inference_output_type we receive following model:
So the effect is that you can determine which input data type the model accepts and returns. This can be important if you work with an embedded camera, as included with the MX8MPEVK. The mipi camera returns 8bit values, so if you want to spare a conversion to float32 int8 input can be handy. But be aware, if you use a model without prediction layers to gain e.g., embeddings, an int8 output will result in very poor performance. Here an output of float32 is recommended. This shows that each problem needs a specific solution.
Second and Third Method — Quantize a Saved Model from *.h5 or *.pb Files
If you already have your model, you most likely have it saved somewhere either as an Keras h5 file or a TensorFlow protocol buffer pb. We will quickly save our model using TF2.3:
import tensorflow as tf f
rom keras_vggface_TF.vggfaceTF import VGGFace
pretrained_model = VGGFace(model='resnet50', include_top=False, input_shape=(224, 224, 3), pooling='avg') # pooling: None, avg or max
# Saving as a protocol buffer
!mkdir -p saved_model
pretrained_model.save('saved_model/my_model')
#saving as a h5 file (if your Tensorflow Version is < 2.0 use this)
!mkdir -p keras_model
pretrained_model.save('keras_model/my_model.h5')
Following the conversion and quantization is very similar as in Method One. The only difference is how we load the model in with the converter. Either load the model and continue as in Method One.
#Load the h5 model
pretrained_model = tf.keras.models.load_model('my_model.h5')
#Then use the converter as before
converter = tf.lite.TFLiteConverter.from_keras_model(pretrained_model)
Or load the h5 model directly. When using TensorFlow version 2 and above you have to use a compatible converter:
# Or load the h5 model from file
# TensorFlow version >2.x
converter = tf.compat.v1.lite.TFLiteConverter.from_keras_model_file(saved_model_dir + h5_modelname) #works now also with TF2.x
# TensorFlowversion < 2.0
converter = tf.lite.TFLiteConverter.from_keras_model_file(saved_model_dir + h5_modelname)
If you load from a TensorFlow pb file use:
# Or load the pb file from the modelfolder
# TensorFlow version >2.x
converter = tf.lite.TFLiteConverter.from_saved_model(saved_model_dir) # the folder contains the pb file and a assets and variables folder
Converting with TensorFlow Versions below 2.0
If you want to convert a model written in TensorFlow version < 1.15.3 using Keras, not all options are available for TFlite conversion and quantization. The best way is to save the model with the TensorFlow version it was created in (e.g., rcmalli keras-vggface was trained in TF 1.13.2). I would suggest not using the “saving and freeze graph” method to create a pb file as the pb files differ between TF1 and TF2. The TFLiteConverter.from_saved_model does not work, creating quite a hassle to achieve quantization. I would suggest using the above mentioned method using Keras:
import keras
….
pretrained_model.save('my_model.h5')
Then convert and quantize your model with a TensorFlow version from 1.15.3 onward. From this version on a lot of functions where added in preparation for TF2. I suggest using the latest version. This will result in the same models presented here.
記事全体を表示