(l-onnx-doc-DequantizeLinear)= # DequantizeLinear (l-onnx-op-dequantizelinear-28)= ## DequantizeLinear - 28 ### Version - **name**: [DequantizeLinear (GitHub)](https://github.com/onnx/onnx/blob/main/docs/Operators.md#DequantizeLinear) - **domain**: `main` - **since_version**: `28` - **function**: `False` - **support_level**: `SupportType.COMMON` - **shape inference**: `True` This version of the operator has been available **since version 28**. ### Summary The linear dequantization operator. It consumes a quantized tensor, a scale, and a zero point to compute the full-precision tensor. The dequantization formula is `y = (x - x_zero_point) * x_scale`. `x_scale` and `x_zero_point` must have the same shape, determining the quantization's granularity: a scalar for per-tensor/per-layer quantization, a 1-D tensor for per-axis quantization, or have a rank identical to the input for blocked quantization. See QuantizeLinear for details on quantization granularity. `x_zero_point` and `x` must have the same type. `x` and `y` must have the same shape. In the case of dequantizing `int32`, there's no zero point (zero point is supposed to be 0). `zero-point` is usually not used in the case of float8 and 4-bit types quantization, but the dequantization formula remains the same for consistency. The output type is determined by the attribute `output_dtype`. If `output_dtype` is not supplied then the output type is the same as `x_scale`. The output type also determines the precision of the multiplication operation. ### Attributes * **axis - INT** (default is `1`): (Optional) The axis of the dequantizing dimension of the input tensor. Used for per-axis and blocked quantization. Negative value means counting dimensions from the back. Accepted range is `[-r, r-1]` where `r = rank(input)`. * **block_size - INT** (default is `0`): (Optional) The size of the quantization block (number of times every scale is replicated). Used only for blocked quantization. The block size is a positive integer. Given `x` shape `(D0, ..., Di, ..., Dn)`, `y_scale` shape `(S0, ... Si, ...Sn)` and `axis=i`, the accepted range is `[ceil(Di/Si), ceil(Di/(Si-1))-1]` * **output_dtype - INT** (default is `0`): (Optional) The output data type. If not supplied, the output data type is inferred from `x_scale` data type (`T2`) ### Inputs Between 2 and 3 inputs. - **x** (heterogeneous) - **T1**: N-D quantized input tensor to be de-quantized. - **x_scale** (heterogeneous) - **T2**: Scale for input `x`. For per-tensor/layer dequantization the scale is a scalar, for per per-axis dequantization it is a 1-D Tensor and for blocked dequantization it has the same shape as the input, except for one dimension in which blocking is performed. - **x_zero_point** (optional, heterogeneous) - **T1**: Zero point for input `x`. Shape must match x_scale. It's optional. Zero point is 0 when it's not specified. ### Outputs - **y** (heterogeneous) - **T3**: N-D full precision output tensor. It has the same shape as input `x`. The data type is specified by the `output_dtype` attribute or, in its absence, the type of `x_scale`. ### Type Constraints * **T1** in ( `tensor(float4e2m1)`, `tensor(float6e2m3)`, `tensor(float6e3m2)`, `tensor(float8e4m3fn)`, `tensor(float8e4m3fnuz)`, `tensor(float8e5m2)`, `tensor(float8e5m2fnuz)`, `tensor(int16)`, `tensor(int2)`, `tensor(int32)`, `tensor(int4)`, `tensor(int8)`, `tensor(uint16)`, `tensor(uint2)`, `tensor(uint4)`, `tensor(uint8)` ): The type of the inputs 'x_zero_point' and 'x'. * **T2** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)`, `tensor(float8e8m0)` ): The type of the input 'x_scale'. * **T3** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)` ): The type of the output 'y'. ### Examples #### default ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], ) # scalar zero point and scale x = np.array([0, 3, 128, 255]).astype(np.uint8) x_scale = np.float32(2) x_zero_point = np.uint8(128) y = np.array([-256, -250, 0, 254], dtype=np.float32) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear", ) ``` #### _axis ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], ) # 1-D tensor zero point and scale of size equal to axis 1 of the input tensor x = np.array( [ [ [[3, 89], [34, 200], [74, 59]], [[5, 24], [24, 87], [32, 13]], [[245, 99], [4, 142], [121, 102]], ], ], dtype=np.uint8, ) x_scale = np.array([2, 4, 5], dtype=np.float32) x_zero_point = np.array([84, 24, 196], dtype=np.uint8) y = ( x.astype(np.float32) - x_zero_point.reshape(1, 3, 1, 1).astype(np.float32) ) * x_scale.reshape(1, 3, 1, 1) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_axis", ) ``` #### _e4m3fn ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.FLOAT8E4M3FN, [5], [0, 0.5, 1, 448, -104]) x_scale = np.float32(2) y = np.array([0.0, 1.0, 2.0, 896.0, -208.0], dtype=np.float32) expect( node, inputs=[x, x_scale], outputs=[y], name="test_dequantizelinear_e4m3fn", ) ``` #### _e4m3fn_float16 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.FLOAT8E4M3FN, [5], [0, 0.5, 1, 448, -104]) x_scale = np.float16(2) y = np.array([0.0, 1.0, 2.0, 896.0, -208.0], dtype=np.float16) expect( node, inputs=[x, x_scale], outputs=[y], name="test_dequantizelinear_e4m3fn_float16", ) ``` #### _e4m3fn_zero_point ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "zero_point"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.FLOAT8E4M3FN, [5], [0, 0.5, 1, 448, -104]) zero_point = make_tensor("zero_point", TensorProto.FLOAT8E4M3FN, [1], [0]) x_scale = np.float32(2) y = np.array([0.0, 1.0, 2.0, 896.0, -208.0], dtype=np.float32) expect( node, inputs=[x, x_scale, zero_point], outputs=[y], name="test_dequantizelinear_e4m3fn_zero_point", ) ``` #### _e5m2 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.FLOAT8E5M2, [5], [0, 0.5, 1, 49152, -96]) x_scale = np.float32(2) y = np.array([0.0, 1.0, 2.0, 98304.0, -192.0], dtype=np.float32) expect( node, inputs=[x, x_scale], outputs=[y], name="test_dequantizelinear_e5m2", ) ``` #### _uint16 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], ) x = np.array([30000, 31000, 32768, 33000]).astype(np.uint16) x_scale = np.float32(2) x_zero_point = np.uint16(32767) y = np.array([-5534.0, -3534.0, 2.0, 466.0], dtype=np.float32) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_uint16", ) ``` #### _int16 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], ) x = np.array([-300, -30, -1025, 1270]).astype(np.int16) x_scale = np.float32(2) x_zero_point = np.int16(-1024) y = np.array([1448.0, 1988.0, -2.0, 4588.0], dtype=np.float32) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_int16", ) ``` #### _uint4 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.UINT4, [5], [0, 1, 7, 10, 15]) x_scale = np.float32(2) x_zero_point = make_tensor("x_zero_point", TensorProto.UINT4, (1,), [1]) y = np.array([-2, 0, 12, 18, 28], dtype=np.float32) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_uint4", ) ``` #### _int4 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.INT4, [5], [0, 1, 7, -4, -8]) x_scale = np.float32(2) x_zero_point = make_tensor("x_zero_point", TensorProto.INT4, (1,), [1]) y = np.array([-2, 0, 12, -10, -18], dtype=np.float32) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_int4", ) ``` #### _uint2 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.UINT2, [4], [0, 1, 2, 3]) x_scale = np.float32(2) x_zero_point = make_tensor("x_zero_point", TensorProto.UINT2, (1,), [1]) y = np.array([-2, 0, 2, 4], dtype=np.float32) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_uint2", ) ``` #### _int2 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.INT2, [4], [0, 1, -1, -2]) x_scale = np.float32(2) x_zero_point = make_tensor("x_zero_point", TensorProto.INT2, (1,), [1]) y = np.array([-2, 0, -4, -6], dtype=np.float32) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_int2", ) ``` #### _float4e2m1 ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], axis=0, ) # scalar zero point and scale x = make_tensor("x", TensorProto.FLOAT4E2M1, [5], [0, 1, -1, 1.5, -4]) x_scale = np.float32(2) x_zero_point = make_tensor("x_zero_point", TensorProto.FLOAT4E2M1, (1,), [0]) y = np.array([0, 2, -2, 3, -8], dtype=np.float32) expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_float4e2m1", ) ``` #### _blocked ```python import numpy as np import onnx node = onnx.helper.make_node( "DequantizeLinear", inputs=["x", "x_scale", "x_zero_point"], outputs=["y"], axis=1, block_size=2, ) x = np.array( [ [ [[3, 89], [34, 200], [74, 59]], [[5, 24], [24, 87], [32, 13]], [[5, 12], [12, 33], [65, 42]], [[245, 99], [4, 142], [121, 102]], ], ], dtype=np.uint8, ) x_scale = np.array( [ [ [[3.0, 2.0], [4.0, 1.0], [2.0, 2.0]], [[5.0, 2.0], [4.0, 3.0], [5.0, 2.0]], ], ], dtype=np.float32, ) x_zero_point = np.array( [ [ [[1, 0], [0, 1], [2, 20]], [[3, 2], [4, 3], [15, 2]], ], ], dtype=np.uint8, ) # x.shape = (1, 4, 3, 2) # x_scale.shape = (1, 2, 3, 2) assert x_scale.shape == x_zero_point.shape block_axis = 1 # The block shape is [x.shape[i] // x_scale.shape[i] for i in range(len(x.shape))] = (1, 2, 1, 1) assert all( x.shape[i] == x_scale.shape[i] for i in range(len(x.shape)) if i != block_axis ) assert x.shape[block_axis] % x_scale.shape[block_axis] == 0 repeats = x.shape[block_axis] // x_scale.shape[block_axis] # Create element-wise scale and zero point x_scale_elementwise = np.repeat(x_scale, repeats=repeats, axis=block_axis) x_zero_point_elementwise = np.repeat( x_zero_point, repeats=repeats, axis=block_axis ) y = ( x.astype(np.float32) - x_zero_point_elementwise.astype(np.float32) ) * x_scale_elementwise expect( node, inputs=[x, x_scale, x_zero_point], outputs=[y], name="test_dequantizelinear_blocked", ) ``` ```{toctree} text_diff_DequantizeLinear_25_28 ``` (l-onnx-op-dequantizelinear-25)= ## DequantizeLinear - 25 ### Version - **name**: [DequantizeLinear (GitHub)](https://github.com/onnx/onnx/blob/main/docs/Operators.md#DequantizeLinear) - **domain**: `main` - **since_version**: `25` - **function**: `False` - **support_level**: `SupportType.COMMON` - **shape inference**: `True` This version of the operator has been available **since version 25**. ### Summary The linear dequantization operator. It consumes a quantized tensor, a scale, and a zero point to compute the full-precision tensor. The dequantization formula is `y = (x - x_zero_point) * x_scale`. `x_scale` and `x_zero_point` must have the same shape, determining the quantization's granularity: a scalar for per-tensor/per-layer quantization, a 1-D tensor for per-axis quantization, or have a rank identical to the input for blocked quantization. See QuantizeLinear for details on quantization granularity. `x_zero_point` and `x` must have the same type. `x` and `y` must have the same shape. In the case of dequantizing `int32`, there's no zero point (zero point is supposed to be 0). `zero-point` is usually not used in the case of float8 and 4-bit types quantization, but the dequantization formula remains the same for consistency. The output type is determined by the attribute `output_dtype`. If `output_dtype` is not supplied then the output type is the same as `x_scale`. The output type also determines the precision of the multiplication operation. ### Attributes * **axis - INT** (default is `1`): (Optional) The axis of the dequantizing dimension of the input tensor. Used for per-axis and blocked quantization. Negative value means counting dimensions from the back. Accepted range is `[-r, r-1]` where `r = rank(input)`. * **block_size - INT** (default is `0`): (Optional) The size of the quantization block (number of times every scale is replicated). Used only for blocked quantization. The block size is a positive integer. Given `x` shape `(D0, ..., Di, ..., Dn)`, `y_scale` shape `(S0, ... Si, ...Sn)` and `axis=i`, the accepted range is `[ceil(Di/Si), ceil(Di/(Si-1))-1]` * **output_dtype - INT** (default is `0`): (Optional) The output data type. If not supplied, the output data type is inferred from `x_scale` data type (`T2`) ### Inputs Between 2 and 3 inputs. - **x** (heterogeneous) - **T1**: N-D quantized input tensor to be de-quantized. - **x_scale** (heterogeneous) - **T2**: Scale for input `x`. For per-tensor/layer dequantization the scale is a scalar, for per per-axis dequantization it is a 1-D Tensor and for blocked dequantization it has the same shape as the input, except for one dimension in which blocking is performed. - **x_zero_point** (optional, heterogeneous) - **T1**: Zero point for input `x`. Shape must match x_scale. It's optional. Zero point is 0 when it's not specified. ### Outputs - **y** (heterogeneous) - **T3**: N-D full precision output tensor. It has the same shape as input `x`. The data type is specified by the `output_dtype` attribute or, in its absence, the type of `x_scale`. ### Type Constraints * **T1** in ( `tensor(float4e2m1)`, `tensor(float8e4m3fn)`, `tensor(float8e4m3fnuz)`, `tensor(float8e5m2)`, `tensor(float8e5m2fnuz)`, `tensor(int16)`, `tensor(int2)`, `tensor(int32)`, `tensor(int4)`, `tensor(int8)`, `tensor(uint16)`, `tensor(uint2)`, `tensor(uint4)`, `tensor(uint8)` ): The type of the inputs 'x_zero_point' and 'x'. * **T2** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)`, `tensor(float8e8m0)` ): The type of the input 'x_scale'. * **T3** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)` ): The type of the output 'y'. ```{toctree} text_diff_DequantizeLinear_24_28 text_diff_DequantizeLinear_24_25 ``` (l-onnx-op-dequantizelinear-24)= ## DequantizeLinear - 24 ### Version - **name**: [DequantizeLinear (GitHub)](https://github.com/onnx/onnx/blob/main/docs/Operators.md#DequantizeLinear) - **domain**: `main` - **since_version**: `24` - **function**: `False` - **support_level**: `SupportType.COMMON` - **shape inference**: `True` This version of the operator has been available **since version 24**. ### Summary The linear dequantization operator. It consumes a quantized tensor, a scale, and a zero point to compute the full-precision tensor. The dequantization formula is `y = (x - x_zero_point) * x_scale`. `x_scale` and `x_zero_point` must have the same shape, determining the quantization's granularity: a scalar for per-tensor/per-layer quantization, a 1-D tensor for per-axis quantization, or have a rank identical to the input for blocked quantization. See QuantizeLinear for details on quantization granularity. `x_zero_point` and `x` must have the same type. `x` and `y` must have the same shape. In the case of dequantizing `int32`, there's no zero point (zero point is supposed to be 0). `zero-point` is usually not used in the case of float8 and 4-bit types quantization, but the dequantization formula remains the same for consistency. The output type is determined by the attribute `output_dtype`. If `output_dtype` is not supplied then the output type is the same as `x_scale`. The output type also determines the precision of the multiplication operation. ### Attributes * **axis - INT** (default is `1`): (Optional) The axis of the dequantizing dimension of the input tensor. Used for per-axis and blocked quantization. Negative value means counting dimensions from the back. Accepted range is `[-r, r-1]` where `r = rank(input)`. * **block_size - INT** (default is `0`): (Optional) The size of the quantization block (number of times every scale is replicated). Used only for blocked quantization. The block size is a positive integer. Given `x` shape `(D0, ..., Di, ..., Dn)`, `y_scale` shape `(S0, ... Si, ...Sn)` and `axis=i`, the accepted range is `[ceil(Di/Si), ceil(Di/(Si-1))-1]` * **output_dtype - INT** (default is `0`): (Optional) The output data type. If not supplied, the output data type is inferred from `x_scale` data type (`T2`) ### Inputs Between 2 and 3 inputs. - **x** (heterogeneous) - **T1**: N-D quantized input tensor to be de-quantized. - **x_scale** (heterogeneous) - **T2**: Scale for input `x`. For per-tensor/layer dequantization the scale is a scalar, for per per-axis dequantization it is a 1-D Tensor and for blocked dequantization it has the same shape as the input, except for one dimension in which blocking is performed. - **x_zero_point** (optional, heterogeneous) - **T1**: Zero point for input `x`. Shape must match x_scale. It's optional. Zero point is 0 when it's not specified. ### Outputs - **y** (heterogeneous) - **T3**: N-D full precision output tensor. It has the same shape as input `x`. The data type is specified by the `output_dtype` attribute or, in its absence, the type of `x_scale`. ### Type Constraints * **T1** in ( `tensor(float4e2m1)`, `tensor(float8e4m3fn)`, `tensor(float8e4m3fnuz)`, `tensor(float8e5m2)`, `tensor(float8e5m2fnuz)`, `tensor(int16)`, `tensor(int32)`, `tensor(int4)`, `tensor(int8)`, `tensor(uint16)`, `tensor(uint4)`, `tensor(uint8)` ): The type of the inputs 'x_zero_point' and 'x'. * **T2** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)`, `tensor(float8e8m0)` ): The type of the input 'x_scale'. * **T3** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)` ): The type of the output 'y'. ```{toctree} text_diff_DequantizeLinear_23_28 text_diff_DequantizeLinear_23_25 text_diff_DequantizeLinear_23_24 ``` (l-onnx-op-dequantizelinear-23)= ## DequantizeLinear - 23 ### Version - **name**: [DequantizeLinear (GitHub)](https://github.com/onnx/onnx/blob/main/docs/Operators.md#DequantizeLinear) - **domain**: `main` - **since_version**: `23` - **function**: `False` - **support_level**: `SupportType.COMMON` - **shape inference**: `True` This version of the operator has been available **since version 23**. ### Summary The linear dequantization operator. It consumes a quantized tensor, a scale, and a zero point to compute the full-precision tensor. The dequantization formula is `y = (x - x_zero_point) * x_scale`. `x_scale` and `x_zero_point` must have the same shape, determining the quantization's granularity: a scalar for per-tensor/per-layer quantization, a 1-D tensor for per-axis quantization, or have a rank identical to the input for blocked quantization. See QuantizeLinear for details on quantization granularity. `x_zero_point` and `x` must have the same type. `x` and `y` must have the same shape. In the case of dequantizing `int32`, there's no zero point (zero point is supposed to be 0). `zero-point` is usually not used in the case of float8 and 4-bit types quantization, but the dequantization formula remains the same for consistency. The output type is determined by the attribute `output_dtype`. If `output_dtype` is not supplied then the output type is the same as `x_scale`. The output type also determines the precision of the multiplication operation. ### Attributes * **axis - INT** (default is `1`): (Optional) The axis of the dequantizing dimension of the input tensor. Used for per-axis and blocked quantization. Negative value means counting dimensions from the back. Accepted range is `[-r, r-1]` where `r = rank(input)`. * **block_size - INT** (default is `0`): (Optional) The size of the quantization block (number of times every scale is replicated). Used only for blocked quantization. The block size is a positive integer. Given `x` shape `(D0, ..., Di, ..., Dn)`, `y_scale` shape `(S0, ... Si, ...Sn)` and `axis=i`, the accepted range is `[ceil(Di/Si), ceil(Di/(Si-1))-1]` * **output_dtype - INT** (default is `0`): (Optional) The output data type. If not supplied, the output data type is inferred from `x_scale` data type (`T2`) ### Inputs Between 2 and 3 inputs. - **x** (heterogeneous) - **T1**: N-D quantized input tensor to be de-quantized. - **x_scale** (heterogeneous) - **T2**: Scale for input `x`. For per-tensor/layer dequantization the scale is a scalar, for per per-axis dequantization it is a 1-D Tensor and for blocked dequantization it has the same shape as the input, except for one dimension in which blocking is performed. - **x_zero_point** (optional, heterogeneous) - **T1**: Zero point for input `x`. Shape must match x_scale. It's optional. Zero point is 0 when it's not specified. ### Outputs - **y** (heterogeneous) - **T3**: N-D full precision output tensor. It has the same shape as input `x`. The data type is specified by the `output_dtype` attribute or, in its absence, the type of `x_scale`. ### Type Constraints * **T1** in ( `tensor(float4e2m1)`, `tensor(float8e4m3fn)`, `tensor(float8e4m3fnuz)`, `tensor(float8e5m2)`, `tensor(float8e5m2fnuz)`, `tensor(int16)`, `tensor(int32)`, `tensor(int4)`, `tensor(int8)`, `tensor(uint16)`, `tensor(uint4)`, `tensor(uint8)` ): The type of the inputs 'x_zero_point' and 'x'. * **T2** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)` ): The type of the input 'x_scale'. * **T3** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)` ): The type of the output 'y'. ```{toctree} text_diff_DequantizeLinear_21_28 text_diff_DequantizeLinear_21_25 text_diff_DequantizeLinear_21_24 text_diff_DequantizeLinear_21_23 ``` (l-onnx-op-dequantizelinear-21)= ## DequantizeLinear - 21 ### Version - **name**: [DequantizeLinear (GitHub)](https://github.com/onnx/onnx/blob/main/docs/Operators.md#DequantizeLinear) - **domain**: `main` - **since_version**: `21` - **function**: `False` - **support_level**: `SupportType.COMMON` - **shape inference**: `True` This version of the operator has been available **since version 21**. ### Summary The linear dequantization operator. It consumes a quantized tensor, a scale, and a zero point to compute the full-precision tensor. The dequantization formula is `y = (x - x_zero_point) * x_scale`. `x_scale` and `x_zero_point` must have the same shape, determining the quantization's granularity: a scalar for per-tensor/per-layer quantization, a 1-D tensor for per-axis quantization, or have a rank identical to the input for blocked quantization. See QuantizeLinear for details on quantization granularity. `x_zero_point` and `x` must have the same type. `x` and `y` must have the same shape. In the case of dequantizing `int32`, there's no zero point (zero point is supposed to be 0). `zero-point` is usually not used in the case of float8 types quantization, but the dequantization formula remains the same for consistency, and `x_scale` still determines the output type. ### Attributes * **axis - INT** (default is `1`): (Optional) The axis of the dequantizing dimension of the input tensor. Used for per-axis and blocked quantization. Negative value means counting dimensions from the back. Accepted range is `[-r, r-1]` where `r = rank(input)`. * **block_size - INT** (default is `0`): (Optional) The size of the quantization block (number of times every scale is replicated). Used only for blocked quantization. The block size is a positive integer. Given `x` shape `(D0, ..., Di, ..., Dn)`, `y_scale` shape `(S0, ... Si, ...Sn)` and `axis=i`, the accepted range is `[ceil(Di/Si), ceil(Di/(Si-1))-1]` ### Inputs Between 2 and 3 inputs. - **x** (heterogeneous) - **T1**: N-D quantized input tensor to be de-quantized. - **x_scale** (heterogeneous) - **T2**: Scale for input `x`. For per-tensor/layer dequantization the scale is a scalar, for per per-axis dequantization it is a 1-D Tensor and for blocked dequantization it has the same shape as the input, except for one dimension in which blocking is performed. - **x_zero_point** (optional, heterogeneous) - **T1**: Zero point for input `x`. Shape must match x_scale. It's optional. Zero point is 0 when it's not specified. ### Outputs - **y** (heterogeneous) - **T2**: N-D full precision output tensor. It has same shape as input `x`. ### Type Constraints * **T1** in ( `tensor(float8e4m3fn)`, `tensor(float8e4m3fnuz)`, `tensor(float8e5m2)`, `tensor(float8e5m2fnuz)`, `tensor(int16)`, `tensor(int32)`, `tensor(int4)`, `tensor(int8)`, `tensor(uint16)`, `tensor(uint4)`, `tensor(uint8)` ): The type of the inputs 'x_zero_point' and 'x'. * **T2** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)` ): 'x_scale' determines the output type. ```{toctree} text_diff_DequantizeLinear_19_28 text_diff_DequantizeLinear_19_25 text_diff_DequantizeLinear_19_24 text_diff_DequantizeLinear_19_23 text_diff_DequantizeLinear_19_21 ``` (l-onnx-op-dequantizelinear-19)= ## DequantizeLinear - 19 ### Version - **name**: [DequantizeLinear (GitHub)](https://github.com/onnx/onnx/blob/main/docs/Operators.md#DequantizeLinear) - **domain**: `main` - **since_version**: `19` - **function**: `False` - **support_level**: `SupportType.COMMON` - **shape inference**: `True` This version of the operator has been available **since version 19**. ### Summary The linear dequantization operator. It consumes a quantized tensor, a scale, and a zero point to compute the full precision tensor. The dequantization formula is `y = (x - x_zero_point) * x_scale`. `x_scale` and `x_zero_point` must have same shape, and can be either a scalar for per-tensor / per layer quantization, or a 1-D tensor for per-axis quantization. `x_zero_point` and `x` must have same type. `x` and `y` must have same shape. In the case of dequantizing int32, there's no zero point (zero point is supposed to be 0). `zero-point` is usually not used in the case of float8e4m3fn, float8e4m3fnuz, float8e5m2, float8e5m2fnuz quantization, but the dequantization formula remains the same for consistency and 'x_scale' still determines the output type. ### Attributes * **axis - INT** (default is `1`): (Optional) The axis of the dequantizing dimension of the input tensor. Used only for per-axis quantization. Negative value means counting dimensions from the back. Accepted range is `[-r, r-1]` where `r = rank(input)`. When the rank of the input is 1, per-tensor quantization is applied, rendering the axis unnecessary in this scenario. ### Inputs Between 2 and 3 inputs. - **x** (heterogeneous) - **T1**: N-D quantized input tensor to be de-quantized. - **x_scale** (heterogeneous) - **T2**: Scale for input 'x'. It can be a scalar, which means a per-tensor/layer dequantization, or a 1-D tensor for per-axis dequantization. - **x_zero_point** (optional, heterogeneous) - **T1**: Zero point for input 'x'. Shape must match x_scale. It's optional. Zero point is 0 when it's not specified. ### Outputs - **y** (heterogeneous) - **T2**: N-D full precision output tensor. It has same shape as input 'x'. ### Type Constraints * **T1** in ( `tensor(float8e4m3fn)`, `tensor(float8e4m3fnuz)`, `tensor(float8e5m2)`, `tensor(float8e5m2fnuz)`, `tensor(int32)`, `tensor(int8)`, `tensor(uint8)` ): Constrain 'x_zero_point' and 'x' to 8-bit integer or float, or /32-bit integer tensor. * **T2** in ( `tensor(bfloat16)`, `tensor(float)`, `tensor(float16)` ): 'x_scale' determines the output type. ```{toctree} text_diff_DequantizeLinear_13_28 text_diff_DequantizeLinear_13_25 text_diff_DequantizeLinear_13_24 text_diff_DequantizeLinear_13_23 text_diff_DequantizeLinear_13_21 text_diff_DequantizeLinear_13_19 ``` (l-onnx-op-dequantizelinear-13)= ## DequantizeLinear - 13 ### Version - **name**: [DequantizeLinear (GitHub)](https://github.com/onnx/onnx/blob/main/docs/Operators.md#DequantizeLinear) - **domain**: `main` - **since_version**: `13` - **function**: `False` - **support_level**: `SupportType.COMMON` - **shape inference**: `True` This version of the operator has been available **since version 13**. ### Summary The linear dequantization operator. It consumes a quantized tensor, a scale, and a zero point to compute the full precision tensor. The dequantization formula is `y = (x - x_zero_point) * x_scale`. `x_scale` and `x_zero_point` must have same shape, and can be either a scalar for per-tensor / per layer quantization, or a 1-D tensor for per-axis quantization. `x_zero_point` and `x` must have same type. `x` and `y` must have same shape. In the case of dequantizing int32, there's no zero point (zero point is supposed to be 0). ### Attributes * **axis - INT** (default is `1`): (Optional) The axis of the dequantizing dimension of the input tensor. Ignored for per-tensor quantization. Negative value means counting dimensions from the back. Accepted range is [-r, r-1] where r = rank(input). ### Inputs Between 2 and 3 inputs. - **x** (heterogeneous) - **T**: N-D quantized input tensor to be de-quantized. - **x_scale** (heterogeneous) - **tensor(float)**: Scale for input 'x'. It can be a scalar, which means a per-tensor/layer dequantization, or a 1-D tensor for per-axis dequantization. - **x_zero_point** (optional, heterogeneous) - **T**: Zero point for input 'x'. Shape must match x_scale. It's optional. Zero point is 0 when it's not specified. ### Outputs - **y** (heterogeneous) - **tensor(float)**: N-D full precision output tensor. It has same shape as input 'x'. ### Type Constraints * **T** in ( `tensor(int32)`, `tensor(int8)`, `tensor(uint8)` ): Constrain 'x_zero_point' and 'x' to 8-bit/32-bit integer tensor. ```{toctree} text_diff_DequantizeLinear_10_28 text_diff_DequantizeLinear_10_25 text_diff_DequantizeLinear_10_24 text_diff_DequantizeLinear_10_23 text_diff_DequantizeLinear_10_21 text_diff_DequantizeLinear_10_19 text_diff_DequantizeLinear_10_13 ``` (l-onnx-op-dequantizelinear-10)= ## DequantizeLinear - 10 ### Version - **name**: [DequantizeLinear (GitHub)](https://github.com/onnx/onnx/blob/main/docs/Operators.md#DequantizeLinear) - **domain**: `main` - **since_version**: `10` - **function**: `False` - **support_level**: `SupportType.COMMON` - **shape inference**: `True` This version of the operator has been available **since version 10**. ### Summary The linear dequantization operator. It consumes a quantized tensor, a scale, a zero point to compute the full precision tensor. The dequantization formula is y = (x - x_zero_point) * x_scale. 'x_scale' and 'x_zero_point' are both scalars. 'x_zero_point' and 'x' must have same type. 'x' and 'y' must have same shape. In the case of dequantizing int32, there's no zero point (zero point is supposed to be 0). ### Inputs Between 2 and 3 inputs. - **x** (heterogeneous) - **T**: N-D quantized input tensor to be de-quantized. - **x_scale** (heterogeneous) - **tensor(float)**: Scale for input 'x'. It's a scalar, which means a per-tensor/layer quantization. - **x_zero_point** (optional, heterogeneous) - **T**: Zero point for input 'x'. It's a scalar, which means a per-tensor/layer quantization. It's optional. 0 is the default value when it's not specified. ### Outputs - **y** (heterogeneous) - **tensor(float)**: N-D full precision output tensor. It has same shape as input 'x'. ### Type Constraints * **T** in ( `tensor(int32)`, `tensor(int8)`, `tensor(uint8)` ): Constrain 'x_zero_point' and 'x' to 8-bit/32-bit integer tensor.