Repository navigation
Implement AVX-512 intrinsics #310
Description
Activity
I think the best instruction set to get started with is probably
avx512fas it has the constructors for types that we can use for all the other sets:["AVX512F"]
-
_mm512_kunpackb -
_mm512_maskz_cvtepu32_ps -
_mm512_castpd128_pd512 -
_mm512_fixupimm_round_ps -
_mm512_mask_cosd_pd -
_mm512_mask_cvtusepi64_epi8 -
_mm_cvtsd_u32 -
_mm_mask_load_sd -
_mm512_mask_min_epu64 -
_mm512_maskz_add_round_ps -
_mm512_atan_ps -
_mm_cvt_roundss_i32 -
_mm512_mask_sqrt_round_ps -
_mm512_maskz_rsqrt14_pd -
_mm512_maskz_mul_epi32 -
_mm512_tanh_pd -
_mm512_maskz_cvt_roundepu32_ps -
_mm512_mask_permutexvar_epi32 -
_mm512_setr_ps -
_mm512_setzero_pd -
_mm512_mask_alignr_epi64 -
_mm512_set_pd -
_mm512_mask_i32gather_epi64 -
_mm512_exp10_ps -
_mm_roundscale_ss -
_mm512_maskz_cvtepu8_epi64 -
_mm512_mask_storeu_epi64 -
_mm512_maskz_cvtepi32_epi64 -
_mm512_mask_cvtepi16_epi64 -
_mm512_sllv_epi64 -
_mm_mask_cmp_round_sd_mask -
_mm512_cvtsepi64_epi16 -
_mm512_maskz_fmadd_round_ps -
_mm512_broadcast_f32x4 -
_mm512_mask_i32gather_pd -
_mm512_cmpgt_epi64_mask -
_mm512_mask_cvtepi64_storeu_epi32 -
_mm512_maskz_unpackhi_pd -
_mm512_mask_fixupimm_pd -
_mm512_mask_erfc_ps -
_mm512_mask_cvtsepi32_storeu_epi8 -
_mm512_maskz_cvtsepi32_epi16 -
_mm512_tand_pd -
_mm_mask_fnmsub_round_ss -
_mm512_mask_expm1_pd -
_mm_mask_min_round_ss -
_mm512_min_ps -
_mm512_permutexvar_ps -
_mm_mask_max_sd -
_mm512_mask_cvtepi32_epi8 -
_mm512_mask_cvtpd_epu32 -
_mm_mask_div_ss -
_mm_maskz_add_round_ss -
_mm_maskz_fmsub_round_ss -
_mm512_maskz_sra_epi64 -
_mm512_nearbyint_pd -
_mm512_min_round_ps -
_mm512_maskz_expand_epi32 -
_mm512_mask_cvtusepi32_epi16 -
_mm512_maskz_cvtepi32_epi8 -
_mm_scalef_round_sd -
_mm512_mask_cvtepi32_pd -
_mm512_mul_epi32 -
_mm_mask_rcp14_sd -
_mm512_maskz_inserti64x4 -
_mm512_trunc_pd -
_mm512_cvtepi64_epi32 -
_mm512_mask_expand_pd -
_mm512_mask_i32scatter_pd -
_mm512_maskz_unpacklo_ps -
_mm_mask_getexp_round_ss -
_mm512_castpd512_pd256 -
_mm512_maskz_loadu_pd -
_mm512_inserti64x4 -
_mm512_mask2_permutex2var_ps -
_mm_mask_fmsub_round_ss -
_mm_mask_fmadd_ss -
_mm512_mask_roundscale_round_pd -
_mm512_maskz_cvt_roundps_pd -
_mm512_mask_cmpge_epi64_mask -
_mm512_maskz_cvtusepi64_epi8 -
_mm512_maskz_cvttpd_epi32 -
_mm512_maskz_shuffle_f32x4 -
_mm512_mask_cvtusepi64_storeu_epi8 -
_mm512_mask_cmple_epu64_mask -
_mm512_mask_erfc_pd -
_mm512_floor_ps -
_mm512_cvtsepi64_epi8 -
_mm512_maskz_cvtps_epu32 -
_mm512_maskz_cvtepi32_epi16 -
_mm512_cvtusepi32_epi8 -
_mm512_sinh_ps -
_mm512_mask_permutex2var_epi64 -
_mm512_maskz_mul_epu32 -
_mm512_mask_cvtepu8_epi64 -
_mm_cvtt_roundsd_si32 -
_mm512_maskz_cvtsepi64_epi32 -
_mm512_zextps128_ps512 -
_mm512_scalef_round_pd -
_mm512_scalef_pd -
_mm_roundscale_round_sd -
_mm512_maskz_load_ps -
_mm512_maskz_sqrt_ps -
_mm512_mask_cvtps_pd -
_mm_getmant_round_sd -
_mm512_cmpge_epu64_mask -
_mm512_maskz_movedup_pd -
_mm512_rint_ps -
_mm512_fmsubadd_ps -
_mm_maskz_load_ss -
_mm_mask_cvtsd_ss -
_mm_cvt_roundi64_sd -
_mm512_cmpgt_epu64_mask -
_mm512_cmple_epu64_mask -
_mm512_cdfnorm_ps -
_mm512_mask_cmpeq_epi64_mask -
_mm512_mask_cvttpd_epu32 -
_mm512_extractf64x4_pd -
_mm512_log2_pd -
_mm512_permutex2var_epi32 -
_mm512_mask_cmpneq_epi64_mask -
_mm512_maskz_max_ps -
_mm512_maskz_getmant_round_ps -
_mm_mask_getmant_sd -
_mm_cmp_ss_mask -
_mm512_cvtps_epi32 -
_mm512_div_round_ps -
_mm512_maskz_mov_pd -
_mm512_cdfnorminv_ps -
_mm512_mask_expandloadu_ps -
_mm_maskz_fmadd_ss -
_mm512_cdfnorminv_pd -
_mm512_rem_epi32 -
_mm512_maskz_set1_epi32 -
_mm512_mask_cvtph_ps -
_mm512_cvtps_pd -
_mm512_mask_cmplt_epu64_mask -
_mm512_maskz_unpacklo_epi32 -
_mm512_maskz_expandloadu_ps -
_mm512_mask_recip_pd -
_mm512_sind_ps -
_mm512_mask_set1_epi32 -
_mm512_mask_permutex2var_pd -
_mm_mask_rsqrt14_sd -
_mm_mask_sub_ss -
_mm512_mask_cvttpd_epi32 -
_mm512_sin_pd -
_mm512_mask_srli_epi64 -
_mm512_mask_nearbyint_ps -
_mm512_mask_i64scatter_pd -
_mm_rsqrt14_sd -
_mm_fmsub_round_sd -
_mm512_mask_shuffle_i32x4 -
_mm_maskz_getexp_ss -
_mm512_permute_ps -
_mm512_maskz_sub_round_pd -
_mm512_mask_div_round_pd -
_mm_cvt_roundss_si64 -
_mm512_maskz_shuffle_i64x2 -
_mm512_cosh_ps -
_mm512_max_round_ps -
_mm512_maskz_ternarylogic_epi32 -
_mm_mask_scalef_round_ss -
_mm512_maskz_fmsub_ps -
_mm_maskz_cvtss_sd -
_mm512_rsqrt14_ps -
_mm512_cvttpd_epu32 -
_mm_mask_sqrt_ss -
_mm512_mask_cvtusepi32_storeu_epi8 -
_mm512_mask_cdfnorm_ps -
_mm512_mask_insertf32x4 -
_mm512_srav_epi64 -
_mm512_mask_i64gather_pd -
_mm512_mask_nearbyint_pd -
_mm512_maskz_sub_pd -
_mm512_max_epu64 -
_mm512_maskz_add_pd -
_mm512_mask_i64gather_ps -
_mm_maskz_max_sd -
_mm512_set1_epi32 -
_mm512_cvtepu16_epi64 -
_mm512_rem_epu64 -
_mm_mask_cmp_round_ss_mask -
_mm512_stream_si512 -
_mm512_mask_unpacklo_ps -
_mm512_fmaddsub_round_pd -
_mm_maskz_min_ss -
_mm512_maskz_cvtepi32_pd -
_mm512_div_epu16 -
_mm_mask_roundscale_round_sd -
_mm_mask_fnmadd_round_sd -
_mm512_mask_sub_epi64 -
_mm512_mask_cvt_roundpd_ps -
_mm_mask_scalef_sd -
_mm512_mask_cvt_roundph_ps -
_mm512_maskz_mov_epi64 -
_mm512_mask_unpackhi_epi64 -
_mm_maskz_fnmadd_sd -
_mm512_mask_cos_pd -
_mm_maskz_fmadd_round_ss -
_mm512_cvtps_epu32 -
_mm512_cmpneq_epi64_mask -
_mm512_set_epi16 -
_mm512_maskz_fmadd_ps -
_mm512_mask_inserti64x4 -
_mm512_maskz_cvtsepi64_epi16 -
_mm512_maskz_slli_epi64 -
_mm512_mask_rsqrt14_ps -
_mm_cvtss_u32 -
_mm512_mask_broadcast_f64x4 -
_mm512_rol_epi64 -
_mm512_maskz_cvtusepi32_epi8 -
_mm512_maskz_rol_epi64 -
_mm512_mask_test_epi64_mask -
_mm512_mask_mul_epi32 -
_mm512_mask_sqrt_ps -
_mm512_maskz_sub_ps -
_mm512_cbrt_pd -
_mm512_rem_epu8 -
_mm_mask_add_ss -
_mm_mask3_fmsub_round_sd -
_mm512_mask_cdfnorminv_ps -
_mm512_ternarylogic_epi64 -
_mm512_cvtusepi64_epi16 -
_mm512_erfc_ps -
_mm_maskz_cvtsd_ss -
_mm512_rem_epi8 -
_mm_cvtt_roundss_si32 -
_mm512_mask_cvt_roundps_ph -
_mm_mask_sub_sd -
_mm512_mask_permutexvar_epi64 -
_mm512_mask_min_pd -
_mm512_maskz_roundscale_round_pd -
_mm512_maskz_mul_ps -
_mm512_mask_insertf64x4 -
_mm512_div_epu32 -
_mm512_scalef_ps -
_mm512_maskz_getexp_round_ps -
_mm512_div_round_pd -
_mm512_tan_ps -
_mm_mask3_fnmadd_ss -
_mm512_mask_permutex_pd -
_mm512_max_epi64 -
_mm_maskz_getmant_sd -
_mm512_mask_cvtepu32_pd -
_mm512_mask_loadu_ps -
_mm512_cvtepu32_ps -
_mm512_unpacklo_epi64 -
_mm_maskz_move_sd -
_mm512_maskz_xor_epi64 -
_mm_mask_rsqrt14_ss -
_mm512_hypot_pd -
_mm512_mask_i64gather_epi32 -
_mm512_maskz_rcp14_ps -
_mm_mask_fixupimm_round_sd -
_mm512_mask_atanh_ps -
_mm512_maskz_sll_epi64 -
_mm512_maskz_rcp14_pd -
_mm512_maskz_insertf32x4 -
_mm512_cvt_roundps_pd -
_mm512_maskz_broadcastq_epi64 -
_mm512_mask_srl_epi64 -
_mm512_maskz_fmaddsub_pd -
_mm512_nearbyint_ps -
_mm_mask_fnmsub_round_sd -
_mm_cvt_roundu32_ss -
_mm_mask_fnmsub_ss -
_mm512_roundscale_round_pd -
_mm512_maskz_add_epi64 -
_mm_maskz_mul_ss -
_mm512_maskz_extractf64x4_pd -
_mm512_maskz_cvt_roundpd_epu32 -
_mm512_mask_cvtepi8_epi32 -
_mm512_maskz_expand_epi64 -
_mm_mask_fmsub_ss -
_mm512_mask_i64scatter_epi32 -
_mm_maskz_getmant_round_sd -
_mm_mask3_fnmadd_round_ss -
_mm_cvtt_roundss_u32 -
_mm512_kxnor -
_mm512_mask_rorv_epi32 -
_mm512_mask_unpacklo_epi64 -
_mm512_roundscale_pd -
_mm_mask3_fmadd_ss -
_mm_mask_sqrt_round_sd -
_mm512_mask_fixupimm_round_ps -
_mm512_mask_cvtps_ph -
_mm512_unpackhi_pd -
_mm512_cvtepi32_epi16 -
_mm512_mask_broadcast_f32x4 -
_mm512_maskz_sra_epi32 -
_mm512_castsi512_si128 -
_mm_mask_min_sd -
_mm512_mask_fixupimm_round_pd -
_mm512_cvtepi16_epi64 -
_mm512_cvttps_epu32 -
_mm512_mask_sra_epi32 -
_mm512_mask_scalef_pd -
_mm512_set1_epi16 -
_mm512_fixupimm_pd -
_mm512_setr_epi32 -
_mm512_maskz_srli_epi32 -
_mm512_cvtsepi32_epi16 -
_mm512_maskz_srav_epi32 -
_mm512_mask_cvttps_epi32 -
_mm512_maskz_scalef_ps -
_mm_cvt_roundss_u32 -
_mm512_rsqrt14_pd -
_mm_fmsub_round_ss -
_mm512_moveldup_ps -
_mm512_sra_epi64 -
_mm512_mask_scalef_ps -
_mm_mask_rcp14_ss -
_mm_maskz_div_ss -
_mm512_broadcastsd_pd -
_mm512_shuffle_f64x2 -
_mm512_set1_pd -
_mm_maskz_fnmsub_sd -
_mm512_mask_cvtusepi64_storeu_epi16 -
_mm512_cmplt_epi32_mask -
_mm_cvt_roundsd_ss -
_mm512_mask_unpacklo_epi32 -
_mm512_invsqrt_ps -
_mm512_cosh_pd -
_mm512_mask_extracti64x4_epi64 -
_mm_maskz_roundscale_sd -
_mm_maskz_min_round_sd -
_mm512_mask_acos_ps -
_mm512_mask_permute_pd -
_mm512_exp2_ps -
_mm512_maskz_srai_epi32 -
_mm512_maskz_sllv_epi64 -
_mm512_mask_rolv_epi64 -
_mm512_mask_ceil_pd -
_mm512_mask_div_pd -
_mm512_mask_cvtepi64_epi32 -
_mm512_srl_epi32 -
_mm512_unpacklo_epi32 -
_mm512_cvttps_epi32 -
_mm_mask_fixupimm_ss -
_mm512_maskz_broadcastd_epi32 -
_mm512_mask_rol_epi32 -
_mm512_i64scatter_pd -
_mm512_ternarylogic_epi32 -
_mm512_maskz_srlv_epi64 -
_mm512_mask_unpacklo_pd -
_mm512_mask_cvtsepi64_epi8 -
_mm512_mask_acos_pd -
_mm512_mask_exp2_ps -
_mm512_maskz_srl_epi64 -
_mm_cvtss_u64 -
_mm_fixupimm_sd -
_mm512_srl_epi64 -
_mm512_shuffle_i64x2 -
_mm512_broadcastq_epi64 -
_mm_maskz_getexp_round_ss -
_mm_rcp14_ss -
_mm_maskz_sub_round_sd -
_mm512_cos_pd -
_mm_maskz_fnmadd_ss -
_mm512_mask_ror_epi64 -
_mm512_maskz_div_pd -
_mm_maskz_fmsub_sd -
_mm512_mask_shuffle_f64x2 -
_mm512_zextpd128_pd512 -
_mm512_mask_ror_epi32 -
_mm512_maskz_cvtusepi64_epi16 -
_mm512_recip_pd -
_mm512_storeu_ps -
_mm512_maskz_getmant_round_pd -
_mm512_cvtpd_epi32 -
_mm_mask_mul_round_ss -
_mm512_ror_epi64 -
_mm_getexp_round_sd -
_mm512_rem_epi64 -
_mm512_maskz_shuffle_i32x4 -
_mm_maskz_fnmadd_round_sd -
_mm512_maskz_cvtt_roundps_epi32 -
_mm512_mask_div_epi32 -
_mm512_shuffle_f32x4 -
_mm512_maskz_roundscale_ps -
_mm512_cmplt_epu64_mask -
_mm512_maskz_sllv_epi32 -
_mm512_mask_cosh_pd -
_mm512_mask_sqrt_round_pd -
_mm512_maskz_cvtt_roundpd_epu32 -
_mm512_asinh_pd -
_mm512_mask_cos_ps -
_mm512_castps128_ps512 -
_mm512_maskz_alignr_epi64 -
_mm512_mask_cdfnorminv_pd -
_mm512_maskz_min_pd -
_mm512_maskz_min_round_pd -
_mm_cvt_roundss_si32 -
_mm512_maskz_fmadd_pd -
_mm_cvt_roundsd_si64 -
_mm512_maskz_broadcast_i32x4 -
_mm512_div_ps -
_mm512_mask_div_epu32 -
_mm_cvtt_roundss_u64 -
_mm512_mask_exp_pd -
_mm512_ceil_ps -
_mm512_mask3_fmsubadd_pd -
_mm512_maskz_cvt_roundps_epi32 -
_mm_sub_round_ss -
_mm_maskz_mul_round_sd -
_mm512_maskz_div_round_ps -
_mm512_i32scatter_epi64 -
_mm512_mask_cvtusepi64_epi32 -
_mm512_maskz_mov_epi32 -
_mm512_setzero -
_mm512_ceil_pd -
_mm512_mask_fmaddsub_ps -
_mm_mask_cmp_ss_mask -
_mm512_mask_exp10_pd -
_mm512_cmplt_epi64_mask -
_mm_cvtu32_ss -
_mm512_mask_sinh_ps -
_mm512_mask_max_round_ps -
_mm_maskz_div_round_ss -
_mm512_mask_loadu_epi32 -
_mm512_roundscale_round_ps -
_mm_mask_fmadd_round_ss -
_mm512_mask_floor_ps -
_mm512_mask_expandloadu_epi64 -
_mm_mask_div_sd -
_mm512_cbrt_ps -
_mm512_mask_cmplt_epi64_mask -
_mm512_maskz_extracti64x4_epi64 -
_mm512_maskz_abs_epi64 -
_mm512_mask_broadcastsd_pd -
_mm512_min_round_pd -
_mm_cvttss_i32 -
_mm512_acos_pd -
_mm512_broadcast_f64x4 -
_mm512_atan2_pd -
_mm512_fmaddsub_pd -
_mm_mask_roundscale_round_ss -
_mm512_rolv_epi64 -
_mm512_maskz_cvtpd_epi32 -
_mm512_fmsubadd_pd -
_mm512_maskz_rsqrt14_ps -
_mm512_cmpeq_epu64_mask -
_mm_rcp14_sd -
_mm512_svml_round_pd -
_mm_mask_getmant_ss -
_mm_maskz_scalef_round_sd -
_mm512_permutex2var_ps -
_mm512_i64scatter_ps -
_mm_mask_sub_round_sd -
_mm512_castsi512_si256 -
_mm512_mask_cvt_roundepi32_ps -
_mm512_mask_sra_epi64 -
_mm512_mask_cvtepu8_epi32 -
_mm_mask_max_round_sd -
_mm512_sqrt_round_pd -
_mm512_maskz_fnmsub_round_pd -
_mm512_maskz_load_epi32 -
_mm_mask3_fnmadd_sd -
_mm512_cvtepu32_pd -
_mm512_mask_cbrt_ps -
_mm512_mask_cvtepi32_storeu_epi8 -
_mm512_cvtusepi32_epi16 -
_mm512_mask_extractf32x4_ps -
_mm_mask_scalef_round_sd -
_mm512_mask_log1p_ps -
_mm_maskz_fmsub_ss -
_mm_maskz_fnmadd_round_ss -
_mm512_maskz_fmsubadd_ps -
_mm512_mask_cvtsepi32_epi8 -
_mm512_mask_cvtepi32_ps -
_mm512_mask_mullox_epi64 -
_mm512_maskz_min_round_ps -
_mm512_cvtusepi64_epi8 -
_mm512_cvtepu8_epi64 -
_mm512_mask_shuffle_f32x4 -
_mm_maskz_roundscale_round_sd -
_mm512_fmaddsub_ps -
_mm_maskz_sub_sd -
_mm512_extracti32x4_epi32 -
_mm512_mask_compress_epi64 -
_mm_maskz_roundscale_round_ss -
_mm512_maskz_permutex2var_ps -
_mm512_rorv_epi64 -
_mm512_cvtepi16_epi32 -
_mm_maskz_max_ss -
_mm512_mask_erf_ps -
_mm512_div_pd -
_mm512_mask_rorv_epi64 -
_mm512_cvt_roundps_epi32 -
_mm512_erf_pd -
_mm512_permutex2var_epi64 -
_mm512_setzero_ps -
_mm512_mask3_fmsubadd_ps -
_mm_mask_min_round_sd -
_mm512_maskz_insertf64x4 -
_mm512_mask_cmpgt_epu64_mask -
_mm512_mask_cvtsepi64_storeu_epi32 -
_mm512_mask_rolv_epi32 -
_mm512_mask_fmsubadd_pd -
_mm_mask_fnmadd_round_ss -
_mm512_knot -
_mm512_maskz_cvtepu16_epi32 -
_mm_mask3_fnmsub_round_sd -
_mm512_castps256_ps512 -
_mm512_maskz_fmaddsub_ps -
_mm_mask3_fnmadd_round_sd -
_mm_maskz_rcp14_sd -
_mm512_maskz_cvtusepi64_epi32 -
_mm512_maskz_and_epi32 -
_mm512_cvt_roundph_ps -
_mm512_maskz_unpackhi_epi32 -
_mm512_mask_rem_epu32 -
_mm512_kand -
_mm512_mask_log1p_pd -
_mm512_maskz_abs_epi32 -
_mm_cvttss_i64 -
_mm512_permutexvar_epi64 -
_mm512_mask_cvtepi64_storeu_epi8 -
_mm_maskz_sub_round_ss -
_mm_cmp_round_sd_mask -
_mm512_ror_epi32 -
_mm_maskz_getmant_ss -
_mm512_loadu_pd -
_mm512_mask_min_epi64 -
_mm512_maskz_compress_epi64 -
_mm512_scalef_round_ps -
_mm512_logb_ps -
_mm_mask_move_ss -
_mm512_mask_asin_ps -
_mm512_testn_epi32_mask -
_mm512_mask_cvtsepi64_epi16 -
_mm512_mask_fmaddsub_round_pd -
_mm_maskz_cvt_roundsd_ss -
_mm512_cvtepi64_epi16 -
_mm512_mask_roundscale_ps -
_mm512_maskz_cvtepu8_epi32 -
_mm_maskz_rcp14_ss -
_mm512_srai_epi64 -
_mm512_maskz_loadu_epi32 -
_mm512_maskz_extracti32x4_epi32 -
_mm512_pow_pd -
_mm512_maskz_fnmadd_round_pd -
_mm512_loadu_ps -
_mm_mask_mul_sd -
_mm512_max_ps -
_mm_maskz_max_round_sd -
_mm512_maskz_cvtps_ph -
_mm512_maskz_cvtusepi32_epi16 -
_mm512_maskz_div_round_pd -
_mm512_maskz_ror_epi64 -
_mm_maskz_rsqrt14_ss -
_mm512_mask_i64gather_epi64 -
_mm512_maskz_roundscale_round_ps -
_mm512_maskz_cvtt_roundps_epu32 -
_mm512_log1p_pd -
_mm512_mask_mul_epu32 -
_mm512_maskz_rolv_epi32 -
_mm512_maskz_mul_pd -
_mm512_permutex_epi64 -
_mm512_cvt_roundpd_ps -
_mm512_mask_i64scatter_epi64 -
_mm512_mask_permute_ps -
_mm_fnmadd_round_sd -
_mm_cvt_roundsd_u64 -
_mm512_cvtsepi32_epi8 -
_mm512_mask_srav_epi64 -
_mm512_mask_cvtepi32_storeu_epi16 -
_mm512_cmpneq_epu64_mask -
_mm512_maskz_slli_epi32 -
_mm512_unpacklo_pd -
_mm512_maskz_compress_epi32 -
_mm512_mask_svml_round_pd -
_mm512_maskz_expandloadu_epi64 -
_mm512_mask_erfcinv_pd -
_mm512_set_epi8 -
_mm512_maskz_fnmadd_ps -
_mm512_maskz_sqrt_round_pd -
_mm512_mask_cvtusepi32_storeu_epi16 -
_mm_maskz_cvt_roundss_sd -
_mm512_mask_shuffle_i64x2 -
_mm512_maskz_min_epu64 -
_mm512_mask_asinh_pd -
_mm_mask_roundscale_ss -
_mm512_mask_fmsubadd_round_ps -
_mm512_broadcastd_epi32 -
_mm512_mask_storeu_pd -
_mm512_set4_pd -
_mm512_mask_sincos_pd -
_mm512_cosd_ps -
_mm_cvtt_roundsd_u64 -
_mm512_tanh_ps -
_mm512_mask_cvtt_roundpd_epi32 -
_mm512_mask_abs_epi64 -
_mm512_i64gather_epi32 -
_mm512_recip_ps -
_mm_cvti64_ss -
_mm_mask_store_sd -
_mm512_maskz_broadcast_f64x4 -
_mm512_maskz_unpacklo_pd -
_mm512_maskz_cvt_roundpd_ps -
_mm_cvttss_u64 -
_mm512_maskz_movehdup_ps -
_mm512_maskz_srlv_epi32 -
_mm512_castsi128_si512 -
_mm512_cvt_roundpd_epu32 -
_mm512_log_ps -
_mm512_maskz_moveldup_ps -
_mm512_i64gather_epi64 -
_mm512_div_epi64 -
_mm512_mask3_fmaddsub_ps -
_mm512_maskz_permutexvar_epi32 -
_mm512_maskz_cvtepi16_epi32 -
_mm512_atanh_pd -
_mm512_mask_erfinv_ps -
_mm512_mask_i32scatter_epi64 -
_mm512_loadu_si512 -
_mm_maskz_add_ss -
_mm512_maskz_cvt_roundpd_epi32 -
_mm512_log10_ps -
_mm512_mask3_fmsubadd_round_ps -
_mm512_erfcinv_pd -
_mm512_unpackhi_ps -
_mm512_mask_acosh_ps -
_mm512_mask_loadu_epi64 -
_mm512_rem_epu16 -
_mm512_cmp_epi64_mask -
_mm512_set4_epi64 -
_mm512_mask_cvtusepi64_epi16 -
_mm512_maskz_xor_epi32 -
_mm512_maskz_fixupimm_round_ps -
_mm_mask_div_round_ss -
_mm_mul_round_ss -
_mm512_set_ps -
_mm512_maskz_fmsub_pd -
_mm512_mask_max_round_pd -
_mm_cvtt_roundsd_i32 -
_mm512_acosh_ps -
_mm512_mask_cvtepu32_ps -
_mm512_cvtepu32_epi64 -
_mm_fixupimm_round_sd -
_mm_cvt_roundu64_sd -
_mm512_tand_ps -
_mm512_mask_slli_epi64 -
_mm512_mask_erf_pd -
_mm512_mask_log_ps -
_mm512_kortestc -
_mm512_mask_cmple_epi64_mask -
_mm512_maskz_getmant_pd -
_mm_mask_getexp_round_sd -
_mm512_cvtepi8_epi64 -
_mm512_add_epi64 -
_mm512_mask_expandloadu_epi32 -
_mm512_asin_pd -
_mm_maskz_roundscale_ss -
_mm512_sinh_pd -
_mm512_maskz_mullo_epi32 -
_mm_mask_fmadd_round_sd -
_mm512_i32gather_epi64 -
_mm512_cvt_roundps_epu32 -
_mm512_mask_rint_pd -
_mm512_mask_logb_pd -
_mm512_mask_asinh_ps -
_mm512_trunc_ps -
_mm512_mask_roundscale_pd -
_mm_cvt_roundi32_ss -
_mm512_maskz_sub_epi32 -
_mm_maskz_move_ss -
_mm512_unpackhi_epi64 -
_mm512_mask_permutexvar_ps -
_mm512_setr4_pd -
_mm_cvtt_roundsd_si64 -
_mm512_mask_broadcast_i64x4 -
_mm512_maskz_cvtt_roundpd_epi32 -
_mm_div_round_sd -
_mm512_mask_expm1_ps -
_mm_cvttsd_u64 -
_mm_min_round_sd -
_mm512_mask_cvt_roundpd_epu32 -
_mm512_mask_acosh_pd -
_mm512_log_pd -
_mm512_maskz_scalef_round_pd -
_mm512_maskz_cvtps_pd -
_mm_cvti32_ss -
_mm512_maskz_min_epi64 -
_mm512_mask_unpackhi_ps -
_mm512_testn_epi64_mask -
_mm512_test_epi64_mask -
_mm512_maskz_broadcast_f32x4 -
_mm_mask3_fmadd_round_sd -
_mm_maskz_sqrt_round_ss -
_mm512_cos_ps -
_mm_fmadd_round_sd -
_mm512_maskz_srli_epi64 -
_mm512_logb_pd -
_mm512_mask_scalef_round_ps -
_mm512_maskz_and_epi64 -
_mm512_i32scatter_pd -
_mm512_maskz_cvtepu16_epi64 -
_mm512_mask_srl_epi32 -
_mm512_cvt_roundepu32_ps -
_mm512_mask_rol_epi64 -
_mm512_mask_fmaddsub_round_ps -
_mm512_mask_cvt_roundps_epi32 -
_mm_cvtu32_sd -
_mm_maskz_load_sd -
_mm512_undefined_ps -
_mm512_mask_log10_pd -
_mm512_i64scatter_epi64 -
_mm512_fixupimm_ps -
_mm512_maskz_cvttps_epu32 -
_mm512_mask_cvtusepi32_epi8 -
_mm512_mask_logb_ps -
_mm_mask_div_round_sd -
_mm512_cvtt_roundpd_epu32 -
_mm512_extracti64x4_epi64 -
_mm_mask_roundscale_sd -
_mm512_exp_pd -
_mm_maskz_fnmsub_ss -
_mm512_maskz_permutex_pd -
_mm512_shuffle_pd -
_mm512_mask_div_ps -
_mm512_sub_epi64 -
_mm512_mask_tan_ps -
_mm512_maskz_fnmadd_pd -
_mm512_mask_cvtsepi64_epi32 -
_mm512_cvtepi32_ps -
_mm512_mask_sind_ps -
_mm512_cvtt_roundps_epi32 -
_mm512_maskz_max_epu64 -
_mm512_int2mask -
_mm_mask_fmsub_round_sd -
_mm512_cvtpd_epu32 -
_mm512_stream_pd -
_mm512_mask_sllv_epi64 -
_mm512_maskz_unpackhi_epi64 -
_mm512_mask2_permutex2var_epi32 -
_mm512_mask_ternarylogic_epi64 -
_mm512_maskz_cvtepi32_ps -
_mm512_maskz_rol_epi32 -
_mm512_expm1_pd -
_mm512_maskz_fmadd_round_pd -
_mm512_maskz_getexp_ps -
_mm512_kxor -
_mm512_mask_invsqrt_pd -
_mm_cvtt_roundss_si64 -
_mm512_mask_abs_epi32 -
_mm512_mask_permutevar_pd -
_mm_cvt_roundsd_i32 -
_mm_fmadd_round_ss -
_mm_add_round_ss -
_mm512_maskz_permute_ps -
_mm_mask3_fnmsub_ss -
_mm512_maskz_cvtepi8_epi32 -
_mm_mask_load_ss -
_mm512_cmple_epi64_mask -
_mm_mask3_fnmsub_round_ss -
_mm512_permutexvar_pd -
_mm_cvt_roundss_i64 -
_mm512_kortestz -
_mm512_maskz_min_epu32 -
_mm512_maskz_andnot_epi64 -
_mm_fnmsub_round_sd -
_mm512_mask_tand_ps -
_mm512_mask3_fmaddsub_pd -
_mm512_mask_fmaddsub_pd -
_mm512_erfinv_ps -
_mm_maskz_fnmsub_round_sd -
_mm512_maskz_fmsub_round_pd -
_mm512_mask_hypot_pd -
_mm512_maskz_cvtsepi32_epi8 -
_mm512_mask_scalef_round_pd -
_mm512_mask_pow_ps -
_mm512_permutevar_ps -
_mm512_mask_srlv_epi64 -
_mm_cmp_sd_mask -
_mm512_cvtt_roundps_epu32 -
_mm512_castps512_ps256 -
_mm512_mask_fmsubadd_round_pd -
_mm512_exp10_pd -
_mm_maskz_getexp_round_sd -
_mm512_mask_broadcastq_epi64 -
_mm512_sll_epi32 -
_mm512_maskz_permutevar_ps -
_mm_cvt_roundsi32_ss -
_mm_cvt_roundss_u64 -
_mm512_mask_cvtusepi64_storeu_epi32 -
_mm_mask_fmadd_sd -
_mm_mask3_fmsub_round_ss -
_mm512_mask_erfcinv_ps -
_mm512_mask_ternarylogic_epi32 -
_mm_mask_fixupimm_round_ss -
_mm_getmant_sd -
_mm512_mask_extractf64x4_pd -
_mm512_mask_cvtepi16_epi32 -
_mm512_maskz_mov_ps -
_mm512_mask_i64scatter_ps -
_mm512_mask_compress_ps -
_mm512_min_epi64 -
_mm512_cvtepi64_epi8 -
_mm512_atanh_ps -
_mm_mask3_fmadd_round_ss -
_mm_maskz_min_round_ss -
_mm_cvtt_roundss_i64 -
_mm512_maskz_inserti32x4 -
_mm_mask3_fmadd_sd -
_mm_scalef_ss -
_mm_maskz_fnmsub_round_ss -
_mm512_mask_moveldup_ps -
_mm512_mask_max_ps -
_mm_roundscale_sd -
_mm512_mask_expand_ps -
_mm512_maskz_cvtpd_ps -
_mm512_i32gather_pd -
_mm512_maskz_expandloadu_pd -
_mm512_mask_trunc_pd -
_mm512_kor -
_mm512_unpackhi_epi32 -
_mm512_mask_expand_epi64 -
_mm512_maskz_cvtepi8_epi64 -
_mm_cvtsd_u64 -
_mm512_mask_cbrt_pd -
_mm512_maskz_permute_pd -
_mm_sub_round_sd -
_mm_comi_round_ss -
_mm_cvtt_roundsd_i64 -
_mm512_mask2_permutex2var_epi64 -
_mm512_maskz_permutexvar_pd -
_mm_add_round_sd -
_mm512_cvtepu16_epi32 -
_mm512_srlv_epi64 -
_mm512_maskz_permutevar_pd -
_mm512_maskz_loadu_epi64 -
_mm_div_round_ss -
_mm512_movedup_pd -
_mm512_mask_srai_epi64 -
_mm512_mask_cvt_roundpd_epi32 -
_mm512_setr_epi64 -
_mm512_invsqrt_pd -
_mm512_maskz_cvtepi64_epi32 -
_mm512_mask_fixupimm_ps -
_mm_mul_round_sd -
_mm512_mask_ceil_ps -
_mm512_setr4_epi64 -
_mm512_shuffle_i32x4 -
_mm512_mask_cvtt_roundps_epu32 -
_mm512_set4_ps -
_mm512_maskz_permutexvar_ps -
_mm512_mask_testn_epi64_mask -
_mm512_mask_cvtsepi32_epi16 -
_mm512_maskz_cvtepi16_epi64 -
_mm512_mask_inserti32x4 -
_mm512_maskz_srai_epi64 -
_mm_cvtsd_i64 -
_mm_cvt_roundsi64_ss -
_mm512_maskz_add_round_pd -
_mm512_setzero_si512 -
_mm512_sra_epi32 -
_mm_getmant_round_ss -
_mm512_mask_extracti32x4_epi32 -
_mm512_maskz_rorv_epi32 -
_mm512_mask_cmp_epi64_mask -
_mm512_maskz_fixupimm_ps -
_mm512_maskz_fnmadd_round_ps -
_mm512_exp2_pd -
_mm512_maskz_max_epi64 -
_mm512_mask3_fmaddsub_round_ps -
_mm512_mask_shuffle_ps -
_mm_fixupimm_ss -
_mm_mask_max_ss -
_mm512_abs_epi64 -
_mm512_zextsi256_si512 -
_mm_mask_scalef_ss -
_mm512_maskz_load_pd -
_mm_maskz_div_sd -
_mm512_acosh_pd -
_mm512_mask3_fmaddsub_round_pd -
_mm_maskz_fixupimm_sd -
_mm512_sqrt_round_ps -
_mm512_mask_atan_ps -
_mm_maskz_getexp_sd -
_mm512_mask_cvtps_epu32 -
_mm512_cvtepi32_epi64 -
_mm512_rorv_epi32 -
_mm512_mask_asin_pd -
_mm512_extractf32x4_ps -
_mm512_exp_ps -
_mm512_rint_pd -
_mm512_kandn -
_mm_cvt_roundsd_i64 -
_mm512_sind_pd -
_mm512_mask_cvtepi64_epi8 -
_mm512_maskz_expand_ps -
_mm_mask_fmsub_sd -
_mm512_mask_sll_epi64 -
_mm512_maskz_fnmsub_round_ps -
_mm512_tan_pd -
_mm512_maskz_cvt_roundph_ps -
_mm512_maskz_add_ps -
_mm512_alignr_epi64 -
_mm512_mask_cvtpd_ps -
_mm512_mullox_epi64 -
_mm512_mask_set1_epi64 -
_mm_maskz_div_round_sd -
_mm512_mask_cvt_roundepu32_ps -
_mm512_maskz_broadcastss_ps -
_mm512_mask_tanh_ps -
_mm512_mask_cosh_ps -
_mm512_maskz_permutex_epi64 -
_mm512_cvtepi8_epi32 -
_mm_maskz_rsqrt14_sd -
_mm_cvttsd_u32 -
_mm_cvtss_i32 -
_mm512_asin_ps -
_mm512_mask_tan_pd -
_mm_mask_fnmadd_ss -
_mm_mask_mul_ss -
_mm_maskz_sqrt_ss -
_mm512_permutex2var_pd -
_mm_scalef_round_ss -
_mm512_maskz_mul_round_pd -
_mm_maskz_sub_ss -
_mm512_maskz_max_round_ps -
_mm512_maskz_load_epi64 -
_mm_cvt_roundsd_si32 -
_mm_maskz_scalef_sd -
_mm512_abs_epi32 -
_mm_getexp_sd -
_mm512_mask_storeu_epi32 -
_mm512_maskz_fmsubadd_round_ps -
_mm512_maskz_andnot_epi32 -
_mm512_castps512_ps128 -
_mm512_maskz_expandloadu_epi32 -
_mm_cvttss_u32 -
_mm_maskz_scalef_round_ss -
_mm512_floor_pd -
_mm_fnmadd_round_ss -
_mm512_maskz_fmsubadd_pd -
_mm512_mask_cvt_roundps_epu32 -
_mm_mask_add_round_ss -
_mm_mask_cvt_roundss_sd -
_mm512_maskz_cvtepu32_epi64 -
_mm512_maskz_roundscale_pd -
_mm512_acos_ps -
_mm512_mask_add_epi64 -
_mm512_cvtps_ph -
_mm512_mask_rcp14_pd -
_mm512_maskz_fnmsub_pd -
_mm512_sll_epi64 -
_mm512_sin_ps -
_mm512_maskz_max_epi32 -
_mm512_mask_rsqrt14_pd -
_mm512_maskz_scalef_round_ps -
_mm_cvtu64_ss -
_mm512_mask_testn_epi32_mask -
_mm512_maskz_scalef_pd -
_mm_fnmsub_round_ss -
_mm_rsqrt14_ss -
_mm512_mask_unpackhi_epi32 -
_mm512_mask_cvtepi64_storeu_epi16 -
_mm_cvtu64_sd -
_mm512_maskz_cvttps_epi32 -
_mm512_setzero_epi32 -
_mm512_div_epi8 -
_mm512_fmaddsub_round_ps -
_mm512_maskz_sqrt_pd -
_mm512_cvt_roundps_ph -
_mm512_mask_cmp_epu64_mask -
_mm512_stream_ps -
_mm_maskz_min_sd -
_mm_maskz_getmant_round_ss -
_mm512_maskz_add_epi32 -
_mm512_maskz_loadu_ps -
_mm512_mask_log10_ps -
_mm512_mask_atan_pd -
_mm_getexp_ss -
_mm512_storeu_pd -
_mm512_broadcastss_ps -
_mm512_log10_pd -
_mm512_mask_recip_ps -
_mm512_maskz_broadcast_i64x4 -
_mm512_zextps256_ps512 -
_mm512_maskz_fmsubadd_round_pd -
_mm512_slli_epi64 -
_mm512_maskz_sqrt_round_ps -
_mm512_maskz_permutexvar_epi64 -
_mm_maskz_fixupimm_round_sd -
_mm512_maskz_getexp_round_pd -
_mm512_mask_compressstoreu_pd -
_mm_cvt_roundu64_ss -
_mm512_mask_cvtps_epi32 -
_mm512_mask_pow_pd -
_mm512_mask_cdfnorm_pd -
_mm512_mask3_fmsubadd_round_pd -
_mm_sqrt_round_sd -
_mm_mask_sqrt_sd -
_mm512_mask_movedup_pd -
_mm512_rcp14_ps -
_mm512_cmpge_epi64_mask -
_mm512_min_pd -
_mm512_undefined -
_mm512_mask_sin_pd -
_mm512_rcp14_pd -
_mm512_mask_cvtt_roundps_epi32 -
_mm512_maskz_compress_ps -
_mm_maskz_mul_sd -
_mm512_mask_log2_pd -
_mm512_maskz_srl_epi32 -
_mm512_maskz_or_epi64 -
_mm512_maskz_shuffle_f64x2 -
_mm512_erfc_pd -
_mm512_div_epi32 -
_mm_maskz_fmadd_sd -
_mm_maskz_scalef_ss -
_mm512_maskz_cvtepi64_epi8 -
_mm512_maskz_mul_round_ps -
_mm512_cvt_roundepi32_ps -
_mm512_rol_epi32 -
_mm512_mask_broadcast_i32x4 -
_mm512_zextpd256_pd512 -
_mm512_mask_sincos_ps -
_mm512_broadcast_i32x4 -
_mm512_cvt_roundpd_epi32 -
_mm512_maskz_getmant_ps -
_mm512_mask_cvtepi8_epi64 -
_mm512_permutevar_pd -
_mm512_mask_floor_pd -
_mm512_max_pd -
_mm_cvt_roundsd_u32 -
_mm512_cvtt_roundpd_epi32 -
_mm512_mask_exp_ps -
_mm512_broadcast_i64x4 -
_mm_cvt_roundss_sd -
_mm512_maskz_expand_pd -
_mm512_mask_cvtepi64_epi16 -
_mm512_mask_shuffle_pd -
_mm512_mask_storeu_ps -
_mm512_permute_pd -
_mm512_mask_cvt_roundps_pd -
_mm512_maskz_max_epu32 -
_mm512_mask_div_round_ps -
_mm512_srli_epi64 -
_mm_cmp_round_ss_mask -
_mm_cvti32_sd -
_mm512_mask_cmpge_epu64_mask -
_mm512_maskz_cvt_roundps_ph -
_mm512_asinh_ps -
_mm512_maskz_extractf32x4_ps -
_mm_maskz_fmadd_round_sd -
_mm_cvti64_sd -
_mm512_fmsubadd_round_ps -
_mm512_maskz_shuffle_epi32 -
_mm512_set_epi64 -
_mm512_mask_sinh_pd -
_mm_cvtt_roundss_i32 -
_mm512_mask_cvtepu32_epi64 -
_mm512_maskz_cvt_roundps_epu32 -
_mm512_log1p_ps -
_mm512_mask_unpackhi_pd -
_mm512_mask_cvtsepi32_storeu_epi16 -
_mm_mask3_fnmsub_sd -
_mm512_min_epu64 -
_mm512_hypot_ps -
_mm512_mask_rem_epi32 -
_mm_mask_fnmsub_sd -
_mm_maskz_max_round_ss -
_mm512_maskz_cvtpd_epu32 -
_mm512_mask_compressstoreu_epi64 -
_mm512_mask_tand_pd -
_mm512_pow_ps -
_mm512_permutex_pd -
_mm512_mask_cmplt_epi32_mask -
_mm512_mask_max_epu64 -
_mm512_mask_cosd_ps -
_mm512_maskz_cvt_roundepi32_ps -
_mm512_maskz_fmaddsub_round_ps -
_mm512_maskz_sll_epi32 -
_mm512_undefined_epi32 -
_mm512_maskz_permutex2var_epi64 -
_mm512_cmp_epu64_mask -
_mm512_mask_exp2_pd -
_mm_roundscale_round_ss -
_mm512_insertf32x4 -
_mm512_maskz_srav_epi64 -
_mm512_maskz_fmaddsub_round_pd -
_mm512_mask_min_round_ps -
_mm512_rem_epi16 -
_mm512_maskz_shuffle_ps -
_mm512_mask_trunc_ps -
_mm_max_round_ss -
_mm_maskz_fmsub_round_sd -
_mm512_mask_max_epi64 -
_mm512_maskz_cvtph_ps -
_mm_mask_max_round_ss -
_mm512_mask_sqrt_pd -
_mm512_mask_cvttps_epu32 -
_mm512_mask_atanh_pd -
_mm512_mask_cvtsepi64_storeu_epi8 -
_mm_cvtt_roundsd_u32 -
_mm512_erfinv_pd -
_mm_cvtsd_i32 -
_mm_mask_add_round_sd -
_mm512_set4_epi32 -
_mm512_mask_broadcastss_ps -
_mm512_stream_load_si512 -
_mm512_inserti32x4 -
_mm_cvt_roundsi64_sd -
_mm512_maskz_shuffle_pd -
_mm512_div_epi16 -
_mm512_maskz_sub_round_ps -
_mm512_mask_cvtepu16_epi64 -
_mm512_roundscale_ps -
_mm_maskz_fixupimm_round_ss -
_mm512_maskz_ternarylogic_epi64 -
_mm512_mask_permutex2var_ps -
_mm512_mask_expand_epi32 -
_mm512_maskz_max_pd -
_mm_mask_cvt_roundsd_ss -
_mm512_maskz_min_ps -
_mm_maskz_mul_round_ss -
_mm512_rolv_epi32 -
_mm512_storeu_si512 -
_mm512_maskz_ror_epi32 -
_mm512_cvtph_ps -
_mm512_maskz_max_round_pd -
_mm512_maskz_fixupimm_round_pd -
_mm_mask_store_ss -
_mm512_castsi256_si512 -
_mm_mask_fnmadd_sd -
_mm512_mask_sll_epi32 -
_mm512_rem_epu32 -
_mm512_maskz_unpackhi_ps -
_mm512_castpd512_pd128 -
_mm512_mask_loadu_pd -
_mm512_mask_compress_epi32 -
_mm512_mask2_permutex2var_pd -
_mm512_cvtsepi64_epi32 -
_mm512_maskz_permutex2var_epi32 -
_mm512_div_epu64 -
_mm512_mask_rint_ps -
_mm512_atan_pd -
_mm512_mask_cmpgt_epi64_mask -
_mm512_mask_permutevar_ps -
_mm_getmant_ss -
_mm_mask_move_sd -
_mm512_fixupimm_round_pd -
_mm512_mask_cvtpd_epi32 -
_mm512_maskz_unpacklo_epi64 -
_mm512_mask_max_pd -
_mm512_sincos_pd -
_mm512_mask_permutex_epi64 -
_mm_fixupimm_round_ss -
_mm_maskz_sqrt_sd -
_mm_mask3_fmsub_ss -
_mm512_set_epi32 -
_mm512_mask_cmpeq_epu64_mask -
_mm512_mask_cvtepu16_epi32 -
_mm512_mask_log_pd -
_mm512_mask_atan2_ps -
_mm_scalef_sd -
_mm512_setr_pd -
_mm512_mask_cvtepi32_epi16 -
_mm_mask_cmp_sd_mask -
_mm512_i64scatter_epi32 -
_mm512_maskz_fnmsub_ps -
_mm512_i64gather_ps -
_mm_mask_mul_round_sd -
_mm512_maskz_alignr_epi32 -
_mm512_cvtpd_ps -
_mm512_kmov -
_mm512_maskz_div_ps -
_mm512_mask_roundscale_round_ps -
_mm512_mul_epu32 -
_mm512_maskz_rolv_epi64 -
_mm512_erf_ps -
_mm512_set1_ps -
_mm512_mask_invsqrt_ps -
_mm512_maskz_or_epi32 -
_mm512_maskz_permutex2var_pd -
_mm_maskz_sqrt_round_sd -
_mm512_mask_cvtt_roundpd_epu32 -
_mm_maskz_add_sd -
_mm512_cvtepi32_pd -
_mm512_sqrt_ps -
_mm512_insertf64x4 -
_mm512_erfcinv_ps -
_mm512_cvtepi32_epi8 -
_mm512_maskz_compress_pd -
_mm512_maskz_cvtepu32_pd -
_mm512_setr4_epi32 -
_mm_maskz_fixupimm_ss -
_mm512_mask_rcp14_ps -
_mm512_sincos_ps -
_mm512_mask_sin_ps -
_mm512_cdfnorm_pd -
_mm512_max_round_pd -
_mm_max_round_sd -
_mm512_mask_expandloadu_pd -
_mm512_maskz_cvtps_epi32 -
_mm512_mask_min_ps -
_mm_sqrt_round_ss -
_mm_mask_sqrt_round_ss -
_mm512_mask_min_round_pd -
_mm512_atan2_ps -
_mm_mask_sub_round_ss -
_mm_mask_add_sd -
_mm512_maskz_broadcastsd_pd -
_mm512_mask_permutexvar_pd -
_mm_mask_getmant_round_ss -
_mm512_maskz_cvtsepi64_epi8 -
_mm512_sqrt_pd -
_mm512_unpacklo_ps -
_mm512_set1_epi64 -
_mm512_cvtepu8_epi32 -
_mm512_undefined_pd -
_mm512_mask_atan2_pd -
_mm512_maskz_set1_epi64 -
_mm512_mask_compressstoreu_epi32 -
_mm512_shuffle_ps -
_mm_min_round_ss -
_mm512_mask_tanh_pd -
_mm512_permutexvar_epi32 -
_mm512_castpd256_pd512 -
_mm_cvt_roundi64_ss -
_mm512_maskz_min_epi32 -
_mm_mask_getmant_round_sd -
_mm_mask_fixupimm_sd -
_mm512_zextsi128_si512 -
_mm_maskz_add_round_sd -
_mm512_mask_movehdup_ps -
_mm_mask3_fmsub_sd -
_mm_cvttsd_i32 -
_mm_mask_cvtss_sd -
_mm512_mask_permutex2var_epi32 -
_mm512_cvtusepi64_epi32 -
_mm512_expm1_ps -
_mm512_maskz_fmsub_round_ps -
_mm512_maskz_cvttpd_epu32 -
_mm512_mask_compressstoreu_ps -
_mm512_mask_sind_pd -
_mm512_mask_cvtepi32_epi64 -
_mm512_div_epu8 -
_mm512_fmsubadd_round_pd -
_mm512_movehdup_ps -
_mm512_mask_cmpneq_epu64_mask -
_mm_mask_getexp_sd -
_mm512_cosd_pd -
_mm512_maskz_rorv_epi64 -
_mm512_cmpeq_epi64_mask -
_mm512_maskz_fixupimm_pd -
_mm512_mask_compress_pd -
_mm_mask_min_ss -
_mm512_mask_cvtsepi64_storeu_epi16 -
_mm512_mask_hypot_ps -
_mm512_maskz_cvtepi64_epi16 -
_mm512_mask_fmsubadd_ps -
_mm512_maskz_sub_epi64 -
_mm_getexp_round_ss -
_mm512_mask2int -
_mm512_mask_erfinv_pd -
_mm_cvtss_i64 -
_mm512_i64gather_pd -
_mm_cvttsd_i64 -
_mm512_mask_exp10_ps -
_mm512_maskz_getexp_pd -
_mm_comi_round_sd -
_mm512_mask_broadcastd_epi32 -
_mm512_cvttpd_epi32 -
_mm512_set1_epi8 -
_mm_mask_getexp_ss -
_mm512_setr4_ps
Reacted by unageek-
Dissecting one of the interesting intrinsics here:
/// Compute the absolute value of packed 8-bit integers in a, /// and store the unsigned results in dst using writemask k (elements /// are copied from src when the corresponding mask bit is not set). __m512i _mm512_mask_abs_epi8 (__m512i src, __mmask64 k, __m512i a);
the
__mmask64type appears, which is a 64-bit mask where LLVM requires us to implement it as a<64 x i1>vector that must be allocated to ak64registers. In AVX-512kregisters are mask registers, and what seems to be more interesting is that AVX-512 does seem to supporti1as a type that is legal to lower to a clearedkregister with the first bit either set or unset...So it would be nice to know how does exactly all of this works in LLVM because
i1types are illegal in all other x86 "targets" (e.g. AVX2). Does anybody know?Another difference with AVX2 is that if we want to use a mask in AVX2 to select values from two
u8x32, the mask is ani8x32with each byte either set or unset but IIUC AVX-512__mmask32is also usable for this, but it requires 32bits instead. It would be nice to know if these two (i8x32as a mask and__mmask32) can interact, and if so, how. going from__mmask32toi8x32can probably be done in LLVM assext <32 x i1> to <32 x i8>and the opposite with atrunc <32 x i8> to <32 x i1>but maybe there is a different way in which these things must be done.This affects boolean vectors / masks, because
bool8x32would need to be casteable tobool1x32and vice-versa.So it would be nice to know how does exactly all of this works in LLVM because i1 types are illegal in all other x86 "targets" (e.g. AVX2). Does anybody know?
I don't understand how it works, but there are possibly relevant slides from the 2017 LLVM meeting, maybe they are useful: https://llvm.org/devmtg/2017-03//assets/slides/avx512_mask_registers_code_generation_challenges_in_llvm.pdf
Another possibly relevant point is that AVX512VL extends the mask registers (and the corresponding intrinsics) to 128- and 256-bit vectors. But, at least when using these from C, LLVM will currently just use blend instructions instead of masks: https://godbolt.org/g/FjU1Xn
But, at least when using these from C, LLVM will currently just use blend instructions instead of masks: https://godbolt.org/g/FjU1Xn
That's a really nice test. Do you know if there is an LLVM bug open for it? I haven't been able to find any.
Hi, has there been any new developments since this was last active? I would like to contribute AVX-512 intrinsics, but I'm not sure what (if anything) is blocking it, so if anyone has any pointers I'd be happy to help!
You can add any intrinsic that does not use
__mmask..types without issues.If you want to add an intrinsic that uses
__mmask..., you would need to add the mask types first. It is unclear what that would take. A#[repr(simd)] struct __mmask64(i64);might just work, or it might fail spectacularly. AFAIK nobody has tried yet.Clang defines masks like
__mmask16as just (https://github.com/llvm-mirror/clang/blob/master/lib/Headers/avx512fintrin.h#L48):typedef unsigned char __mmask8; typedef unsigned short __mmask16; typedef unsigned int __mmask32; typedef unsigned long long __mmask64;
So maybe just a wrapper struct without
#[repr(simd)]would be enough:pub struct __mmask8(u8); pub struct __mmask16(u16); pub struct __mmask32(u32); pub struct __mmask64(u64);
Cool! I'll give it a try some time this week, RustConf permitting.
Should AVX-512 intrinsics be split into modules corresponding to their feature flag?
This seems sensible except that I'm not sure how it should interact with the AVX512VL extension, since it seems weird to have the 512/256/128-bit versions of the same intrinsic in different places.
@hdevalence we currently split the functionality in modules corresponding to their target-feature flag and/or cpuid flag. I expect avx512f, avx512vl, etc. to be their own modules like they are in clang.
This stuff is decided on a 1:1 basis though, whoever sets the PR can get the conversation started. Are there any technical reasons to split it in any other way?
Hmm, but the VL flag is orthogonal to the other flags, so for instance the
_mm256_madd52hi_epu64intrinsic requires IFMA and VL. Where should it live?@hdevalence in clang they live in an
avx512ifmavlheader... avx-512 is complicated :/ many intrinsics require two features...EDIT: typically the ones that require
avx512f+avx512{something_else}live in the{something_else}module though.avx-512 is complicated :/
no kidding... looking at the AVX-512 Venn diagram:

it seems like the only CPUs that don't have VL extensions for all of their supported AVX-512 instructions are the Xeon Phi cores, which I think are all cancelled now, so it seems like the common case will be that if a CPU supports an instruction it will almost certainly support the VL extensions for it.
In that case, maybe it makes sense to split the intrinsics into modules
avx512f,avx512ifma, etc., and then within those modules separately gate the VL variants on theavx512vlflag. This is still correct in the edge case that VL is not present, but seems like a more logical grouping... I think clang maybe can't really do this because C doesn't have a module system.Does this seem like a sensible arrangement?
Reacted by n-soda, Kengo Sawatsu, zachmatson, Lesmiscore, Daira-Emma Hopwood and Jonathan BirkReacted by Aaron Hillit seems like the only CPUs that don't have VL extensions for all of their supported AVX-512 instructions are the Xeon Phi cores, which I think are all cancelled now,
I don't think we should worry about these. Some of these did not support SSE4.2 and IIRC AVX2 either (only AVX-512), and we can't target them with LLVM IIRC.
Does this seem like a sensible arrangement?
Sure. If once we start this way we discover that putting these into their own modules makes things clearer, we can always do that later.
73 remaining items
I had a look in the compiler and it seems that this is a bug in the implementation of
simd_select_bitmask: it should acceptu8inputs when the number of lanes is less than 8.simd_bitmaskalready supports this by returningu8when the number of lanes is less than 8.@minybot @bjorn3 Would one of you be willing to make a PR to fix this in rustc? The relevant code is here: https://github.com/rust-lang/rust/blob/f3c923a13a458c35ee26b3513533fce8a15c9c05/compiler/rustc_codegen_llvm/src/intrinsic.rs#L1272
There is another solution without touching simd_select_bitmask.
Use cast. Take _mm512_mask_extractf32x4_ps (__m128 src, __mmask8 k, __m512 a, int imm8) as an example.
a->(32x4); Cast to (32x16); Cast to (32x8); do bitmask; Cast to (32x4).
There is no cast128_to_256 directly. only 128_to_512, 512_to_256. 512_to_128.I just went ahead and fixed the issue in rust-lang/rust#77504.
I just went ahead and fixed the issue in rust-lang/rust#77504.
I test it, and it works when the mask size is 4.
For Mask operation in avx512 such as _kadd_mask32, it adds two masks.
According to https://travisdowns.github.io/blog/2019/12/05/kreg-facts.html, the Mask has its own hardware register.
Is there anyway to make sure _kadd_mask32 will generate "kaddd" instruction?No, but it's fine since we don't guarantee a particular instruction is used for an intrinsic: we leave it to LLVM to decide whether it is better to use a
kaddinstruction or a normaladdinstruction.While working on a private project, I needed masked loading, so I wanted to prepare a PR with implementations for
_mm512_mask_load_epi32and the like. Reading https://github.com/rust-lang/stdarch/blob/master/crates/core_arch/avx512f.md, I found the following:- _mm512_mask_load_epi32 //need i1
- _mm512_maskz_load_epi32 //need i1
What is the "need i1" part? I have not found any explanation there.
Currently, I am tempted to implement masked loading like in (as an example)
/// Load packed 32-bit integers from memory into dst using writemask k (elements are copied from src when the corresponding mask bit is not set). mem_addr must be aligned on a 64-byte boundary or a general-protection exception may be generated. /// /// [Intel's documentation](https://software.intel.com/sites/landingpage/IntrinsicsGuide/#text=_mm512_mask_load_epi32&expand=3305) #[inline] #[target_feature(enable = "avx512f")] #[cfg_attr(test, assert_instr(vmovdqa32))] pub unsafe fn _mm512_mask_load_epi32(src: __m512i, k: __mmask16, mem_addr: *const i32) -> __m512i { let loaded = ptr::read(mem_addr as *const __m512i).as_i32x16(); let src = src.as_i32x16(); transmute(simd_select_bitmask(k, loaded, src)) }
which follows how
_mm512_maskz_mov_epi32and_mm512_load_epi32are implemented. If this sounds correct, I might make a PR in the next days.This is incorrect since
_mm512_mask_load_epi32must not cause page faults on the parts of the vector that are masked off. Your version will still cause these page faults.To support this properly we need to call an LLVM intrinsic directly. However this intrinsic uses a vector of
i1as argument, which we cannot represent with Rust types. We need additional support in the compiler to call LLVM intrinsics that take a vector ofi1as a parameter.Makes sense. Many thanks for the explanation.
Another possible implementation for
_mm512_mask_load_epi32would using theasmfeature. I have successfully used the following implementation:#[inline] pub unsafe fn _mm512_mask_loadu_epi32(src: __m512i, mask: __mmask16, ptr: *const i32) -> __m512i { let mut result: __m512i = src; asm!( "vmovdqu32 {io}{{{k}}}, [{p}]", p = in(reg) ptr, k = in(kreg) mask, io = inout(zmm_reg) result, options(nostack), options(pure), options(readonly) ); result }
If such an implementation would be ok maintenance wise I could try preparing a PR that adds the missing avx512f this way.
If such an implementation would be ok maintenance wise I could try preparing a PR that adds the missing avx512f this way.
Sounds good!
Just coming from the discussion: rust-lang/portable-simd#28.
Regarding the separation of avx512f intrinsics and and target_feature=avx512f, now, I have enough interest and time to investigate it.
My particular case of interest is using zmm_reg for inline assembly (so need for avx512f intrinsics), but target_feature=avx512f is not stable yet. If it helps the stabilisation of target_feature, I am willing to work on it under some guidance.
@Amanieu What do you think?I expect that we will be stabilizing AVX-512 soon, thanks to the hard work of many people in implementing the full set of AVX-512 intrinsics in stdarch.
Reacted by Tony Arcieri, Robert Knight, mert-kurttutan, Dmitry Rusakov, Al Johri, Sayantan Chakraborty and Joshua FergusonReacted by Tony Arcieri, Dmitry Rusakov, Al Johri and Adnan ChaumetteReacted by Tony Arcieri, Al Johri, Radzivon Bartoshyk and Dmitry RusakovI believe this is a nice time to bring up the topic of
avx512vp2intersect😅 - it is stuck due to noi1support in rustc. Is there any other way we can implement them? That would truly complete the avx512 setWith #2081 the AVX512 set is finally complete ❤️
General instructions for this can be found at #40, but the list of AVX-512 intrinsics is quite large! This is intended to help track progress but you'll likely want to talk to us out of band to ensure that everything is coordinated.
Intrinsic lists: https://gist.github.com/alexcrichton/3281adb58af7f465cebee49759ae3164