Repository navigation
Tracking issue for WebAssembly SIMD support #74372
Description
Activity
- addedO-wasmTarget: WASM (WebAssembly), http://webassembly.org/Target: WASM (WebAssembly), http://webassembly.org/C-tracking-issueCategory: An issue tracking the progress of sth. like the implementation of an RFCCategory: An issue tracking the progress of sth. like the implementation of an RFC
on Jul 15, 2020 - addedA-SIMDArea: SIMD (Single Instruction Multiple Data)Area: SIMD (Single Instruction Multiple Data)
on Jul 15, 2020 - added a commit that references this issue
on Jul 28, 2020 Is there a way to enable wasm SIMD globally so that it can be adopted by the autovectorizer?
Emscripten seems to have such a feature: https://emscripten.org/docs/porting/simd.html
To enable SIMD, pass the -msimd128 flag at compile time. This will also turn on LLVM’s autovectorization passes, so no source modifications are necessary to benefit from SIMD.
Is there a way to enable wasm SIMD globally so that it can be adopted by the autovectorizer?
Enabling
simd128viatarget-featureseems to do the trick, this simple case gets vectorized: https://godbolt.org/z/hcaxxfReacted by est31Consider switching *_{any,all}_true to returning a bool
They really should, it doesn't seem like the compiler can reason about it... though maybe that is because the compiler doesn't even inline any of the intrinsics at all, which seems like a decently big problem:
A lot of the
load/storeintrinsics feel redundant: the memory access itself is better done in Rust code rather than by the intrinsic itself.Examples:
v128_load/v128_storeare literally just*mand*m = a. Also, they are inconsistent with the Clang version which allows unaligned access while the Rust version doesn't.v128_loadN_splatis justiNxM_splat(*x).v128_loadN_laneis justiNxM_replace_lane(a, *val).v128_storeN_laneis just*m = iNxM_extract_lane(v).v128_loadN_zerodoesn't have a direct intrinsic equivalent, but probably should have.
I would keep
v128_loadNxM_{s,u}though since these do not directly translate to Rust code.I agree
v128_{load,store}are redundant, especially because it's something Rust has to support anyway because you can take the address of any value at any time. I figured it was worth adding them for completeness, but you've got a good point that it may be best to use theptr::{read,write}_unalignedintrinsics instead of just raw pointer reads/writes.For
v128_{load,store}N_laneas well asv128_loadN_zeroI think unfortunately at this time need to expose the intrinsics. LLVM doesn't seem to optimize the equivalent paterns into the relevant instruction, although I'm sure that at some point in the future it will gain the ability to do so. I glanced at the discussions adding some of those intrinsics awhile back and I think the rationale was that compilers like V8 aren't great at pattern-matching multiple instructions to optimize the load/splat for example, and I think that ended up motivating the addition of the instructions. (I'm not 100% sure on this point though). In general though I think the idea is that compilers aren't always good enough to fuse the operations so there's a special instruction for it. (and compilers in this case I guess is both the one compiling wasm to native code but also LLVM compiling source code to wasm).I've been adding most of the intrinsics so far but unfortunately I haven't had the opportunity to write a compiled-to-wasm thing which actually uses the instructions. In that sense I'm mostly shooting in the dark as to what the APIs should be. So far I've tried to maximize availability (ensuring there's a function-per-instruction, regardless of how silly it is to have) and stick close to the most standardized piece, the spec names/type signatures. I did this to stay in the spirit of the x86 intrinsics which are copied verbatim from the spec and we provide virtually no abstractions to make them nice to use (even in some cases where it would be easy to do so). Overall, though, I'm not sure if this is the best tradeoff for wasm. WebAssembly isn't the same as x86 in this regard, so we may get more bang for our buck by putting more thought and effort into what a usable API would be (at the cost of time for stabilization, of course).
tl;dr; I'd like to ask if a @rust-lang/libs team member would be willing to
propose an FCP on this issue for stabilization.
Ok I think enough things have landed now that I'd like to propose that this
issue is considered for stabilization. I've read over the OP and it's still
quite accurate, the only difference is that the WebAssembly SIMD
proposal is now at phase 4 which is WebAssembly's "close to
stabilization". Additionally browsers will start shipping this feature ungated
in the future, with Chrome already scheduled to do so and also
implemented in
Firefox.
Consequently the instruction set that we're targeting is stable at this point
and I think we're in a position where we can stabilize the Rust intrinsics
themselves.To recap, this tracking proposal is for
v128-related functions in the
core::arch::wasm32
module (also available in
stdof course). The intrinsics andv128type in this module are intended to
expose the functionality of the WebAssembly simd
proposal which has both an English
description
as well as a proposed formal
specification.The
v128type and the new instructions added to WebAssembly can be used very
similarly to existing CPU intrinsics that Rust has forx86andx86_64.
Unlikex86_64, however, WebAssembly does not have a standard like the Intel
Intrinsics Guide or precedent in other languages of how to bind the SIMD support
in the language itself. Consequently a decision to stabilize these intrinsics in
Rust will be sort of blazing a new trail rather than following existing
precedent.The design princicples for the SIMD intrinsics in the
wasm32module I'd
propose stabilizing are:-
One type,
v128, is exposed. This mirrors the underlying type added to
WebAssembly itself. We could alternatively add types likei32x4instead of
just a singularv128type, but this would require extra conversions to also
be defined and given the precedent ofx86_64SIMD support in Rust thearch
module's primary goal is "easiest for compilers to implement and the library
team to stabilize" instead of "easiest for developers to use". -
Instructions are exposed with an
unsafefunction that has an appropriate
#[target_feature]annotation on it. WebAssembly is unlike other targets
where theunsafehere is not necessary. Unlikex86_64which isn't
guaranteed to raise a CPU fault on invalid intrinsics executed, WebAssembly
won't even begin to execute code unless the entire module validates.
These intrinsics could be safe if not for the current design of the
#[target_feature]attributes, and ideally these would be made safe in a
future Rust release as a non-breaking change. -
Instruction type signatures closely mirror the instruction's type signature
but use Rust native types where possible. For examplei16x8_all_truereturns
aboolinstead of ani32like the instruction does. -
The
v128.constinstruction is exposed as a variety of constructors (see
below), but none of which require constant arguments. The compiler (LLVM) is
left to select the most appropriate instruction (or sequence of instructions)
for creating a vector. This may bev128.const, but it may also not be. -
Instruction function names and type signatures are inspired by their
instruction name and instruction signature, but are not required to follow it.
WebAssembly instruction names follow WebAssembly conventions (e.g. only having
onei32type instead of twoi32andu32types) which aren't necessarily
conventions for Rust as well. Furthermore WebAssembly exposes types like a
Rustboolas ani32. This means that each instruction is considered in
isolation of how to best expose it to Rust.Some general principles for naming intrinsics are:
- Conveniences are exposed when the instruction is extremely low level. For
exampleu64x2is exposed which is typically lowered tov128.const. - Suffixes like
_sand_ufor instructions are baked into the funciton
names to to indicate the signededness of the vector argument. - Rust functions are provided for both signed and unsigned names of vectors
(with corresponding argument types).
Unlike x86_64 where there are thousands of intrinsics that can't be
hand-verified and scrutinized, WebAssembly is exposing ~200 intrinsics which
is expected to be reasonable enough to bikeshed intrinsics individually and
find the best names. A table of how instructions are exposed is: - Conveniences are exposed when the instruction is extremely low level. For
Wasm Instruction Rust Function(s) v128.loadv128_loadv128.load8x8_si16x8_load_extend_i8x8v128.load8x8_ui16x8_load_extend_u8x8v128.load16x4_si32x4_load_extend_i16x4v128.load16x4_ui32x4_load_extend_u16x4v128.load32x2_si64x2_load_extend_i32x2v128.load32x2_ui64x2_load_extend_u32x2v128.load8_splatv128_load8_splatv128.load16_splatv128_load16_splatv128.load32_splatv128_load32_splatv128.load64_splatv128_load64_splatv128.load32_zerov128_load32_zerov128.load64_zerov128_load64_zerov128.storev128_storev128.load8_lanev128_load8_lanev128.load16_lanev128_load16_lanev128.load32_lanev128_load32_lanev128.load64_lanev128_load64_lanev128.store8_lanev128_store8_lanev128.store16_lanev128_store16_lanev128.store32_lanev128_store32_lanev128.store64_lanev128_store64_lanev128.consti8x16,u8x16,i16x8,u16x8,i32x4,u32x4,i64x2,u64x2,f32x4,f64x2i8x16.shufflei8x16_shuffle,i16x8_shuffle,i32x4_shuffle,i64x2_shufflei8x16.extract_lane_si8x16_extract_lanei8x16.extract_lane_uu8x16_extract_lanei16x8.extract_lane_si16x8_extract_lanei16x8.extract_lane_uu16x8_extract_lanei32x4.extract_lanei32x4_extract_lane,u32x4_extract_lanei64x2.extract_lanei64x2_extract_lane,u64x2_extract_lanef32x4.extract_lanef32x4_extract_lanef64x2.extract_lanef64x2_extract_lanei8x16.replace_lanei8x16_replace_lane,u8x16_replace_lanei16x8.replace_lanei16x8_replace_lane,u16x8_replace_lanei32x4.replace_lanei32x4_replace_lane,u32x4_replace_lanei64x2.replace_lanei64x2_replace_lane,u64x2_replace_lanef32x4.replace_lanef32x4_replace_lanef64x2.replace_lanef64x2_replace_lanei8x16.swizzlei8x16_swizzlei8x16.splati8x16_splat,u8x16_splati16x8.splati16x8_splat,u16x8_splati32x4.splati32x4_splat,u32x4_splati64x2.splati64x2_splat,u64x2_splatf32x4.splatf32x4_splatf64x2.splatf64x2_splati8x16.eqi8x16_eqi8x16.nei8x16_nei8x16.lt_si8x16_lti8x16.lt_uu8x16_lti8x16.gt_si8x16_gti8x16.gt_uu8x16_gti8x16.le_si8x16_lei8x16.le_uu8x16_lei8x16.ge_si8x16_gei8x16.ge_uu8x16_gei16x8.eqi16x8_eqi16x8.nei16x8_nei16x8.lt_si16x8_lti16x8.lt_uu16x8_lti16x8.gt_si16x8_gti16x8.gt_uu16x8_gti16x8.le_si16x8_lei16x8.le_uu16x8_lei16x8.ge_si16x8_gei16x8.ge_uu16x8_gei32x4.eqi32x4_eqi32x4.nei32x4_nei32x4.lt_si32x4_lti32x4.lt_uu32x4_lti32x4.gt_si32x4_gti32x4.gt_uu32x4_gti32x4.le_si32x4_lei32x4.le_uu32x4_lei32x4.ge_si32x4_gei32x4.ge_uu32x4_gei64x2.eqi64x2_eqi64x2.nei64x2_nei64x2.lt_si64x2_lti64x2.gt_si64x2_gti64x2.le_si64x2_lei64x2.ge_si64x2_gef32x4.eqf32x4_eqf32x4.nef32x4_nef32x4.ltf32x4_ltf32x4.gtf32x4_gtf32x4.lef32x4_lef32x4.gef32x4_gef64x2.eqf64x2_eqf64x2.nef64x2_nef64x2.ltf64x2_ltf64x2.gtf64x2_gtf64x2.lef64x2_lef64x2.gef64x2_gev128.notv128_notv128.andv128_andv128.andnotv128_andnotv128.orv128_orv128.xorv128_xorv128.bitselectv128_bitselectv128.any_truev128_any_truei8x16.absi8x16_absi8x16.negi8x16_negi8x16.popcnti8x16_popcnti8x16.all_truei8x16_all_truei8x16.bitmaski8x16_bitmaski8x16.narrow_i16x8_si8x16_narrow_i16x8i8x16.narrow_i16x8_uu8x16_narrow_i16x8i8x16.shli8x16_shli8x16.shr_si8x16_shri8x16.shr_uu8x16_shri8x16.addi8x16_addi8x16.add_sat_si8x16_add_sati8x16.add_sat_uu8x16_add_sati8x16.subi8x16_subi8x16.sub_sat_si8x16_sub_sati8x16.sub_sat_uu8x16_sub_sati8x16.min_si8x16_mini8x16.min_uu8x16_mini8x16.max_si8x16_maxi8x16.max_uu8x16_maxi8x16.avgr_uu8x16_avgri16x8.extadd_pairwise_i8x16_si16x8_extadd_pairwise_i8x16i16x8.extadd_pairwise_i8x16_ui16x8_extadd_pairwise_u8x16i16x8.absi16x8_absi16x8.negi16x8_negi16x8.qmulr_sat_si16x8_q15mulr_sati16x8.all_truei16x8_all_truei16x8.bitmaski16x8_bitmaski16x8.narrow_i32x4_si16x8_narrow_i32x4i16x8.narrow_i32x4_uu16x8_narrow_i32x4i16x8.extend_low_i8x16_si16x8_extend_low_i8x16i16x8.extend_high_i8x16_si16x8_extend_high_i8x16i16x8.extend_low_i8x16_ui16x8_extend_low_u8x16i16x8.extend_high_i8x16_ui16x8_extend_high_u8x16i16x8.shli16x8_shli16x8.shr_si16x8_shri16x8.shr_uu16x8_shri16x8.addi16x8_addi16x8.add_sat_si16x8_add_sati16x8.add_sat_uu16x8_add_sati16x8.subi16x8_subi16x8.sub_sat_si16x8_sub_sati16x8.sub_sat_uu16x8_sub_sati16x8.muli16x8_muli16x8.min_si16x8_mini16x8.min_uu16x8_mini16x8.max_si16x8_maxi16x8.max_uu16x8_maxi16x8.avgr_uu16x8_avgri16x8.extmul_low_i8x16_si16x8_extmul_low_i8x16i16x8.extmul_high_i8x16_si16x8_extmul_high_i8x16i16x8.extmul_low_i8x16_ui16x8_extmul_low_u8x16i16x8.extmul_high_i8x16_ui16x8_extmul_high_u8x16i32x4.extadd_pairwise_i16x8_si32x4_extadd_pairwise_i16x8i32x4.extadd_pairwise_i16x8_ui32x4_extadd_pairwise_u16x8i32x4.absi32x4_absi32x4.negi32x4_negi32x4.all_truei32x4_all_truei32x4.bitmaski32x4_bitmaski32x4.extend_low_i16x8_si32x4_extend_low_i16x8i32x4.extend_high_i16x8_si32x4_extend_high_i16x8i32x4.extend_low_i16x8_ui32x4_extend_low_u16x8i32x4.extend_high_i16x8_ui32x4_extend_high_u16x8i32x4.shli32x4_shli32x4.shr_si32x4_shri32x4.shr_uu32x4_shri32x4.addi32x4_addi32x4.subi32x4_subi32x4.muli32x4_muli32x4.min_si32x4_mini32x4.min_uu32x4_mini32x4.max_si32x4_maxi32x4.max_uu32x4_maxi32x4.dot_i16x8_si32x4_dot_i16x8i32x4.extmul_low_i16x8_si32x4_extmul_low_i16x8i32x4.extmul_high_i16x8_si32x4_extmul_high_i16x8i32x4.extmul_low_i16x8_ui32x4_extmul_low_u16x8i32x4.extmul_high_i16x8_ui32x4_extmul_high_u16x8i64x2.absi64x2_absi64x2.negi64x2_negi64x2.all_truei64x2_all_truei64x2.bitmaski64x2_bitmaski64x2.extend_low_i32x4_si64x2_extend_low_i32x4i64x2.extend_high_i32x4_si64x2_extend_high_i32x4i64x2.extend_low_i32x4_ui64x2_extend_low_u32x4i64x2.extend_high_i32x4_ui64x2_extend_high_u32x4i64x2.shli64x2_shli64x2.shr_si64x2_shri64x2.shr_uu64x2_shri64x2.addi64x2_addi64x2.subi64x2_subi64x2.muli64x2_muli64x2.extmul_low_i32x4_si64x2_extmul_low_i32x4i64x2.extmul_high_i32x4_si64x2_extmul_high_i32x4i64x2.extmul_low_i32x4_ui64x2_extmul_low_u32x4i64x2.extmul_high_i32x4_ui64x2_extmul_high_u32x4f32x4.ceilf32x4_ceilf32x4.floorf32x4_floorf32x4.truncf32x4_truncf32x4.nearestf32x4_nearestf32x4.absf32x4_absf32x4.negf32x4_negf32x4.sqrtf32x4_sqrtf32x4.addf32x4_addf32x4.subf32x4_subf32x4.mulf32x4_mulf32x4.divf32x4_divf32x4.minf32x4_minf32x4.maxf32x4_maxf32x4.pminf32x4_pminf32x4.pmaxf32x4_pmaxf64x2.ceilf64x2_ceilf64x2.floorf64x2_floorf64x2.truncf64x2_truncf64x2.nearestf64x2_nearestf64x2.absf64x2_absf64x2.negf64x2_negf64x2.sqrtf64x2_sqrtf64x2.addf64x2_addf64x2.subf64x2_subf64x2.mulf64x2_mulf64x2.divf64x2_divf64x2.minf64x2_minf64x2.maxf64x2_maxf64x2.pminf64x2_pminf64x2.pmaxf64x2_pmaxi32x4.trunc_sat_f32x4_si32x4_trunc_sat_f32x4i32x4.trunc_sat_f32x4_uu32x4_trunc_sat_f32x4f32x4.convert_i32x4_sf32x4_convert_i32x4f32x4.convert_i32x4_uf32x4_convert_u32x4i32x4.trunc_sat_f64x2_s_zeroi32x4_trunc_sat_f64x2_zeroi32x4.trunc_sat_f64x2_u_zerou32x4_trunc_sat_f64x2_zerof64x2.convert_low_i32x4_sf64x2_convert_low_i32x4f64x2.convert_low_i32x4_uf64x2_convert_low_u32x4f32x4.demote_f64x2_zerof32x4_demote_f64x2_zerof64x2.promote_low_f32x4f64x2_promote_low_f32x4It's worth nothing that Clang has a header file for WebAssembly
intrinsics in the same way it has one for x86_64. The Rust names do not
precisely match the names exposed in C. The first shift is that Rust functions
are not prefixed withwasm_(since Rust has awasm32module unlike C does).
Even accounting for this difference Rust has other naming differences:wasm_i16x8_load8x8vsi16x8_load_extend_i8x8(and related)wasm_u16x8_load8x8vsi16x8_load_extend_u8x8(and related)wasm_i8x16_makevsi8x16(and related)wasm_i8x16_constvsi8x16(and related)- C lacks a
wasm_u8x16_splatfunction (only hasi8x16) - C lacks a
wasm_u8x16_replace_lanefunction (only hasi8x16) - C lacks a
wasm_u32x4_extract_lanefunction (only hasi32x4) wasm_u32x4_extend_low_u16x8vsi32x4_extend_low_u16x8(and related)wasm_u16x8_extmul_high_u8x16vsi16x8_extmul_high_u8x16(and related)
The intention behind these design decisions is that all the functionality of the
proposal should be exposed in Rust, but at the same time we give ourselves
wiggle room where possible to add conveniences. The conveniences aren't super
principled other than evaluating intrinsics one-by-one. There are other possible
points naturally on the design spectrum for this module, for example providing
many types likeu8x16with first-class methods and such. Functions taking
memory arguments could arguably take safe references instead of raw
pointers as well. As @Amanieu pointed
out
other intrinsics are unlikely to ever get used or can otherwise be trivially
composed from other intrinsics (relying on LLVM to optimize, which it doesn't
always do just yet). My hope is that by getting a few more eyeballs on this we
can figure out a good balance for Rust and what we'd like the intrinsics to be.Implementation Status
It's worth touching on the implementation status of this proposal currently as
well. At this time thestdarch
repository
has an implementation for all of the above intrinsics. After
#84654 lands LLVM will support all
instructions excepti64x2.abs(just a minor LLVM issue) and the final update
tostdarchto match LLVM will be in
rust-lang/stdarch#1131. Note that actually updating the
stdarchsubmodule in rust-lang/rust happening in
#83278 is blocked. This means that the
current documentation for
wasm does not match this
proposal, despite this proposal all being implemented.It's also worth mentioning that this is all a relatively new feature in LLVM.
LLVM 12 an prior do not support the WebAssembly simd proposal as-is. Just before
reaching stage 4 a number of opcodes were renumbered, so LLVM versions 12 and
prior will produce invalid WebAssembly binaries using the simd feature. A number
of backports have been performed onto Rust's LLVM 12 fork so the Rust fork of
LLVM fully supports WebAssembly SIMD, but this means that until LLVM 13 is
released it's unlikely that cross-language-LTO or similar interop will be
supported between C and Rust.Reacted by Christopher Serr, est31, Sven, Daniel Liu, Josh Triplett and Andrew Gallant-
i8x16, u8x16, i16x8, u16x8, i32x4, u32x4, i64x2, u64x2, f32x4, f64x2
I'm still somewhat skeptical of those being the names for the
v128.constconstructor functions as that means no such types could ever live in this module, and idk if we truly want to block off that possibility forever. I agree that for now just a v128 type and these free standing functions is fine, but maybe those constructor functions could be named with a_constsuffix like the actual instruction (or e.g._makelike C seems to use).There was a small bikeshed here about those names where
_constwas a bit of a misnomer since the arguments aren't required to be constant. Other prefixes like_makeor_newseem fine to me!I don't think, though, that this prevents us from having dedicated types for each lane width because functions go into the value namespace and structs go into the type namespace.
Seems like it prevents the type from being a tuple struct though: https://play.rust-lang.org/?version=stable&mode=debug&edition=2018&gist=40b1e2abd2ade54dd598db826ed594bf
But as long as using a non-tuple struct for those types doesn't cause any problems such as repr(simd) or so not working anymore, then I guess it's fine.
46 remaining items
I've posted the final PR for stabilization to #86204
- added a commit that references this issue
on Jun 11, 2021 What will happen if code dependent on SIMD intrinsics will be compiled without the necessary target feature being enabled? Will compiler generate respective SIMD instructions no questions asked (well, it will not inline them, but it's not important for the question)? Shouldn't it generate a compilation error? Otherwise generated WASM may silently become unusable in runtimes without SIMD support (imagine an incorrect change somewhere deep in a project's dependency tree).
As the documentation indicates:
This means to generate a binary without SIMD you’ll need to avoid both options above plus calling into any intrinsics in this module.
where the "both options" are using
-Ctarget-feature=+simd128or#[target_feature(enable = "simd128")].The compiler will always generate simd instructions if you call intrinsics in
std::arch::wasm32, regardless of your compilation settings for the rest of the codegen unit. This is the same as for all other platform intrinsics, for example, in thex86andx86_64modules.This is the same as for all other platform intrinsics, for example, in the x86 and x86_64 modules.
Yes, but on other platforms those intrinsics are
unsafe, so there is a higher barrier for making mistake. Also having a respective instruction in a generated binary does not cause CPU to reject that binary, i.e. on x86 it's fine to have unchecked SIMD instructions in dead branches. In a certain sense using WASM SIMD intrinsic effectively changes target. So i wonder if the current handling of them is too lax.Also
#[target_feature(enable = "simd128")]is useless in my understanding. Since there is no runtime detection (and by your words there probably will not be one), WASM target features currently should be enabled for a whole project. I don't see a use case for enabling it only for selected functions.Yes, but on other platforms those intrinsics are unsafe, so there is a higher barrier for making mistake.
Isn't this a violation of RFC 2045? It seems to me that RFC 2396's requirement to make calling functions with target features on functions that don't have target features unsafe has not been implemented?
Since there is no runtime detection (and by your words there probably will not be one)
@newpavlov the docs you linked mention runtime detection proposals.
the docs you linked mention runtime detection proposals.
Yes, I know and this is why I asked this question on reddit. But it looks like that @alexcrichton has acted under assumption that dynamic detection will not be added to WASM:
Currently WebAssembly doesn't have any form of dynamic detectino for supported features, and even on the horizon I'm not sure there's really anything viable for adding this in a form that looks like x86. The closest equivalent for WebAssembly is conditional sections but afaik that proposal is sort of dead in the water right now and doesn't have a future. There are possible alternatives of a roughly similar shape, but it's all about selection at compile time instead of runtime. Basically I think it's tough to answer precisely what would happen if wasm gets dynamic detection because it's unclear to me what it means to get dynamic detection.
BTW I want to discuss this point from the same comment a bit:
The whole point of #[target_feature] and unsafe is that we don't know what happens if a CPU executes code it doesn't understand, but the whole point of WebAssembly is this never happens.
In my understanding, unsafety of
target_featurenot caused only by concerns of CPU behavior (after all, outside of some very edge scenarios you usually will get a simple invalid opcode exception). It's probably not even the top reason, see this @gnzlbg's comment referenced in thetarget_featureRFC. Are we sure that those codegen concerns are not applicable to WASM?There's some more discussion on #84988 for reference, but my thoughts on this topic are:
- Wasm engines, forever and all of time, will reject modules they do not understand. This means if you have a simd instruction in their an the wasm engine doesn't understand it, then a correct engine will never run any code and simply reject the module.
- If dynamic feature detection is added then it will be added in such a way that the same modules can be executed differently on two engines (one for example with simd and without). If you have correctly configured your module and dynamic detection then the module that doesn't support simd will never see the simd code. This means it's never executed. If you incorrectly configured the dynamic feature detection then the engine that doesn't support simd will see simd instructions it does not understand. From the first point this means that the runtime will reject the module.
- Some concerns on the original
target_featureRFC were indeed connected to "assemblers and things can assume weird things". AFAIK those are generally theoretical concerns but still ones we take seriously. Again WebAssembly is different because the concern worried about here is that instructions for the wrong architecture leak into other parts of the program (e.g. an assembler realizes that you're using avx2 things so it assumes it can just keep using those, even though it's on a different path that's not supposed to execute when avx2 is not detected or something like that). As always wasm engines will either run the module or reject it entirely.
Basically there's gotchas and subtelties about generating modules in WebAssembly that do use simd, don't use simd, or try to dynamically detect simd (assuming some sort of future proposal), but these are all build time concerns. Whatever happens the final module will have some defined semantics based on the instructions its using (which are presumably all stable in the upstream wasm spec with clearly-defined semantics). This means that you'll either run the module on an engine understanding all instructions, in which case everything will behave exactly as expected, or you won't, in which case nothing will execute at all.
Overall there is no situation where a unknown wasm instruction is executed. There's situations your build isn't what you expect, but that's not UB that's a configuration issue.
Wasm engines, forever and all of time, will reject modules they do not understand.
IIUC with the
features.supportedproposal engines will accept feature blocks which they do not understand, but instead of parsing it, they will replace it as a whole with theunreachableinstruction. It's effectively analogue of theud2instruction from x86. The question is now: is it allowed for safe Rust code to trigger architecture's "unreachable" instructions? AFAIK in the case of x86 the answer is no. Do we allow it for WASM?Also I wonder how in the presence of runtime feature detection compiler will handle functions which use SIMD intrinsics, but do not provide a non-SIMD fallback. Will it simply wrap every intrinsic with a
features.supportedcheck and useunreachablein fallback branches?There's situations your build isn't what you expect, but that's not UB that's a configuration issue.
Yes, it's not UB which would cause memory corruption, but it's still a real pitfall. Is it possible to make use of
simd128SIMD intrinsics without globally enabled target feature a compilation error? It would completely solve my concerns (at the very least until some kind dynamic detection will be added, but I really hope that we wil lget something similar to target restriction contexts before that). Also I think it would significantly reduce probability of unpleasant surprises caused by such misconfiguration, especially considering that its source may be deep in a dependency tree.- On Thu, Jul 29, 2021 at 03:50:32PM -0700, Artyom Pavlov wrote: >Wasm engines, forever and all of time, will reject modules they do not understand. With the `features.supported` proposal engines will accept feature blocks which they do not understand, but instead of parsing it, they will replace the whole with the `unreachable` instruction. It's effectively analogue of the `ud2` instruction from x86. The question is now: is it allowed for safe Rust code to trigger "unreachable" instructions? AFAIK in the case of x86 the answer is no. Do we allow it for WASM?As far as I know, WebAssembly's "unreachable" is more like [`unreachable!`](https://doc.rust-lang.org/std/macro.unreachable.html) (which is entirely safe), not like [`unreachable_unchecked`](https://doc.rust-lang.org/std/hint/fn.unreachable_unchecked.html) (which is unsafe and can cause UB).
Yeah one of the reasons to make target_feature unsafe was because some instruction sets might override the meaning of instructions or do other UB when they encounter instructions they don't support. On wasm, there is a safe guarantee that it'll trap so it's well defined.
- added a commit that references this issue
on Aug 5, 2023
I'm opening this as a tracking issue for the SIMD intrinsics in the
{std,core}::arch::wasm32module. Eventually we're going to want to stabilize these intrinsics for the WebAssembly target, so I think it's good to have a canonical place to talk about them! I'm also going to update the#![unstable]annotations to point to this issue to direct users here if they want to use these intrinsics.The WebAssembly simd proposal is currently in "phase 3". I would say that we probably don't want to consider stabilizing these intrinsics until the proposal has at least reached "phase 4" where it's being standardized, because there are still changes to the proposal happening over time (small ones at this point, though). As a brief overview, the WebAssembly simd proposal adds a new type,
v128, and a suite of instructions to perform data processing with this type. The intention is that this is readily portable to a lot of architectures so usage of SIMD can be fast in lots of places.For rust stabilization purposes the code for all these intrinsics lives in the rust-lang/stdarch git repository. All code lives in
crates/core_arch/src/wasm32/simd128.rs. I've got a large refactoring and sync queued up for that module, so I'm going to be writing this issue with the assumption that it will land mostly as designed there.Currently the design principles for the SIMD intrinsics are:
memory_size,memory_growandunreachableintrinsics, most intrinsics are named after the instruction that it represents. There is generally a 1:1 mapping with new instructions added to WebAssembly and intrinsics in the module.#[target_feature(enable = "simd128")]which forces them all to beunsafev128.constis exposed through a suite ofconstfunctions, one for each vector type (but not unsigned, just signed integers). Additionally the arguments are not actually required to be constant, so it's expected that the compiler will make the best choice about how to generate a runtime vector.v8x16_shuffleand*_{extract,replace}_laneuse const generics to represent constant arguments. This is different from x86_64 which uses the older#[rustc_args_required_const]attribute.v16x8,v32x4, andv64x2as conveniences instead of only providingv8x16_shuffle. All of them are implemented in terms of thev8x16.shuffleinstruction, however.v128type, not a type for each size of vector that intrinsics operate withextract_laneintrinsics return the value type associated with the intrinsic name, they do not all returni32unlike the actual WebAssembly instruction. This means that we do not haveextract_lane_sandextract_lane_uintrinsics because the compiler will select the appropriate one depending on the context.It's important to note that clang has an implementation of these intrinsics in the
wasm_simd128.hheader. The current design of the Rustwasm32module is different in that:wasm_*isn't used.v128, is exposed instead of types for each size/kind of vectorwasm_i16x8_load_8x8andwasm_u16x8_load_8x8while Rust hasi16x8_load8x8_sandi16x8_load8x8_u.Most of these differences are largely stylistic, but there are some that are conveniences (like other forms of shuffles) which might be nice to expose in Rust as well. All the conveniences still compile down to one instruction, it's just different how users specify in code how the instruction is generated. I believe it should be possible for conveniences to live outside the standard library as well, however.
How SIMD will be used
If the SIMD proposal were to move to stage 4 today I think we're in a really good spot for stabilization. #74320 is a pretty serious bug we will want to fix before full stabilization but I don't believe the fix will be hard to land in LLVM (I've already talked with some folks on that side).
Other than that SIMD-in-wasm is different from other platforms where a binary with SIMD will refuse to run on engines that do not have SIMD support. In that sense there is no runtime feature detection available to SIMD consumers. (at least not natively)
After rust-lang/stdarch#874 lands programs will simply use
#[target_feature(enable = "...")]orRUSTFLAGSand everything should work. The SIMD intrinsics will always be exposed from the standard library (but the standard library itself will not use them) and available to users. If programs don't use the intrinsics then SIMD won't get emitted, otherwise when used the binary will usev128.Open Questions
A set of things we'll need to settle on before stabilizing (and this will likely expand over time) is:
*_load_*and*_store_*instructions. Primarily the instructions that load 64 bits (8x8, 16x4, ...) I'm unsure of on the types of their pointer arguments.v8x16_shuffleand lane managment instructions.i8x16_extract_lane_sis ok (e.g. havingi8x16_extract_lanereturningi8is all we need), same fori16x8.#[target_feature]"requires unsafe" rules for these WebAssembly intrinsics. Intrinsic likef32x4_splathave no fundamental reason they need to beunsafe. The only reason they're unsafe is because#[target_feature]is used on them to ensure that SIMD instructions are generated in LLVM.*_{any,all}_trueto returning abool