Repository navigation
Poor performance returning enums larged than a word. Possibly poor code generation? #19864
Description
Activity
So I looked at this a little deeper. From the LLVM IR, it seems that when LLVM SROAs the
Resultstruct, it ends up spliting thememcopys andmemsets for the moves. Unfortunately, it includes the padding elements in the struct, resulting in bloated code and unaligned loads and stores.This would be fixed with some TBAA metadata emitted, so is essentially #6736.
This is something I've noticed as well.
Ok, so I did a basic implementation of
tbaa.structand it doesn't look like it fixes the issue.- addedI-slowIssue: Problems and improvements with respect to performance of generated code.Issue: Problems and improvements with respect to performance of generated code.
on Dec 17, 2014 Ok, so I managed to fix up the codegen, hopefully it's still valid. It improved speed a little, but part of the issue is that IoError is 8 words, or 64 bytes on x86-64.
pub struct IoError { pub kind: IoErrorKind // 2 words because a single variant has a uint associated with it pub desc: &'static str // 2 words - the size of a string slice pub detail: Option<String> // 4 words - One for the discriminant 3 for the String. }
It's no surprise that the code using
IoErroris so much slower when the data is so much bigger.I was talking with @mcpherrrin and @huonw, and they mentioned that we are shrinking the enum tag down to the smallest integer size, which would result in padding and the unsigned stores and loads you found. I wonder if disabling the shrinking would speed things up.
@erickt heh, that's exactly the change I made. Instead of just using the smallest integer type, I use the smallest to figure out what the appropriate alignment is, then use the alignment size to set the discriminant type. For some reason LLVM was struggling to figure out that the padding wasn't being written to or read from.
For the record, with @alexcrichton example the slower example gets twice as fast with the newer codegen. It's still much slower than the faster example, but that's hardly surprising.
Even with @Aatch's patches, we're just doubling the IoError case from 40MB/s to 80MB/s. The destructor we're getting from the
description: Option<Str>is really hurting us. Removing that gets us to 287MB/s. We can convertdesc: &'static strinto a function, which gets rid of another word and to 339MB/s. Finally, if we removeIoErrorKind::ShortWrite(uint)as suggested in the io reform rfc, we getIoErrordown into a word. This gets us to 731MB/s. Finally, if implement enum compression, then we'd be able to reduce a type likeResult<(), IoError>down into a single word, which would get us to the 1200MB/s number.#[deriving(Show, PartialEq, Eq)] enum MyError3Kind { EndOfFile, Error, _Error1, //(uint), } #[deriving(Show, PartialEq, Eq)] struct MyError3 { kind: MyError3Kind, //desc: &'static str, //description: Option<String>, } impl Error for MyError3 { fn is_eof(&self) -> bool { self.kind == MyError3Kind::EndOfFile } } #[bench] fn bench_foo11_enum_smaller_error(b: &mut test::Bencher) { let bytes = generate_bytes(); b.bytes = bytes.len() as u64; b.iter(|| { let mut rdr = bytes.as_slice(); let iter = Foo11::new(|buf| -> Result<(), MyError3> { match rdr.push(BUFFER_SIZE, buf) { Ok(_) => Ok(()), Err(io::IoError { kind: io::EndOfFile, .. }) => Ok(()), Err(_) => { Err(MyError3 { kind: MyError3Kind::Error, //desc: "", //description: Some("foo".to_string()), }) } } }); for (idx, item) in iter.enumerate() { assert_eq!(idx as u8, item.unwrap()); } }) }
CC @mitsuhiko
Rust implements large return types via an out pointer similar to the manual version in C.
@huonw: That's already how the platform ABIs handle large return values anyway.
The x86_64 ABI uses a higher threshold than Rust's choice of 1 word though.
- addedC-enhancementCategory: An issue proposing an enhancement or a PR with one.Category: An issue proposing an enhancement or a PR with one.
on Jul 22, 2017 Triage: updated code:
#![feature(test)] extern crate test; use std::iter::repeat; const N: usize = 100000; #[derive(Clone)] struct Foo; #[derive(Clone)] struct Bar { name: &'static str, desc: Option<String>, other: Option<String>, } #[bench] fn b1(b: &mut test::Bencher) { b.iter(|| { let r: Result<u8, Foo> = Ok(1u8); repeat(r).take(N).map(|x| test::black_box(x)).count() }); } #[bench] fn b2(b: &mut test::Bencher) { b.iter(|| { let r: Result<u8, Bar> = Ok(1u8); repeat(r).take(N).map(|x| test::black_box(x)).count() }); }
updated benchmarks:
running 2 tests test b1 ... bench: 25,617 ns/iter (+/- 2,836) test b2 ... bench: 1,093,834 ns/iter (+/- 50,578)- addedT-compilerRelevant to the compiler team, which will review and decide on the PR/issue.Relevant to the compiler team, which will review and decide on the PR/issue.
on Mar 30, 2020 That benchmark on
rustc 1.53.0-nightly (07e0e2ec2 2021-03-24)for me returns:running 2 tests test b1 ... bench: 25,082 ns/iter (+/- 26) test b2 ... bench: 501,483 ns/iter (+/- 1,637)If
b2bench will be changed to#[bench] fn b2(b: &mut test::Bencher) { b.iter(|| { let r: Result<u8, Bar> = Ok(1u8); iter::repeat_with(||r.as_ref()).take(N).map(|x| test::black_box(x)).count() }); }results will be
running 2 tests test b1 ... bench: 25,081 ns/iter (+/- 44) test b2 ... bench: 50,160 ns/iter (+/- 155)but this can be not related to actual problem.
- addedC-bugCategory: This is a bug.Category: This is a bug.and removedC-enhancementCategory: An issue proposing an enhancement or a PR with one.Category: An issue proposing an enhancement or a PR with one.
on May 24, 2023 With rustc 1.75.0-nightly (249624b 2023-10-20):
running 2 tests test b1 ... bench: 21,169 ns/iter (+/- 377) test b2 ... bench: 108,729 ns/iter (+/- 1,838)- addedC-optimizationCategory: An issue highlighting optimization opportunities or PRs implementing suchCategory: An issue highlighting optimization opportunities or PRs implementing suchA-codegenArea: Code generationArea: Code generationand removedC-bugCategory: This is a bug.Category: This is a bug.
on Mar 1, 2025
I've discovered an issue with
IoError, and really returning any enums that are larger than 1 word, are running an order of magnitude slower than returning an error enum that's 1 word size. Here's my test case:Here are the results:
On a related note, @alexcrichton just found a similar case with:
The assembly for
b1is about 3 instructions, but the assembly forb2has a ton ofmovinstructions.