Repository navigation
docs installation slow on NFS home #1867
Description
Activity
What's fascinating here is that while the install took minutes, there's only 6 seconds of syscalls there. I can only assume we're only getting CPU cost, not wait latency logged by strace.
However, profiles support should resolve this issue for you (PR #1673) so let's revisit this after the release of a
rustupcontaining that.I don't think I'd assume that; here's the
timeoutput for a re-installation of nightlyreal 2m38.806s user 0m16.254s sys 0m6.602sI'll be watching that PR with great interest!
Yep, so 2m38s of wallclock time, of which 16s user time (likely decompression, hashing, etc) and 6.6s of syscall CPU time. The remaining 2m14s or so will be waiting for the storage/network. :D
Reacted by Ben Kimock and TrevorYep, you are entirely blocked on high latency IO. Can I get another strace please, with -T this time? Attach it somewhere I can download it, it will be long.
Pending verification from that trace bug...
There are several things we can do, some of which are low hanging.We can introduce threading into your environment; if the open+write calls are low latency and close is slow, which client-checked but server enforced filesystems often have, then the current threading will be effective, we just need to enable it for you.
We can stop doing a second walk over the unpacked files - tar-rs already chmods (fchmod now) the files as it unpacks; we actually completely reset the permissions afterwards. A little bit of smarts (e.g. doing a fchmod on the handle from tar-rs for the common case and then only doing a fchmod for the uncommon case from the manifest) would let us drastically reduce syscalls); better still might be to tell tar-rs to not set chmod's at all as another feature - we'll want to be guided by the profile data.
- added a commit that references this issue
on May 26, 2019 @saethlin I haven't reviewed the strace yet, and its late, but that patch of mine is the lowest hanging fruit that may help, I'd be interested to know whether it has any impact for you at all.
Ok, so interesting that we get open() rather than openat, but -
open("/home/kimockb/.rustup/tmp/4duqw9jkaf9kmt3r_dir/rust-docs/share/doc/rust/html/core/core_arch/x86/avx/_mm256_broadcast_sd.v.html", O_WRONLY|O_CREAT|O_EXCL|O_CLOEXEC, 0666) = 9 <0.002960> write(9, "<!DOCTYPE html>\n<html lang=\"en\">"..., 349) = 349 <0.000012> fchmod(9, 0644) = 0 <0.001859> close(9) = 0 <0.000009> ... chmod("/home/kimockb/.rustup/toolchains/nightly-x86_64-unknown-linux-gnu/share/doc/rust/html/tutorial.html", 0644) = 0 <0.001082>about 3ms to create the inode, about 2ms to fchmod it, and writes and closes are nearly free, and chmod is also ~1ms.
20k items, so 1ms == 20 seconds. so 60 seconds in open(), 40 seconds in fchmod(), and 20 seconds() in the final chmods.
So, the things we need to do to make it fast for HPC filesystems are:
- don't block on inode creation
- don't apply chmod's twice
- and do the final chmod threaded or in the initial IO
@saethlin how many cores do you have available? Clearly we can oversubscribe as your actual CPU time is near zero, but I'm wondering if a heuristic of machine-cpu's which we use for dealing with defender will be reasonable here too (and save on code path proliferation)
Usually I have 32 cores available.
Just to be clear, any changes you make in the name of installing docs faster specifically on HPC systems (as opposed to everywhere) will be moot as soon as we can disable docs. I'm also not really keen on having an installation process that uses threads to drive enormous load onto the shared filesystem. I've witnessed write latencies of a few seconds because of misuse and I don't want to see rustup usage be responsible for that.
composefs/tar-rs#207 will permit tackling the chmod duplication. Its worth noting (for @alexcrichton if no-one else that there are several reasons we don't trust the distribution archives: they have been wrong in the past; we are still used to install those old archives, and when we start streaming-installs of archives we'll only verify signatures at the end so we want to trust as little as possible.
@saethlin so there are two parts here, one is actually doing the least work possible. Like - we used to do a full additional copy of the files which would have made older releases even slower; fixed already. Eliminating a full second pass of chmod metadata updates will reduce load on the metadata server (which HPC system are you using BTW? gluster? cephfs?).
The second part, once we have reduce the work done to the bare minimum with no waste is about scheduling: do we schedule the work to complete as rapidly as possible, or do we throttle it?
I understand your concern and share it: I certainly don't want to create headaches for folk by driving up a backlog somewhere. But lets put this in context - we have under a second of IO to a single platter of spinning rust disks (consumer disks have exceeded 200MBps for a decade or more now). Lets not assume that we are going to cause a problem until we have data showing it. I know for instance that with ext4 you want to create the files relatively closely together on a busy system, because ext4 is going to try to keep them together, and if the fs is too full (and high activity is moving what is used around a lot) it won't be able to. https://www.kernel.org/doc/html/latest/filesystems/ext4/overview.html#block-and-inode-allocation-policy
So my proposal is this: I write various things, and you try them out - I don't have your HPC cluster with its particular tradeoffs. If they cause havoc in test, we don't push them into master. If they don't cause havoc, we can consider moving forward.
We're on Lustre, not entirely sure what version though.
Your proposal sounds good. I'm learning about
iostat, hopefully I can watch rustup run and learn something or just keep an eye on the system.As a for instance rust-docs/share/doc/rust/html/rust-by-example/fn/closures/closure_examples:
strace for closure_examples
mkdir("/home/kimockb/.rustup/tmp/4duqw9jkaf9kmt3r_dir/rust-docs/share/doc/rust/html/rust-by-example/fn/closures/closure_examples", 0777) = 0 <0.001906> chmod("/home/kimockb/.rustup/tmp/4duqw9jkaf9kmt3r_dir/rust-docs/share/doc/rust/html/rust-by-example/fn/closures/closure_examples", 0755) = 0 <0.001278> stat("/home/kimockb/.rustup/tmp/4duqw9jkaf9kmt3r_dir/rust-docs/share/doc/rust/html/rust-by-example/fn/closures/closure_examples", {st_mode=S_IFDIR|0755, st_size=0, ...}) = 0 <0.000020> openat(AT_FDCWD, "/home/kimockb/.rustup/toolchains/nightly-x86_64-unknown-linux-gnu/share/doc/rust/html/rust-by-example/fn/closures/closure_examples", O_RDONLY|O_NONBLOCK|O_DIRECTORY|O_CLOEXEC) = 12 <0.000223> lstat("/home/kimockb/.rustup/toolchains/nightly-x86_64-unknown-linux-gnu/share/doc/rust/html/rust-by-example/fn/closures/closure_examples", {st_mode=S_IFDIR|0755, st_size=1024, ...}) = 0 <0.000013> chmod("/home/kimockb/.rustup/toolchains/nightly-x86_64-unknown-linux-gnu/share/doc/rust/html/rust-by-example/fn/closures/closure_examples", 0755) = 0 <0.001041>This has a total syscall time of
mkdir 0.001906 chmod 0.001278 stat 0.000020 openat 0.000223 lstat 0.000013 chmod 0.001041 -------------- 0.004481Which is representative of the extra work here - we really only need to set umask appropriately, do the mkdir once, then never touch the dir again. That would halve the amount of work that the HPC file system is being asked to do.
Reacted by Ben Kimockhttp://www.nics.tennessee.edu/computing-resources/file-systems/io-lustre-tips seems like it has good recommendations here.
In particular, @saethlin can you
lfs setstripe ~/.rustup -s 1m -i -1 -c 1 rm -rf ~/.rustup/tmpAnd test current released rustup ? I'm not expecting any radical change, just curious about any impact this might have. The stdlib and rustc etc are obviously stripable, and we could look at special casing docs etc to make just them single-stripe, but this is more along the lines of information gathering.
(Also TBH I think just some sensible docs in README.md might be useful here).
Oh dangit, I was mistaken. Our
/homepartitions aren't Lustre, only the big data partitions are.╰ ➤ mount | grep $HOME 10.13.152.221:/home/kimockb on /home/kimockb type nfs (rw,relatime,vers=3,rsize=1048576,wsize=1048576,namlen=255,hard,noacl,proto=tcp,timeo=600,retrans=2,sec=sys,mountaddr=10.13.152.221,mountvers=3,mountport=2049,mountproto=tcp,local_lock=none,addr=10.13.152.221)I'll sure find a use for those lustre tips anyway.
Ahha! ok, so if ~ is on plain old NFS, we're probably facing a classic single-machine multi user situation - but its also likely that that machine actually has decent IO backing it, if we're not nuts - we're just 1ms away on the network, so anything that has to cache-bust or be synchronous blocks. I'll have some non-aggressive test branch for you soon.
- changed the title
[-]docs installation slow on HPC system(s)[/-][+]docs installation slow on NFS home[/+]on Jun 6, 2019 Rustup 1.20 has been released which supports a
minimalprofile to not installrust-docsby default. Is this acceptable enough to close this issue?Absolutely. Thanks so much for all your hard work on this.
Problem
rustup docs installation is remarkably slow on an HPC system (probably most HPC systems). In case it's useful to anyone, I'm observing this on HiPerGator, which uses the Lustre filesystem.
I've attached some
strace -coutput below along with what rustup printed during the update. The 433.6 KiB/s is probably mis-stating the situation; for a few seconds at a time the install speed drops to between 50 and 8.0 KiB/s as if it's being throttled by the filesystem.Possible Solution(s)
It would be great to opt out of docs altogether; I'm told this is a feature under consideration which would be great. There's no way for me to use the HTML docs in an HPC environment; I use the local ones on the machine I'm ssh'd from.
Notes
Output of
rustup --version: rustup 1.18.3 (435397f 2019-05-22)Output of
rustup show:strace -coutput while running arustup update nightly