Skip to content

CI: dist-s390x-linux build went from 40min. to 160min with new LLVM pass manager #89609

Description

@hkratz

CI time quadrupled when the new pass manager was enabled:

Right now that does not really matter much because Github's Apple CI runners use very old hardware and take about the same time, but once that is fixed this will block a faster CI. No idea if that platform has any significant userbase.

Maybe disable the new PM for that platform or report it upstream?

Activity

  1. workingjubilee commented on Oct 6, 2021

    @workingjubilee
    Member

    This regression absolutely needs to be reported upstream, as LLVM dev talk has mentioned completely removing the old pass manager within 1~2 versions.

  2. hkratz commented on Oct 6, 2021

    @hkratz
    ContributorAuthor

    Most time is apparently being spent compiling rustc_ast_lowering:

    Building stage1 compiler artifacts (x86_64-unknown-linux-gnu -> s390x-unknown-linux-gnu)
    ...
    [RUSTC-TIMING] rustc_ast_lowering test:false 6556.620
    

    https://github.com/rust-lang-ci/rust/runs/3708445190?check_suite_focus=true#step:26:12286

    This crate usually takes less than 30 seconds to build.

  3. added
    O-SystemZTarget: SystemZ processors (s390x)
    I-compiletimeIssue: Problems and improvements with respect to compile times.
    A-LLVMArea: Code generation parts specific to LLVM. Both correctness bugs and optimization-related issues.
    T-compilerRelevant to the compiler team, which will review and decide on the PR/issue.
    on Oct 6, 2021
  4. hkratz commented on Oct 7, 2021

    @hkratz
    ContributorAuthor

    The build timeout in #88379 (comment) might be caused by this as well.

  5. estebank commented on Oct 7, 2021

    @estebank
    Contributor

    Most time is apparently being spent compiling rustc_ast_lowering

    That is very surprising and am wondering what we could be doing for that not to be O(n).

  6. Mark-Simulacrum commented on Oct 7, 2021

    @Mark-Simulacrum
    Member

    cc @rust-lang/wg-llvm

    I am a little tempted to suggest a patch that passes -Zno-new-pass-manager or w/e on s390x, but that seems pretty unfortunate as well, given the relatively quick deprecation timeline.

  7. camelid commented on Oct 7, 2021

    @camelid
    Member

    Is there anything I can do to unblock #88379? Or will it require the patch that @Mark-Simulacrum is suggesting?

  8. camelid commented on Oct 7, 2021

    @camelid
    Member

    Just checking, has the regression been reported upstream yet? I didn't see a report when I searched the LLVM Bugzilla. (I don't have an LLVM Bugzilla account, so I can't report it.)

  9. cuviper commented on Oct 7, 2021

    @cuviper
    Member

    I'm going to try to find the culprit, but in the meantime to unblock folks that hack could be an extra target condition in should_use_new_llvm_pass_manager. Note that the host arch is apparently not the issue, since these are all running on x86_64, cross-compiling to various targets.

  10. cuviper commented on Oct 8, 2021

    @cuviper
    Member

    I have the dist-s390x-linux docker build running locally, with LLVM debuginfo, and perf top shows a couple hotspots.

    First it was sitting here at about 31% of the samples:

    llvm::Instruction* const* llvm::SmallVectorTemplateCommon<llvm::Instruction*, void>::reserveForParamAndGetAddressImpl<llvm::SmallVectorTemplateBase<llvm::Instruction*, true> >(llvm::SmallVectorTemplateBase<llvm::Instruction*, true>*, llvm::Instruction* const&, unsigned long) at /checkout/src/llvm-project/llvm/include/llvm/ADT/SmallVector.h:217
     (inlined by) llvm::SmallVectorTemplateBase<llvm::Instruction*, true>::reserveForParamAndGetAddress(llvm::Instruction*&, unsigned long) at /checkout/src/llvm-project/llvm/include/llvm/ADT/SmallVector.h:522
     (inlined by) llvm::SmallVectorTemplateBase<llvm::Instruction*, true>::push_back(llvm::Instruction*) at /checkout/src/llvm-project/llvm/include/llvm/ADT/SmallVector.h:547
     (inlined by) PushDefUseChildren at /checkout/src/llvm-project/llvm/lib/Analysis/ScalarEvolution.cpp:4382
     (inlined by) llvm::ScalarEvolution::forgetLoop(llvm::Loop const*) at /checkout/src/llvm-project/llvm/lib/Analysis/ScalarEvolution.cpp:7487
    

    Then it moved away from that to a new hotspot at 28%:

    llvm::ilist_detail::node_options<llvm::Instruction, false, false, void>::const_pointer llvm::ilist_detail::NodeAccess::getValuePtr<llvm::ilist_detail::node_options<llvm::Instruction, false, false, void> >(llvm::ilist_node_impl<llvm::ilist_detail::node_options<llvm::Instruction, false, false, void> > const*) at /checkout/src/llvm-project/llvm/include/llvm/ADT/ilist_node.h:184
     (inlined by) llvm::ilist_detail::SpecificNodeAccess<llvm::ilist_detail::node_options<llvm::Instruction, false, false, void> >::getValuePtr(llvm::ilist_node_impl<llvm::ilist_detail::node_options<llvm::Instruction, false, false, void> > const*) at /checkout/src/llvm-project/llvm/include/llvm/ADT/ilist_node.h:229
     (inlined by) llvm::ilist_iterator<llvm::ilist_detail::node_options<llvm::Instruction, false, false, void>, true, true>::operator*() const at /checkout/src/llvm-project/llvm/include/llvm/ADT/ilist_iterator.h:139
     (inlined by) llvm::simple_ilist<llvm::Instruction>::back() const at /checkout/src/llvm-project/llvm/include/llvm/ADT/simple_ilist.h:141
     (inlined by) llvm::BasicBlock::getTerminator() const at /checkout/src/llvm-project/llvm/lib/IR/BasicBlock.cpp:151
    

    Next it moved to this at 49%:

    llvm::MachineBasicBlock** std::__copy_move<false, false, std::random_access_iterator_tag>::__copy_m<std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**>(std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**) at /usr/include/c++/5/bits/stl_algobase.h:340
     (inlined by) llvm::MachineBasicBlock** std::__copy_move_a<false, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**>(std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**) at /usr/include/c++/5/bits/stl_algobase.h:402
     (inlined by) llvm::MachineBasicBlock** std::__copy_move_a2<false, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**>(std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**) at /usr/include/c++/5/bits/stl_algobase.h:440
     (inlined by) llvm::MachineBasicBlock** std::copy<std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**>(std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**) at /usr/include/c++/5/bits/stl_algobase.h:472
     (inlined by) llvm::MachineBasicBlock** std::__uninitialized_copy<true>::__uninit_copy<std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**>(std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**) at /usr/include/c++/5/bits/stl_uninitialized.h:93
     (inlined by) llvm::MachineBasicBlock** std::uninitialized_copy<std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**>(std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**) at /usr/include/c++/5/bits/stl_uninitialized.h:126
     (inlined by) void llvm::SmallVectorTemplateBase<llvm::MachineBasicBlock*, true>::uninitialized_copy<std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**>(std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, llvm::MachineBasicBlock**) at /checkout/src/llvm-project/llvm/include/llvm/ADT/SmallVector.h:490
     (inlined by) void llvm::SmallVectorImpl<llvm::MachineBasicBlock*>::append<std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, void>(std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >) at /checkout/src/llvm-project/llvm/include/llvm/ADT/SmallVector.h:652
     (inlined by) llvm::MachineBasicBlock** llvm::SmallVectorImpl<llvm::MachineBasicBlock*>::insert<std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, void>(llvm::MachineBasicBlock**, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >, std::reverse_iterator<__gnu_cxx::__normal_iterator<llvm::MachineBasicBlock**, std::vector<llvm::MachineBasicBlock*, std::allocator<llvm::MachineBasicBlock*> > > >) at /checkout/src/llvm-project/llvm/include/llvm/ADT/SmallVector.h:851
     (inlined by) llvm::LiveVariables::MarkVirtRegAliveInBlock(llvm::LiveVariables::VarInfo&, llvm::MachineBasicBlock*, llvm::MachineBasicBlock*, llvm::SmallVectorImpl<llvm::MachineBasicBlock*>&) at /checkout/src/llvm-project/llvm/lib/CodeGen/LiveVariables.cpp:112
    

    It's still going, but I'm not sure this is helpful enough to keep watching this way...

  11. cuviper commented on Oct 8, 2021

    @cuviper
    Member

    See also #89524, which is not s390x specific.

  12. camelid commented on Oct 8, 2021

    @camelid
    Member

    Based on those results, it looks to me (albeit as someone who's not knowledgeable about LLVM internals) like maybe it has some really large data structures that are slow to process. Could that be it?

  13. 7 remaining items

  14. added
    I-prioritizeIssue needs a team member to assess the impact. Will be replaced by P-{low,medium,high,critical}
    on Oct 10, 2021
  15. cuviper commented on Oct 12, 2021

    @cuviper
    Member
  16. apiraino commented on Oct 14, 2021

    @apiraino
    Contributor

    Assigning priority as discussed in the Zulip thread of the Prioritization Working Group.

    @rustbot label -I-prioritize +P-high

  17. added
    P-highHigh priority
    and removed
    I-prioritizeIssue needs a team member to assess the impact. Will be replaced by P-{low,medium,high,critical}
    on Oct 14, 2021
  18. self-assigned this
    on Oct 21, 2021
  19. cuviper commented on Oct 21, 2021

    @cuviper
    Member

    Status: #89666 disabled newPM when targeting s390x, but this is a temporary solution because oldPM will be removed eventually. The underlying issue does seem to be excessive inlining, similar to #89524 and others from newPM.

  20. nikic commented on Jan 7, 2022

    @nikic
    Contributor

    If anyone is wondering why this one only occurs on SystemZ: https://github.com/llvm/llvm-project/blob/main/llvm/lib/Target/SystemZ/SystemZTargetTransformInfo.h#L39 Apparently this target just increases all inlining threshold by a factor of three...

  21. cuviper commented on Jan 7, 2022

    @cuviper
    Member

    Aha, thanks! The only other targets that change that multiplier are NVPTX (5) and AMDGPU (11).

    AMD: See, NVIDIA only goes to 5 -- This one goes to 11! 🤘

  22. self-assigned this
    on Mar 9, 2022
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

A-LLVMArea: Code generation parts specific to LLVM. Both correctness bugs and optimization-related issues.I-compiletimeIssue: Problems and improvements with respect to compile times.O-SystemZTarget: SystemZ processors (s390x)P-highHigh priorityT-compilerRelevant to the compiler team, which will review and decide on the PR/issue.regression-from-stable-to-nightlyPerformance or correctness regression from stable to nightly.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions