Skip to content

[Relay] Dead code elimination pass blows up call stack #4534

Description

@halluci293

This was not obvious to me because code that I was previously working with successfully suddenly started segfaulting when I upgraded my TVM build to the 0.6.0 release, but:

  • for even relatively simple feed forward graphs (a few convolutional layers + fully connected + output), the gradient function from relay becomes fairly complex especially if using higher order mode
  • for higher order mode, when running a PartialEvaluate() + DeadCodeElimination() pass, the dead code elimination segfaults under Ubuntu default ulimit, but passes when increasing ulimit to something very large
  • looking at the coredump in gdb, the stack when the segfault happens is several thousand frames deep inside the recursive node traversal here: https://github.com/apache/incubator-tvm/blob/master/src/relay/pass/dead_code.cc#L131
  • for first order mode gradients, the VM compilation step also runs several DeadCodeElimination passes and for gradients of larger models (especially tensorflow models which insert a lot of transpose operations to make conv2d layers NCHW), the same stack overflow happens

I think this is a regression from earlier versions, but:

  • it would be preferable to remove the recursion inside DCE, or fix it so it's tail recursive and the compiler can optimize away previous stack frames, or transform to an explicit stack, so users don't get this cryptic segfault
  • if that's not possible, the documentation should include a note that users should increase their ulimit

Activity

  1. MarisaKirisame commented on Dec 25, 2019

    @MarisaKirisame
    Contributor

    @jroesch care to take a look?

  2. trevor-m commented on Jan 7, 2020

    @trevor-m
    Contributor

    I was getting a segfault during relay.build() while trying to run resnet152_v1 with the script below. Smaller models worked fine. Once I increased my machine's stack limit using ulimit -s unlimited, the segfaults stopped. The stack limit was 8192 kilobytes originally. Might be related?

    import numpy as np
    import tvm
    from tvm import relay
    from tvm.contrib import graph_runtime
    import mxnet
    from mxnet.gluon.model_zoo.vision import get_model
    input_shape = (1, 3, 224, 224)
    block = get_model('resnet152_v1', pretrained=True)
    mod, params = relay.frontend.from_mxnet(block, shape={'data': input_shape}, dtype='float32')
    with relay.build_config(opt_level=3):
        graph, lib, params = relay.build(mod, "cuda", params=params)
    mod = graph_runtime.create(graph, lib, ctx=tvm.gpu(0))
    mod.set_input(**params)
    i_data = np.random.uniform(0, 1, input_shape).astype('float32')
    for i in range(10):
        mod.run(data=i_data)
    
  3. YunLexi commented on Mar 3, 2020

    @YunLexi

    @trevor-m
    Hi, I get the same problem when I try to use relay.build() to build a resnet101, with target as Cuda, but it works fine if I change the model to resnet18. Have you solved this problem?

  4. ANSHUMAN87 commented on Mar 4, 2020

    @ANSHUMAN87
    Contributor

    @YunLexi : The actual issue in dead code elimination pass is fixed withhttps://github.com//pull/4053.
    I think this might be some other issue.
    Can you share the piece of code to reproduce the issue?
    NOTE: I tried the code shared by @trevor-m , but the issue did not occur in my workspace, even with stack limit as 8192.

  5. YunLexi commented on Mar 8, 2020

    @YunLexi

    @ANSHUMAN87
    I just try to run the following tutorial with resnet101_v1, https://github.com/apache/incubator-tvm/blob/master/tutorials/frontend/from_mxnet.py, the program hangs at https://github.com/apache/incubator-tvm/blob/87faaf12f3d2b792bacccadeb369236ab5c5b45b/python/tvm/contrib/nvcc.py#L95 not doing anything, but this tutorial works fine if I change the model to resnet18_v1 or resnet50_v1 with cuda as target.

  6. ANSHUMAN87 commented on Mar 9, 2020

    @ANSHUMAN87
    Contributor

    @YunLexi : I executed from_mxnet.py with cuda as target and with model as resnet101_v1. It runs fine. I think issue is not in TVM. Issue is in your CUDA setup. Can you try uninstall and install freshly again. I think it will solve.

  7. halluci293 commented on Mar 9, 2020

    @halluci293
    ContributorAuthor

    I think the conversation here has diverged from the original problem.

    @ANSHUMAN87 I don't think the original problem I was running into has been solved. To be clear, this isn't a bug in the implementation (it's not infinite recursion, just very deep recursion, because if I remove the stack limit on my system the code works), but it is a problematic implementation. The nature of the dead code elimination implementation requires extensive recursive calls to visit every node of the graph. Since this is implemented as naive recursion, for a large enough graph (like the kind you get when using auto-diff to generate gradient functions Relay) it is easy to exhaust the default stack limit. It is not indicated anywhere obvious in the documentation that users should increase their stack limit, and the resulting segfault when this happens can be really confusing to understand if you don't consider stack overflow.

    As I mentioned in the original issue, there are two things that should be done:

    1. make it obvious in the documentation that users should increase their stack size if they encounter segfaults of this nature, or better yet, implement something like a recursion counter inside that graph traversal that warns users when they're reaching high levels of recursion in their graphs that could trigger stackoverflow
    2. fix the stack overflow, by either figuring out how to optimize the recursion implementation so that the compiler can perform tail call optimization, or convert the recursive code to iteration/reimplement with an explicit stack
  8. ANSHUMAN87 commented on Mar 9, 2020

    @ANSHUMAN87
    Contributor

    @swu : Thank you! I have clearly understood your issue in your original report. Have you tried in the latest code in Tvm, where the PR I mentioned is merged? Please crosscheck. I believe you should not encounter the issue again. I am eagerly waiting for your response. Thanks!

  9. ANSHUMAN87 commented on Mar 9, 2020

    @ANSHUMAN87
    Contributor

    @swu : if still you face the issue. We can find proper solution for it. Thanks!

  10. halluci293 commented on Mar 9, 2020

    @halluci293
    ContributorAuthor

    @ANSHUMAN87 yes, I can verify that the version of TVM I am using does have #4053 (I am using 0.6.0 release which was cut 2 months after that patch was merged, but I also just double checked src/relay/pass/dead_code.cc to verify that the changes from that patch are there).

    This version still gives me stack overflow when trying to build gradient functions unless I ulimit -s unlimited. I don't have a readily available model to share, but I'm trying to build a gradient function for a model that I converted from tensorflow with 5 convolution layers + 3 dense layers, so it's not gigantic.

  11. MarisaKirisame commented on Mar 9, 2020

    @MarisaKirisame
    Contributor

    I can manage an explicit stack or write it in CPS+Trampoline style to remove the blowup.
    The DCE pass need rework. For now I suggest not touching it.
    I am on other project but I will get back to training soon.

  12. tqchen commented on Jul 21, 2020

    @tqchen
    Member

    Close for now as it is potentially fixed by #4886, please feel free to open another thread

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions