Repository navigation
[Relay] Dead code elimination pass blows up call stack #4534
Description
Activity
@jroesch care to take a look?
I was getting a segfault during
relay.build()while trying to run resnet152_v1 with the script below. Smaller models worked fine. Once I increased my machine's stack limit usingulimit -s unlimited, the segfaults stopped. The stack limit was 8192 kilobytes originally. Might be related?import numpy as np import tvm from tvm import relay from tvm.contrib import graph_runtime import mxnet from mxnet.gluon.model_zoo.vision import get_model input_shape = (1, 3, 224, 224) block = get_model('resnet152_v1', pretrained=True) mod, params = relay.frontend.from_mxnet(block, shape={'data': input_shape}, dtype='float32') with relay.build_config(opt_level=3): graph, lib, params = relay.build(mod, "cuda", params=params) mod = graph_runtime.create(graph, lib, ctx=tvm.gpu(0)) mod.set_input(**params) i_data = np.random.uniform(0, 1, input_shape).astype('float32') for i in range(10): mod.run(data=i_data)@trevor-m
Hi, I get the same problem when I try to use relay.build() to build a resnet101, with target as Cuda, but it works fine if I change the model to resnet18. Have you solved this problem?@YunLexi : The actual issue in dead code elimination pass is fixed withhttps://github.com//pull/4053.
I think this might be some other issue.
Can you share the piece of code to reproduce the issue?
NOTE: I tried the code shared by @trevor-m , but the issue did not occur in my workspace, even with stack limit as 8192.@ANSHUMAN87
I just try to run the following tutorial with resnet101_v1, https://github.com/apache/incubator-tvm/blob/master/tutorials/frontend/from_mxnet.py, the program hangs at https://github.com/apache/incubator-tvm/blob/87faaf12f3d2b792bacccadeb369236ab5c5b45b/python/tvm/contrib/nvcc.py#L95 not doing anything, but this tutorial works fine if I change the model to resnet18_v1 or resnet50_v1 with cuda as target.@YunLexi : I executed from_mxnet.py with cuda as target and with model as resnet101_v1. It runs fine. I think issue is not in TVM. Issue is in your CUDA setup. Can you try uninstall and install freshly again. I think it will solve.
I think the conversation here has diverged from the original problem.
@ANSHUMAN87 I don't think the original problem I was running into has been solved. To be clear, this isn't a bug in the implementation (it's not infinite recursion, just very deep recursion, because if I remove the stack limit on my system the code works), but it is a problematic implementation. The nature of the dead code elimination implementation requires extensive recursive calls to visit every node of the graph. Since this is implemented as naive recursion, for a large enough graph (like the kind you get when using auto-diff to generate gradient functions Relay) it is easy to exhaust the default stack limit. It is not indicated anywhere obvious in the documentation that users should increase their stack limit, and the resulting segfault when this happens can be really confusing to understand if you don't consider stack overflow.
As I mentioned in the original issue, there are two things that should be done:
- make it obvious in the documentation that users should increase their stack size if they encounter segfaults of this nature, or better yet, implement something like a recursion counter inside that graph traversal that warns users when they're reaching high levels of recursion in their graphs that could trigger stackoverflow
- fix the stack overflow, by either figuring out how to optimize the recursion implementation so that the compiler can perform tail call optimization, or convert the recursive code to iteration/reimplement with an explicit stack
@swu : Thank you! I have clearly understood your issue in your original report. Have you tried in the latest code in Tvm, where the PR I mentioned is merged? Please crosscheck. I believe you should not encounter the issue again. I am eagerly waiting for your response. Thanks!
@swu : if still you face the issue. We can find proper solution for it. Thanks!
@ANSHUMAN87 yes, I can verify that the version of TVM I am using does have #4053 (I am using 0.6.0 release which was cut 2 months after that patch was merged, but I also just double checked src/relay/pass/dead_code.cc to verify that the changes from that patch are there).
This version still gives me stack overflow when trying to build gradient functions unless I
ulimit -s unlimited. I don't have a readily available model to share, but I'm trying to build a gradient function for a model that I converted from tensorflow with 5 convolution layers + 3 dense layers, so it's not gigantic.I can manage an explicit stack or write it in CPS+Trampoline style to remove the blowup.
The DCE pass need rework. For now I suggest not touching it.
I am on other project but I will get back to training soon.Close for now as it is potentially fixed by #4886, please feel free to open another thread
This was not obvious to me because code that I was previously working with successfully suddenly started segfaulting when I upgraded my TVM build to the 0.6.0 release, but:
I think this is a regression from earlier versions, but: