Repository navigation
Insert yield checks at appropriate places #524
Description
Activity
There's the issue of never yielding, and the issue of fairness. For 0.1 let's just try to make sure we don't tie up a scheduler forever in an iloop.
I added a yield after send, which I thinks makes it less likely that one might create an iloop that never yields. Do we want to try to make back edges yield for 0.1? I say no.
I will note that when this happens, we will probably need to change
rust_task_yield_fail()in rust_task.cpp to not always fail if a task yields in an atomic section. instead, it should check if the yield was explicit or compiler-inserted, and fail if the first or silently ignore if the second.discussed at workweek, this is going to be accomplished "natively" by work stealing + keeping at least one spare thread whenever there are more tasks than threads.
Maybe I don't understand what that means, but won't that just result in "extra threads" being allocated until there are as many threads as tasks?
I'm pretty sure this is still an issue, so reopening. With the new runtime written in rust, we could have a
#[no_yield_checks]attribute for crate-level or file-level that we'd put in the scheduler.The only overhead of a thread compared to a Rust task is the context switching at arbitrary points. Inserting yield checks would be much slower than just using 1:1 threading, at least on Linux, so I don't think it makes sense to do this. You're better off with context switches than yield checks in critical loops.
If I remember right, the plan was to insert checks on back-edges in the CFG (presumably including tailcalls), and to have them only actually yield 1 in BIGNUM times. It would be worth profiling but I think that would still save significantly over kernel-mode context switches..
@bblum So long as the stealing and thread spawning behavior is rate limited (i.e. after the spare thread sees one of the schedulers hasn't seen a yield in K ms), the scenario you're describing would occur only when a user is exclusively making non-yielding (i.e., no i/o, no sleeping, fully cpu bound) tasks. And that's the case where multiplying threads to equal tasks is probably appropriate behavior: to saturate all the available cores with computation, as best the os kernel can.
It is possible someone will not want this behavior in some case. If they make so many cpu-spinning tasks the os literally can't handle the overhead, or perhaps they want a fixed number of rust threads even though they want to overload them. There are other mechanisms users can employ to achieve these ends in these cases. We decided on the strategy we did because it seemed like the more appropriate default, and avoids the worst problems (systemic taxes, artificial blocking or starvation). Most of the time, tasks do i/o or enter a potential yield point (say, malloc) regularly.
Somewhat off-topic: I don't think
mallocis really a potential yield point, because with jemalloc it's lock-free for allocations under 4K and only hits kernel synchronization when it actually has to make a system call (allocations over 4K, and occasionally to increase the pool size for small ones). It's never really blocking.12 remaining items
- added a commit that references this issue
on Mar 7, 2023 - added a commit that references this issue
on Jul 10, 2024 - added a commit that references this issue
on Jun 3, 2025 - added a commit that references this issue
on Sep 8, 2025 - added 8 commits that reference this issue
on Sep 26, 2026