Skip to content

Write a parallel deque for work stealing #4877

Description

@brson

The work stealing algorithm uses a deque. Their algorithm has some properties that might make further optimizations possible later (one end is used only by a single thread). We only need something simple to start with, a locked vector that just pushes and pops and shifts and unshifts. Using a circular buffer would be better, lock-free better still (maybe).

Also probably relevant is the paper on data locality in work stealing though I haven't read it yet.

There are some useful data structures for atomically reference counted types and mutexes in core::private.

Activity

  1. brson commented on Feb 10, 2013

    @brson
    ContributorAuthor

    Related to #3095

  2. brson commented on Feb 15, 2013

    @brson
    ContributorAuthor

    Since I'm not sure that lock-free deques even exist I've scaled back the scope of this slightly.

  3. nikomatsakis commented on Feb 17, 2013

    @nikomatsakis
    Contributor

    The deques used in work-stealing typically have a special property that one thread (the owner) only PUSHES and POP and other threads (the thieves) only ever DEQUEUE. Nobody ever queues. This lets you get better efficiency for multi-threading, but means that they are not multi-purpose.

  4. brson commented on Feb 20, 2013

    @brson
    ContributorAuthor

    Somebody pointed out that the 'Chase / Lev' deque is a lock-free parallel deque for work stealing. Sounds good to me.

  5. ILyoan commented on Apr 30, 2013

    @ILyoan
    Contributor

    Is there any progress on this?

  6. brson commented on May 2, 2013

    @brson
    ContributorAuthor

    @ILyoan No. There is a type in place at rt::work_queue but it is not implemented.

  7. brson commented on May 16, 2013

    @brson
    ContributorAuthor

    A related paper http://www.cs.bgu.ac.il/~hendlerd/papers/dynamic-size-deque.pdf, "A Dynamic-Sized Nonblocking Work Stealing Deque", by Hendler, Lev, Moir, Shavit

  8. emberian commented on Jul 12, 2013

    @emberian
    Contributor

    @Aatch was working on this at some point but stopped.

  9. toddaaro commented on Jul 25, 2013

    @toddaaro
    Contributor

    http://www.di.ens.fr/~zappa/readings/ppopp13.pdf

    This recent paper is a wonderfully detailed description of the atomic memory issues involved with the data structure. Includes pseudocode for a C11 memory model version with every atomic operation specified.

  10. cartazio commented on Sep 4, 2013

    @cartazio
    Contributor

    a good example implementation of such a scheduler can be found in this haskell lib: https://github.com/ekmett/structures/blob/master/src/Control/Concurrent/Deque.hs

  11. brson commented on Oct 17, 2013

    @brson
    ContributorAuthor

    I am wondering whether we really want to use a work-stealing deque as or per-thread work queue. Something that allows for random access would prevent the starvation problems we have to hack around and would be more fair. Is there such a data structure that is lock free?

  12. toffaletti commented on Nov 6, 2013

    @toffaletti
    Contributor

    I've got a work-in-progress implementation of chase-lev deque. I started working from the C11 paper, but I found some bugs and omissions in their implementation so I'm having to go back to the original paper. Any help would be appreciated. I'm currently trying to wrap my head around Section 4 of the paper, which discusses how to adapt the algorithm for growing/shrinking the array without a garbage collector.

    https://github.com/toffaletti/rust-code/blob/master/chase_lev_deque.rs

    Even if this isn't used for the scheduler, I've been told it might be useful for Servo rendering work.

  13. cartazio commented on Nov 6, 2013

    @cartazio
    Contributor

    if you want to see a worked out chase lev implementation thats pretty readable, look at the deque branch of edward kmett's structures lib https://github.com/ekmett/structures/blob/deque/src/Control/Concurrent/Deque.hs

    it uses some other primops you can see defined here http://hackage.haskell.org/package/atomic-primops-0.4/docs/Data-Atomics-Counter-Reference.html

    some of that may not be relevant for the no gc context, but at least gives a pretty readable working implementation to look at explicitly

  14. toffaletti commented on Nov 7, 2013

    @toffaletti
    Contributor

    Thanks, @cartazio I will take a look. I haven't seen a fully working implementation that does array resizing and reclaiming without a GC. The original paper is light on specifics, just outlining a solution and then in a footnote saying "It is straightforward, however, to use the same solution for reclaiming buffers also when growing."

    I decided to use Relacy Race Detector to help get a correct implementation because helgrind and drd report too many false positives. That attempt is here: https://github.com/toffaletti/chase-lev

    It currently passes the tests, but I've gone the extremely heavy handed route of making all atomic access use memory_order_seq_cst. I'll work on relaxing that and fixing any other bugs.

  15. 7 remaining items

  16. added a commit that references this issue on Nov 29, 2013
    57f4a05
  17. added a commit that references this issue on Nov 29, 2013
    a70f9d7
  18. added a commit that references this issue on Nov 29, 2013
    dd1184e
  19. added a commit that references this issue on Aug 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    A-runtimeArea: std's runtime and "pre-main" init for handling backtraces, unwinds, stack overflowsC-enhancementCategory: An issue proposing an enhancement or a PR with one.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions