Skip to content

Tracking issue for collections reform part 2 (RFC 509) #19986

Description

@aturon

Activity

  1. added
    B-RFC-approvedBlocker: Approved by a merged RFC but not yet implemented.
    on Dec 18, 2014
  2. aturon commented on Dec 18, 2014

    @aturon
    ContributorAuthor
  3. Gankra commented on Dec 19, 2014

    @Gankra
    Contributor

    TODO:

    Easy

    Moderate

    • add append to:
      • DList
      • Vec
      • RingBuf
      • BTreeMap
      • BTreeSet
      • Bitv
      • BitvSet
      • VecMap
    • add split_off:
      • DList
      • Vec
      • RingBuf
      • BTreeMap
      • BTreeSet
      • Bitv
      • BitvSet
      • VecMap
    • Upgrade the Entry API - has PR
    • Add truncate and resize to RingBuf (not specified, but makes sense to do) - has PR
    • Add swap_pop to RingBuf (not specified, but makes sense to do) - has PR
  4. Gankra commented on Dec 19, 2014

    @Gankra
    Contributor
  5. aturon commented on Dec 19, 2014

    @aturon
    ContributorAuthor

    Note: I'm about to post a PR that does the deprecation for Vec methods.

  6. cgaebel commented on Dec 19, 2014

    @cgaebel

    cc @bfops I think you were working on upgrading the entry api?

  7. csouth3 commented on Dec 19, 2014

    @csouth3
    Contributor

    I can take care of resize for Vec.

  8. bfops commented on Dec 20, 2014

    @bfops
    Contributor

    Yeah I'm working on the entry API. Seems to be coming along!

  9. pczarn commented on Dec 21, 2014

    @pczarn
    Contributor

    I'll start work on these methods for RingBuf:

    Add truncate and resize to RingBuf (not specified, but makes sense to do)
    Add swap_pop to RingBuf (not specified, but makes sense to do)

    Did you mean swap_remove?

  10. Gankra commented on Dec 21, 2014

    @Gankra
    Contributor

    Yes, sorry.

  11. gamazeps commented on Dec 21, 2014

    @gamazeps
    Contributor

    I'm starting the misc stabilization, should be over in a 2-3 days :)

  12. csouth3 commented on Dec 21, 2014

    @csouth3
    Contributor

    So, some of the misc stuff is done in #20053, but I didn't touch Bitv::{get,set}, didn't touch HashMap's and HashSet's iterators (either the documentation, or actually implementing DoubleEndedIterator and ExactSizeIterator for all of the appropriate iterators using the next_back implemented as next strategy noted in the RFC (if this is even what the RFC intends)), and I didn't attempt the tedium of moving std::vec to std::collections::vec, sooo....there's certainly plenty of misc left to do but wanted to give you heads up so your time doesn't end up getting wasted duplicating work :)

  13. gamazeps commented on Dec 21, 2014

    @gamazeps
    Contributor

    thx :)

  14. pczarn commented on Dec 22, 2014

    @pczarn
    Contributor

    Oh, did you mean swap_back_remove and swap_front_remove? What about grow for ringbuf?

  15. 51 remaining items

  16. xosmig commented on Apr 22, 2016

    @xosmig
    Contributor

    I've started working on append and split_off methods for BTreeMap and BTreeSet.

  17. apasel422 commented on Apr 22, 2016

    @apasel422
    Contributor

    @xosmig #32466 already covers BTreeMap::append.

  18. xosmig commented on Apr 22, 2016

    @xosmig
    Contributor

    @apasel422 OK. Then only split_off.

  19. xosmig commented on Apr 28, 2016

    @xosmig
    Contributor

    Hello!
    I've faced with problems.
    To recalc sizes of trees after split_off using O(log n) operations, it's necessary to know exact size of each subtree.
    So, there are some different solutions:

    1. add size field to the LeafNode structure.

    I think, benefits are great.

    First, it lets us to do very flexible splits and merges with only O(log n) operations.
    Now it's possible to merge trees using O(log n) operations, but it it not so interesting without fast split.

    Second, some other interesting and usefull functions need it.
    For example: kth_element or "get the number of elements between two values". Both with O(log n) operations.

    Third, as I understand (but i am not totally sure), it lets us to implement the rope ( https://en.wikipedia.org/wiki/Rope_(data_structure) ) in easy way. Just as BTreeMap<(), V>.

    But overhead is notable (4 or 8 bytes for each node, depending on machine word size). I think, it is possible to embed this sizes without memory overhead for 64-bit machines (due to aligning).

    1. Implement split_off with O(log n + "size of the one of the results")

    2.1) Split with O(log n) operations and then calculate size of the one of the results.

    2.2) There is another algorithm which doesn't use the fast algorithm with the same complexity, but (as I understand) with greater constant.

    1. Implement split_off with O(log n + "size of the smallest one") but with greater constant

    Split using O(log n) operations and then calculate sizes of the trees via 2 alternating iterators, but stop when one of the sizes is counted.

    Which implementation should I use?

  20. xosmig commented on Apr 29, 2016

    @xosmig
    Contributor

    @apasel422, what do you think?

  21. apasel422 commented on Apr 29, 2016

    @apasel422
    Contributor

    @xosmig This isn't my area of expertise, but I'd recommend opening a discussion thread on https://internals.rust-lang.org/ or submitting a pull request for whichever option is easiest and CCing @jooert and @gereeter. You might also want to take a look at #26227.

  22. xosmig commented on Apr 29, 2016

    @xosmig
    Contributor

    @apasel422 The main question is "Is this overhead okay or not?". So, who can decide it?

  23. gereeter commented on Apr 29, 2016

    @gereeter
    Contributor

    First, note that option one can be slightly optimized by only storing composite sized in internal nodes.

    I'm not a big fan of option one - I expect split_off to be a fairly rare operation, and slowing down everything else and increasing the memory footprint solely for its benefit does not seem like a good idea. Since we explicitly support large collections (as large as can fit in memory, basically), the size field would have to be a usize, which cannot be combined with other fields. The application to ropes is interesting, but really wants a separate implementation - the BTreeMap interface is totally wrong for the job (though the node module might be useful). kth_element and measuring the length of a range seems useful, but again, I see them as niche use cases that really want a specialized sequence data structure.

    I don't have strong feelings between options two and three. I think that a variation on option two, where the choice of which side to count is made based on which edge was descended at the root node, would probably be the best. It doesn't have the overhead of option three, but still wont take long if you split near the edges.

    There is also a fourth option, namely dropping support for constant time calls to len(). This is a somewhat more radical change, but it would easily allow for split_off in O(log n) and avoid the question of how to calculate size completely. Moreover, while it might inconvenience users, it is easy for users to wrap a BTreeMap with something that does track length, recovering the behavior for anyone who actually needs accurate lengths. I suspect that there is much higher demand for is_empty than there is for len.

  24. xosmig commented on Apr 30, 2016

    @xosmig
    Contributor

    @gereeter
    About rope: it can be implemented as BTree with an implicit key. It needs just 2 more usefull methods for BTreeMap: split_nth (splits off the first number of elements) and erase_nth.

  25. xosmig commented on Apr 30, 2016

    @xosmig
    Contributor

    @gereeter
    "First, note that option one can be slightly optimized by only storing composite sized in internal nodes."
    Yes, that's realy great point! There are about (a rough estimate, in fact, on average, more) 6 / 7 leafs. So only 1 / 7 of all nodes should contain 1 extra usize. Also there are about n / (1.5B) = n / 9 nodes on the average, where n is a number of elements in the container. So, only about 1 / 63 n extra usizes. It's only about 1 bit per element on a 64-bit machine and a half bit on a 32-bit machine. It's nothing, isn't it?
    Now the size of InternalNode is about 11 * (size_of::<V>() + size_of::<K>()) + 112 on a 64-bit architecture. So, 8 extra bytes would be okay I think.

  26. xosmig commented on Apr 30, 2016

    @xosmig
    Contributor

    If I didn't make any blunder, I am assured that each Internal node should contain itselfs subtree size.

  27. xosmig commented on Apr 30, 2016

    @xosmig
    Contributor

    There are some average values above. The worst case is: (1 / 5)n nodes and 1 / 6 internal nodes. So (1 / 30)n extra usizes, about 2 extra bits per key on a 64-bit architecture and about 1 extra bit per key on a 32-bit architecture. It's also nothing.

  28. added a commit that references this issue on Jun 2, 2016
    e752aa8
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    A-collectionsArea: `std::collections`B-RFC-approvedBlocker: Approved by a merged RFC but not yet implemented.E-easyCall for participation: Easy difficulty. Experience needed to fix: Not much. Good first issue.T-libs-api[DEPRECATED; DO NOT USE]

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions