Repository navigation
Tracking issue for collections reform part 2 (RFC 509) #19986
Description
Activity
- addedB-RFC-approvedBlocker: Approved by a merged RFC but not yet implemented.Blocker: Approved by a merged RFC but not yet implemented.
on Dec 18, 2014 cc @gankro
TODO:
Easy
- Deprecate Vec::from_elem, Vec::from_fn, Vec::grow, Vec::grow_fn; in favour of iterator adaptors
- Deprecate the ListInsertion trait defined in DList in favour of an inherent method on its iterator
- Add reserve_len to VecMap
- Add resize to Vec - csouth3
- misc stabilization - At least partially done, need to look at it again
Moderate
- add append to:
- DList
- Vec
- RingBuf
- BTreeMap
- BTreeSet
- Bitv
- BitvSet
- VecMap
- add split_off:
- DList
- Vec
- RingBuf
- BTreeMap
- BTreeSet
- Bitv
- BitvSet
- VecMap
- Upgrade the Entry API - has PR
- Add truncate and resize to RingBuf (not specified, but makes sense to do) - has PR
- Add swap_pop to RingBuf (not specified, but makes sense to do) - has PR
Note: I'm about to post a PR that does the deprecation for
Vecmethods.cc @bfops I think you were working on upgrading the entry api?
I can take care of resize for Vec.
Yeah I'm working on the entry API. Seems to be coming along!
I'll start work on these methods for RingBuf:
Add truncate and resize to RingBuf (not specified, but makes sense to do)
Add swap_pop to RingBuf (not specified, but makes sense to do)Did you mean
swap_remove?Yes, sorry.
I'm starting the misc stabilization, should be over in a 2-3 days :)
So, some of the misc stuff is done in #20053, but I didn't touch
Bitv::{get,set}, didn't touchHashMap's andHashSet's iterators (either the documentation, or actually implementingDoubleEndedIteratorandExactSizeIteratorfor all of the appropriate iterators using thenext_backimplemented asnextstrategy noted in the RFC (if this is even what the RFC intends)), and I didn't attempt the tedium of movingstd::vectostd::collections::vec, sooo....there's certainly plenty of misc left to do but wanted to give you heads up so your time doesn't end up getting wasted duplicating work :)thx :)
Oh, did you mean
swap_back_removeandswap_front_remove? What aboutgrowfor ringbuf?51 remaining items
I've started working on
appendandsplit_offmethods forBTreeMapandBTreeSet.@apasel422 OK. Then only
split_off.Hello!
I've faced with problems.
To recalc sizes of trees aftersplit_offusing O(log n) operations, it's necessary to know exact size of each subtree.
So, there are some different solutions:- add
sizefield to theLeafNodestructure.
I think, benefits are great.
First, it lets us to do very flexible splits and merges with only O(log n) operations.
Now it's possible to merge trees using O(log n) operations, but it it not so interesting without fast split.Second, some other interesting and usefull functions need it.
For example:kth_elementor "get the number of elements between two values". Both with O(log n) operations.Third, as I understand (but i am not totally sure), it lets us to implement the
rope( https://en.wikipedia.org/wiki/Rope_(data_structure) ) in easy way. Just as BTreeMap<(), V>.But overhead is notable (4 or 8 bytes for each node, depending on machine word size). I think, it is possible to embed this sizes without memory overhead for 64-bit machines (due to aligning).
- Implement
split_offwith O(log n + "size of the one of the results")
2.1) Split with O(log n) operations and then calculate size of the one of the results.
2.2) There is another algorithm which doesn't use the fast algorithm with the same complexity, but (as I understand) with greater constant.
- Implement
split_offwith O(log n + "size of the smallest one") but with greater constant
Split using O(log n) operations and then calculate sizes of the trees via 2 alternating iterators, but stop when one of the sizes is counted.
Which implementation should I use?
- add
@apasel422, what do you think?
@xosmig This isn't my area of expertise, but I'd recommend opening a discussion thread on https://internals.rust-lang.org/ or submitting a pull request for whichever option is easiest and CCing @jooert and @gereeter. You might also want to take a look at #26227.
@apasel422 The main question is "Is this overhead okay or not?". So, who can decide it?
First, note that option one can be slightly optimized by only storing composite sized in internal nodes.
I'm not a big fan of option one - I expect
split_offto be a fairly rare operation, and slowing down everything else and increasing the memory footprint solely for its benefit does not seem like a good idea. Since we explicitly support large collections (as large as can fit in memory, basically), the size field would have to be ausize, which cannot be combined with other fields. The application to ropes is interesting, but really wants a separate implementation - theBTreeMapinterface is totally wrong for the job (though the node module might be useful).kth_elementand measuring the length of a range seems useful, but again, I see them as niche use cases that really want a specialized sequence data structure.I don't have strong feelings between options two and three. I think that a variation on option two, where the choice of which side to count is made based on which edge was descended at the root node, would probably be the best. It doesn't have the overhead of option three, but still wont take long if you split near the edges.
There is also a fourth option, namely dropping support for constant time calls to
len(). This is a somewhat more radical change, but it would easily allow forsplit_offinO(log n)and avoid the question of how to calculate size completely. Moreover, while it might inconvenience users, it is easy for users to wrap a BTreeMap with something that does track length, recovering the behavior for anyone who actually needs accurate lengths. I suspect that there is much higher demand foris_emptythan there is forlen.@gereeter
Aboutrope: it can be implemented asBTreewith an implicit key. It needs just 2 more usefull methods forBTreeMap:split_nth(splits off the firstnumberof elements) anderase_nth.@gereeter
"First, note that option one can be slightly optimized by only storing composite sized in internal nodes."
Yes, that's realy great point! There are about (a rough estimate, in fact, on average, more)6 / 7leafs. So only1 / 7of all nodes should contain 1 extra usize. Also there are about n / (1.5B) = n / 9 nodes on the average, where n is a number of elements in the container. So, only about 1 / 63 n extra usizes. It's only about 1 bit per element on a 64-bit machine and a half bit on a 32-bit machine. It's nothing, isn't it?
Now the size ofInternalNodeis about11 * (size_of::<V>() + size_of::<K>()) + 112on a 64-bit architecture. So, 8 extra bytes would be okay I think.If I didn't make any blunder, I am assured that each Internal node should contain itselfs subtree size.
There are some average values above. The worst case is:
(1 / 5)nnodes and1 / 6internal nodes. So(1 / 30)nextra usizes, about 2 extra bits per key on a 64-bit architecture and about 1 extra bit per key on a 32-bit architecture. It's also nothing.- added a commit that references this issue
on Jun 2, 2016
RFC 509