Repository navigation
Tracking issue for char encoding methods #27784
Description
Activity
- addedT-libs-api[DEPRECATED; DO NOT USE][DEPRECATED; DO NOT USE]B-unstableBlocker: Implemented in the nightly compiler and unstable.Blocker: Implemented in the nightly compiler and unstable.
on Aug 13, 2015 How about returning enums like
enum OneOrTwo { One(u16), Two(u16, u16) }orenum Utf16Encoding { SingleCodeUnit(u16), SurrogatePair(u16, u16) }?Certainly possible, but there's also the question of ergonomics here in terms of what to do with that after you've got the information.
I think that one form or another of this functionality that doesn’t require allocation should be exposed. Returning an iterator is nicer than taking
&mut [_], but I don’t know about performance.I’ve suggested taking
&mut [u8; 4]/&mut [u16; 2]and returningusizeonce. The major downside is inability to convert from slice to array.Taking anything else than slice also makes following use case not as elegant as it is now:
let mut buffer = Vec::with_capacity(alot); let mut idx = 0; loop { idx += some_char().encode_utf8(&mut buffer[idx..]).unwrap(); }
- added a commit that references this issue
on Aug 27, 2015 How about returning something that both is an iterator and dereferences to a slice?
struct Utf8Char { bytes: [u8; 4], position: usize, } impl Deref for Utf8Char { type Target = [u8]; fn deref(&self) -> &[u8] { &self.bytes[self.position..] } } impl Iterator for Utf8Char { type Item = u8; fn next(&mut self) -> Option<u8> { if self.position < self.bytes.len() { let byte = self.bytes[self.position]; self.position += 1; Some(byte) } else { None } } }
(“Short” code points have zeros as padding at the start of the array.)
… and similarly for UTF-16, but with
[u16; 2]instead of[u8; 4].@SimonSapin That looks really sweet to me!
@SimonSapin In your
derefmethod, I think that should be&self.bytes[..self.position], right?@SimonSapin What do you think about also exposing
decode_{utf8,utf16}methods? Basically, if you have some bytes and want the next encodedcharout of it, today I think you need to decode into a string and then callchars, which is a bit roundabout (and does extra work I believe).No,
derefreturns the slice that hasn’t been consumed by the iterator yet. For code points that have less than 4 bytes to begin with, padding is at the start of the array, not the end.We already have
char::decode_utf16that takes and returns iterators.For UTF-8 I do want to expose a decoder that’s more low-level than what we currently have, but I’m not sure what it should look like. I have some experiments at https://github.com/SimonSapin/rust-utf8
padding is at the start of the array, not the end.
Ah! That was what I missed. Thanks for the clarification.
I have some experiments at https://github.com/SimonSapin/rust-utf8
Interesting. That is much more complex than I had thought it would be! (I hadn't considered returning additional info about incomplete sequences.)
Most of the complexity comes from self-imposed constraints:
- Support “chunked” decoding so you can start processing, say, an HTML document before it’s finished downloading from the network. The bytes for a single
charcan be split across chunks. - Make it possible to emit
&strslices that borrow&[u8]input bytes whenever possible, to avoid copying too many bytes.
I don’t know how much of that should be in the standard library.
But when the standard library gets performance improvement like #30740 (and perhaps more in the future with SIMD or something?), ideally they’d be in a low-level algorithm that everything else builds on top of.
- Support “chunked” decoding so you can start processing, say, an HTML document before it’s finished downloading from the network. The bytes for a single
82 remaining items
@rfcbot resolved panic-vs-not-panic
Apart from the panicking. I'm a bit confused right now about what the actual API/signature is going to be. The one that returns
&mut str?Yeah
encode_utf8looks like:fn encode_utf8(self, dst: &mut [u8]) -> &mut str
and
encode_utf16looks like:fn encode_utf16(self, dst: &mut [u16]) -> &mut [u16]
Alright!
🔔 This is now entering its final comment period, as per the review above. 🔔
psst @alexcrichton, I wasn't able to add the
final-comment-periodlabel, please do so.- addedfinal-comment-periodIn the final comment period and will be merged soon unless new substantive objections are raised.In the final comment period and will be merged soon unless new substantive objections are raised.
on Nov 12, 2016 The final comment period is now complete.
Reacted by Michael Nitschinger
This is a tracking issue for the unstable
unicodefeature and thechar::encode_utf{8,16}methods.The interfaces here are a little wonky but are done for performance. It's not clear whether these need to be exported or not or if there's a better method to do so through iterators.