Skip to content

Tracking issue for char encoding methods #27784

Description

@alexcrichton

This is a tracking issue for the unstable unicode feature and the char::encode_utf{8,16} methods.

The interfaces here are a little wonky but are done for performance. It's not clear whether these need to be exported or not or if there's a better method to do so through iterators.

Activity

  1. added
    T-libs-api[DEPRECATED; DO NOT USE]
    B-unstableBlocker: Implemented in the nightly compiler and unstable.
    on Aug 13, 2015
  2. SimonSapin commented on Aug 13, 2015

    @SimonSapin
    Contributor

    How about returning enums like enum OneOrTwo { One(u16), Two(u16, u16) } or enum Utf16Encoding { SingleCodeUnit(u16), SurrogatePair(u16, u16) }?

  3. alexcrichton commented on Aug 13, 2015

    @alexcrichton
    MemberAuthor

    Certainly possible, but there's also the question of ergonomics here in terms of what to do with that after you've got the information.

  4. SimonSapin commented on Aug 13, 2015

    @SimonSapin
    Contributor

    I think that one form or another of this functionality that doesn’t require allocation should be exposed. Returning an iterator is nicer than taking &mut [_], but I don’t know about performance.

  5. nagisa commented on Aug 13, 2015

    @nagisa
    Member

    I’ve suggested taking &mut [u8; 4]/&mut [u16; 2] and returning usize once. The major downside is inability to convert from slice to array.

    Taking anything else than slice also makes following use case not as elegant as it is now:

    let mut buffer = Vec::with_capacity(alot);
    let mut idx = 0;
    loop {
         idx += some_char().encode_utf8(&mut buffer[idx..]).unwrap();
    }
  6. added a commit that references this issue on Aug 27, 2015
  7. SimonSapin commented on Oct 28, 2015

    @SimonSapin
    Contributor

    How about returning something that both is an iterator and dereferences to a slice?

    struct Utf8Char {
        bytes: [u8; 4],
        position: usize,
    }
    
    impl Deref for Utf8Char {
        type Target = [u8];
        fn deref(&self) -> &[u8] { &self.bytes[self.position..] }
    }
    
    impl Iterator for Utf8Char {
        type Item = u8;
        fn next(&mut self) -> Option<u8> {
            if self.position < self.bytes.len() {
                let byte = self.bytes[self.position];
                self.position += 1;
                Some(byte)
            } else {
                None
            }
        }
    }

    (“Short” code points have zeros as padding at the start of the array.)

    … and similarly for UTF-16, but with [u16; 2] instead of [u8; 4].

  8. BurntSushi commented on Nov 8, 2015

    @BurntSushi
    Member

    @SimonSapin That looks really sweet to me!

  9. BurntSushi commented on Jan 20, 2016

    @BurntSushi
    Member

    @SimonSapin In your deref method, I think that should be &self.bytes[..self.position], right?

  10. BurntSushi commented on Jan 20, 2016

    @BurntSushi
    Member

    @SimonSapin What do you think about also exposing decode_{utf8,utf16} methods? Basically, if you have some bytes and want the next encoded char out of it, today I think you need to decode into a string and then call chars, which is a bit roundabout (and does extra work I believe).

  11. SimonSapin commented on Jan 21, 2016

    @SimonSapin
    Contributor

    No, deref returns the slice that hasn’t been consumed by the iterator yet. For code points that have less than 4 bytes to begin with, padding is at the start of the array, not the end.

    We already have char::decode_utf16 that takes and returns iterators.

    For UTF-8 I do want to expose a decoder that’s more low-level than what we currently have, but I’m not sure what it should look like. I have some experiments at https://github.com/SimonSapin/rust-utf8

  12. BurntSushi commented on Jan 21, 2016

    @BurntSushi
    Member

    padding is at the start of the array, not the end.

    Ah! That was what I missed. Thanks for the clarification.

    I have some experiments at https://github.com/SimonSapin/rust-utf8

    Interesting. That is much more complex than I had thought it would be! (I hadn't considered returning additional info about incomplete sequences.)

  13. SimonSapin commented on Jan 21, 2016

    @SimonSapin
    Contributor

    Most of the complexity comes from self-imposed constraints:

    • Support “chunked” decoding so you can start processing, say, an HTML document before it’s finished downloading from the network. The bytes for a single char can be split across chunks.
    • Make it possible to emit &str slices that borrow &[u8] input bytes whenever possible, to avoid copying too many bytes.

    I don’t know how much of that should be in the standard library.

    But when the standard library gets performance improvement like #30740 (and perhaps more in the future with SIMD or something?), ideally they’d be in a low-level algorithm that everything else builds on top of.

  14. 82 remaining items

  15. brson commented on Nov 10, 2016

    @brson
    Contributor

    @rfcbot resolved panic-vs-not-panic

  16. Kimundi commented on Nov 11, 2016

    @Kimundi
    Contributor

    Apart from the panicking. I'm a bit confused right now about what the actual API/signature is going to be. The one that returns &mut str ?

  17. alexcrichton commented on Nov 11, 2016

    @alexcrichton
    MemberAuthor

    @Kimundi

    Yeah encode_utf8 looks like:

    fn encode_utf8(self, dst: &mut [u8]) -> &mut str

    and encode_utf16 looks like:

    fn encode_utf16(self, dst: &mut [u16]) -> &mut [u16]
  18. Kimundi commented on Nov 11, 2016

    @Kimundi
    Contributor

    Alright!

  19. rfcbot commented on Nov 12, 2016

    @rfcbot

    🔔 This is now entering its final comment period, as per the review above. 🔔

    psst @alexcrichton, I wasn't able to add the final-comment-period label, please do so.

  20. added
    final-comment-periodIn the final comment period and will be merged soon unless new substantive objections are raised.
    on Nov 12, 2016
  21. rfcbot commented on Nov 22, 2016

    @rfcbot

    The final comment period is now complete.

  22. added a commit that references this issue on Dec 26, 2016
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    B-unstableBlocker: Implemented in the nightly compiler and unstable.E-help-wantedCall for participation: Help is requested to fix this issue.T-libs-api[DEPRECATED; DO NOT USE]final-comment-periodIn the final comment period and will be merged soon unless new substantive objections are raised.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions