Flower Field explains char different from Rust docs

In the Flower Field exercise, under Performance hint, it is suggested that due to char being utf8, there is some (runtime?) check to see if a char is 1, 2, 3 or 4 times a u8. But the Rust docs say:

char is guaranteed to have the same size, alignment, and function call ABI as u32 on all platforms.

and

char is always four bytes in size.

Maybe it was meant there is a performance cost for checking that every char is a valid Unicode code point, but the &str is already guaranteed to be valid Unicode so that check has already been done anyway.

Possibly the Performance hint meant something else, that would impact performance. At least for me this was not clear. If something else was meant, what is it?

What’s meant by “char” in this context is not the Rust type char, which is indeed always 4 bytes. It’s a character as stored in the string, which, according to utf8 can have different lengths.

The input.chars() iterator yields chars. To do that, It actually has to read the utf8 encoded string and decide at which byte offsets the character boundaries are. This has a runtime cost.

Hopefully this PR should avoid confusion in the future.

1 Like

You may refer to UTF8 code point sequence instead of characters, as there is more to a code point than being “a character”. But that might be too specific for an append file.