Repository navigation
wc: count characters as bytes in single-byte locales - #15140
darkraider01 wants to merge 10 commits into
Conversation
| writeln!(stdout) | ||
| } | ||
|
|
||
| static IS_C_LOCALE: LazyLock<bool> = LazyLock::new(|| { |
There was a problem hiding this comment.
Byte counting is limited to C/POSIX because other non-UTF-8 locales can still use multibyte encodings. Those locales keep their existing counting paths.
|
GNU testsuite comparison: |
Merging this PR will regress 4 benchmarks
Warning Please fix the performance issues or acknowledge them on CodSpeed. Performance Changes
Tip Investigate this regression by commenting Comparing Footnotes
|
6523468 to
0a7e171
Compare
|
sorry but it needs to be rebased |
|
right, sorry, on it !!!! |
| .bench_values(|args| black_box(uumain(args))); | ||
| } | ||
|
|
||
| #[divan::bench(args = [100_000])] |
There was a problem hiding this comment.
Should be done in a different pr
8c5a337 to
aee391f
Compare
|
i'll rebase this pr again, once #15139 lands. |
In single-byte locales (where MB_CUR_MAX == 1), characters should be counted as bytes for
wc -m.Previously, this was inferred by checking if the locale was named "C" or "POSIX", which does not match GNU semantics on platforms where the C locale resolves to UTF-8, nor does it handle other single-byte locales (such as ISO-8859-1).
This updates
wcto detect whether the active locale encoding is genuinely single-byte via native libc locale properties (MB_CUR_MAX == 1), counting bytes directly for-macross all dispatch paths while preserving multibyte counting for UTF-8 and other multibyte locales.Once #15139 merges, I'll rebase this follow-up ^^
Closes #14928.