Skip to content

wc: count characters as bytes in single-byte locales - #15140

Draft
darkraider01 wants to merge 10 commits into
uutils:mainfrom
darkraider01:wc-c-locale-character-counts
Draft

darkraider01 wants to merge 10 commits into
uutils:mainfrom
darkraider01:wc-c-locale-character-counts

Conversation

@darkraider01

@darkraider01 darkraider01 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

In single-byte locales (where MB_CUR_MAX == 1), characters should be counted as bytes for wc -m.

Previously, this was inferred by checking if the locale was named "C" or "POSIX", which does not match GNU semantics on platforms where the C locale resolves to UTF-8, nor does it handle other single-byte locales (such as ISO-8859-1).

This updates wc to detect whether the active locale encoding is genuinely single-byte via native libc locale properties (MB_CUR_MAX == 1), counting bytes directly for -m across all dispatch paths while preserving multibyte counting for UTF-8 and other multibyte locales.

Once #15139 merges, I'll rebase this follow-up ^^

Closes #14928.

Comment thread src/uu/wc/src/wc.rs Outdated
writeln!(stdout)
}

static IS_C_LOCALE: LazyLock<bool> = LazyLock::new(|| {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Byte counting is limited to C/POSIX because other non-UTF-8 locales can still use multibyte encodings. Those locales keep their existing counting paths.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

GNU testsuite comparison:

Congrats! The gnu test tests/wc/wc is no longer failing!

@codspeed

codspeed Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Merging this PR will regress 4 benchmarks

⚡ 16 improved benchmarks
❌ 4 regressed benchmarks
✅ 310 untouched benchmarks
⏩ 123 skipped benchmarks1

Warning

Please fix the performance issues or acknowledge them on CodSpeed.

Performance Changes

Mode Benchmark BASE HEAD Efficiency
❌ Simulation five_38_bit_primes 1.7 s 2.2 s -21.78%
❌ Simulation wc_chars_large_line_count[100000] 2.3 ms 2.8 ms -18.01%
❌ Simulation sort_dictionary_order[500000] 1.9 s 2.1 s -6.54%
❌ Simulation rm_recursive_tree 13.7 ms 14.4 ms -5.02%
⚡ Simulation cut_fields_custom_delim 65.9 ms 51.9 ms +27%
⚡ Simulation cut_fields_tab 57.4 ms 45.5 ms +26.24%
⚡ Simulation cut_bytes 18.1 ms 15.5 ms +16.69%
⚡ Simulation check_sorted_utf8_locale 491.1 ms 421.8 ms +16.45%
⚡ Simulation sort_numeric_utf8_locale 40.8 ms 36.1 ms +12.98%
⚡ Simulation merge_single_file_utf8_locale 133.4 ms 119.3 ms +11.85%
⚡ Simulation cut_characters 25.4 ms 22.9 ms +10.88%
⚡ Simulation three_39_bit_primes 545.9 ms 514.8 ms +6.05%
⚡ Simulation merge_pre_sorted_files_utf8_locale 251.9 ms 238 ms +5.82%
⚡ Simulation merge_pre_sorted_files 252.5 ms 238.6 ms +5.82%
⚡ Simulation sort_german_de_locale 295 ms 282.1 ms +4.6%
⚡ Simulation sort_case_sensitive[500000] 334.9 ms 321 ms +4.32%
⚡ Simulation sort_mixed_utf8_locale 82.1 ms 79.4 ms +3.52%
⚡ Simulation sort_reverse_utf8_locale 81.9 ms 79.2 ms +3.45%
⚡ Simulation sort_spill_to_tmp_files_utf8_locale 332.1 ms 321 ms +3.45%
⚡ Simulation sort_unique_utf8_locale 84.9 ms 82.2 ms +3.33%

Tip

Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.


Comparing darkraider01:wc-c-locale-character-counts (aee391f) with main (8f6a838)2

Open in CodSpeed

Footnotes

  1. 123 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

  2. No successful run was found on main (ca3e965) during the generation of this report, so 8f6a838 was used instead as the comparison base. There might be some changes unrelated to this pull request in this report. ↩

@darkraider01
darkraider01 force-pushed the wc-c-locale-character-counts branch from 6523468 to 0a7e171 Compare October 6, 2026 17:56
@darkraider01 darkraider01 changed the title wc: count characters as bytes in C and POSIX locales wc: count characters as bytes in single-byte locales Oct 6, 2026
@sylvestre

Copy link
Copy Markdown
Contributor

sorry but it needs to be rebased

@darkraider01

Copy link
Copy Markdown
Contributor Author

right, sorry, on it !!!!

Comment thread src/uu/wc/benches/wc_bench.rs Outdated
.bench_values(|args| black_box(uumain(args)));
}

#[divan::bench(args = [100_000])]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should be done in a different pr

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

done!

@darkraider01

Copy link
Copy Markdown
Contributor Author

i'll rebase this pr again, once #15139 lands.
/cc @sylvestre

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

wc: should report correct count on multibyte input

2 participants