Fix/pr37 artifact actions v4 - #38
Merged
Merged
Conversation
- use setuptools.command.build_ext instead of the already dead distutils - Update opensuse-build.patch to explicitly link against libstdc++ fixing build failure under Python 3.14 (undefined symbol _ZTVN10__cxxabiv117__class_type_infoE) Fixes: #36
…ng bug v2.15.0's GitHub Action fails on windows-2022 runners with Invalid --only='""' because PowerShell 7.4 changed how empty-string CLI args are quoted (pypa/cibuildwheel#1740). Fixed upstream in v2.16.5.
PyPy wheels on these platforms aren't reliably buildable; only build PyPy on Linux. Trims the now-unreachable macOS branch out of the setuptools<72 PyPy override accordingly.
wbarnha
pushed a commit
that referenced
this pull request
Jul 22, 2026
The issue #57 CC-News corpus is almost entirely valid UTF-8, so cChardet's UTF-8 fast path answers before the detection engine ever runs. That left the exact code the encoding-only overlay changes -- and the freedesktop uchardet engine that regressed in work item #38 -- unexercised by CI, which is why the earlier overlay's UTF-8 mislabeling of GBK/EUC-KR text was not caught. Add a deterministic, network-free non-UTF-8 corpus and wire it into both the test suite and the benchmark job: - benchmarks/make_nonutf8_corpus.py generates labeled single- and multi-byte non-UTF-8 documents (ISO-8859-15, Windows-1251, Shift_JIS, EUC-JP, GBK, Big5, EUC-KR) from short committed multilingual snippets. - benchmark.py scores detection by decode-equivalence when the corpus is labeled, and counts UTF-8 mislabels (non-UTF-8 bytes reported as UTF-8 -- the dangerous case that makes a downstream open(encoding=...) mojibake). - compare.py renders the accuracy column and gains --max-utf8-mislabel-rate; the plain-list CC-News path is unchanged. - benchmark.yml runs the three builds over the non-UTF-8 corpus and gates on a UTF-8 mislabel rate of 0.02 plus a lenient 0.50 throughput floor. The freedesktop engine is inherently slower than the old PyYoshi engine on non-UTF-8, so the mislabel rate -- not raw throughput -- is the signal that uniquely catches the overlay regression, while the throughput floor still trips if the language-model slowdown is reintroduced. - src/tests/test_nonutf8_detection.py asserts, on every PR, that no multibyte document is mislabeled as UTF-8, that multibyte accuracy stays >= 99%, and that the issue #33 Big5 false positive stays suppressed. Verified against the pre-fix overlay: the new test and the benchmark gate both fail on it (multibyte UTF-8 mislabels, 18.7% overall mislabel rate) and both pass on the current overlay (0 mislabels). Full suite: 130 passed, 1 skipped.
wbarnha
pushed a commit
that referenced
this pull request
Jul 22, 2026
The first CI run measured the current build at 0.51x of v2.2.1 on the non-UTF-8 corpus, right at the 0.50 floor -- flaky. That corpus of many small documents does not reproduce uchardet #38's large-stream language-model blow-up (the per-call overhead dominates, so the language pass costs only ~1.6x here, not orders of magnitude), so the throughput ratio cannot robustly separate a good build from a regressed one. Drop the floor to 0.35 as a loose backstop for a catastrophic slowdown; the UTF-8 mislabel rate remains the primary, crisp regression signal.
This was referenced Jul 22, 2026
This was referenced Aug 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.