Skip to content

Fix/pr37 artifact actions v4 - #38

Merged
wbarnha merged 10 commits into
masterfrom
fix/pr37-artifact-actions-v4
Jul 13, 2026
Merged

wbarnha merged 10 commits into
masterfrom
fix/pr37-artifact-actions-v4

Conversation

@wbarnha

@wbarnha wbarnha commented Jul 7, 2026

Copy link
Copy Markdown
Member

No description provided.

mcepl and others added 4 commits June 25, 2026 10:23
- use setuptools.command.build_ext instead of the already dead distutils
- Update opensuse-build.patch to explicitly link against libstdc++
  fixing build failure under Python 3.14 (undefined symbol
  _ZTVN10__cxxabiv117__class_type_infoE)

Fixes: #36
Copilot AI and others added 5 commits July 11, 2026 14:23
…ng bug

v2.15.0's GitHub Action fails on windows-2022 runners with
Invalid --only='""' because PowerShell 7.4 changed how empty-string
CLI args are quoted (pypa/cibuildwheel#1740). Fixed upstream in v2.16.5.
PyPy wheels on these platforms aren't reliably buildable; only build
PyPy on Linux. Trims the now-unreachable macOS branch out of the
setuptools<72 PyPy override accordingly.
@wbarnha
wbarnha merged commit 7940bad into master Jul 13, 2026
6 checks passed
@wbarnha
wbarnha deleted the fix/pr37-artifact-actions-v4 branch July 13, 2026 04:06
wbarnha pushed a commit that referenced this pull request Jul 22, 2026
The issue #57 CC-News corpus is almost entirely valid UTF-8, so cChardet's
UTF-8 fast path answers before the detection engine ever runs. That left the
exact code the encoding-only overlay changes -- and the freedesktop uchardet
engine that regressed in work item #38 -- unexercised by CI, which is why the
earlier overlay's UTF-8 mislabeling of GBK/EUC-KR text was not caught.

Add a deterministic, network-free non-UTF-8 corpus and wire it into both the
test suite and the benchmark job:

- benchmarks/make_nonutf8_corpus.py generates labeled single- and multi-byte
  non-UTF-8 documents (ISO-8859-15, Windows-1251, Shift_JIS, EUC-JP, GBK,
  Big5, EUC-KR) from short committed multilingual snippets.
- benchmark.py scores detection by decode-equivalence when the corpus is
  labeled, and counts UTF-8 mislabels (non-UTF-8 bytes reported as UTF-8 --
  the dangerous case that makes a downstream open(encoding=...) mojibake).
- compare.py renders the accuracy column and gains --max-utf8-mislabel-rate;
  the plain-list CC-News path is unchanged.
- benchmark.yml runs the three builds over the non-UTF-8 corpus and gates on
  a UTF-8 mislabel rate of 0.02 plus a lenient 0.50 throughput floor. The
  freedesktop engine is inherently slower than the old PyYoshi engine on
  non-UTF-8, so the mislabel rate -- not raw throughput -- is the signal that
  uniquely catches the overlay regression, while the throughput floor still
  trips if the language-model slowdown is reintroduced.
- src/tests/test_nonutf8_detection.py asserts, on every PR, that no multibyte
  document is mislabeled as UTF-8, that multibyte accuracy stays >= 99%, and
  that the issue #33 Big5 false positive stays suppressed.

Verified against the pre-fix overlay: the new test and the benchmark gate
both fail on it (multibyte UTF-8 mislabels, 18.7% overall mislabel rate) and
both pass on the current overlay (0 mislabels). Full suite: 130 passed,
1 skipped.
wbarnha pushed a commit that referenced this pull request Jul 22, 2026
The first CI run measured the current build at 0.51x of v2.2.1 on the
non-UTF-8 corpus, right at the 0.50 floor -- flaky. That corpus of many
small documents does not reproduce uchardet #38's large-stream
language-model blow-up (the per-call overhead dominates, so the language
pass costs only ~1.6x here, not orders of magnitude), so the throughput
ratio cannot robustly separate a good build from a regressed one. Drop the
floor to 0.35 as a loose backstop for a catastrophic slowdown; the UTF-8
mislabel rate remains the primary, crisp regression signal.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants