Fix UnicodeEncodeError in write_calibration_table on Windows - #32707
Open
Raaif (Raaif-Yousuf) wants to merge 1 commit into
Open
Raaif (Raaif-Yousuf) wants to merge 1 commit into
Raaif (Raaif-Yousuf) wants to merge 1 commit into
Conversation
write_calibration_table() opens calibration.json and calibration.cache with the platform default text encoding. On Windows that default is the process ANSI code page (e.g. cp1252), not UTF-8, so any tensor name containing a character outside that code page raises UnicodeEncodeError while writing calibration.cache, since tensor names are written as raw text there (the JSON file escapes non-ASCII via ensure_ascii and so doesn't hit this, but is fixed for consistency and defense in depth). Tensor names come straight from the ONNX graph and can be non-ASCII (CJK names, exporter-generated identifiers with special characters, etc.), so this reliably breaks the documented TensorRT calibration table workflow on Windows for such models. Pin both writes to encoding="utf-8" and add a regression test that calls write_calibration_table with a non-ASCII tensor name and checks both output files.
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
Contributor
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The focused portability fix is correct and includes adequate regression coverage.
Review effort: Balanced
Findings: None
What changed in this PR
Pins calibration text outputs to UTF-8, preventing Windows locale-dependent UnicodeEncodeError failures.
Changes:
- Writes JSON and cache files using UTF-8.
- Adds regression coverage for non-ASCII tensor names.
| File | Description |
|---|---|
onnxruntime/python/tools/quantization/quant_utils.py |
Specifies UTF-8 for calibration text files. |
onnxruntime/test/python/quantization/test_calibration.py |
Tests Unicode tensor-name export. |
💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Xavier Dupré (xadupre)
approved these changes
Sep 21, 2026
Author
|
@microsoft-github-policy-service agree |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
write_calibration_table()openscalibration.jsonandcalibration.cachewith a bareopen(path, "w"), so Python picks the encoding fromlocale.getpreferredencoding(). On Windows that is the process ANSI code page (cp1252 on an English install), not UTF-8. Tensor names are written intocalibration.cacheas raw, unescaped text, so any tensor name containing a character outside that code page aborts the whole export withUnicodeEncodeError.This pins both text writes to
encoding="utf-8"and adds a regression test.Minimal repro on Windows 11 against the released
onnxruntime==1.30.0wheel:The failure leaves a truncated
calibration.cachebehind, andcalibration.flatbuffersis never written, so the TensorRT INT8 calibration workflow just stops.Motivation and Context
Tensor names come straight from the ONNX graph, so they are whatever the exporter produced and are not restricted to ASCII. The calibration table is a text format consumed by other tools, so its encoding should not depend on the machine that happens to write it; UTF-8 is the only sensible fixed choice, and it is a no-op for the ASCII names that make up the common case.
calibration.jsondoes not crash today becausejson.dumpsdefaults toensure_ascii=Trueand escapes non-ASCII to\uXXXX, but it is written with the same locale-dependent encoding and is fixed here for consistency.Verified on Windows 11 with Python 3.12 against a clean
onnxruntime==1.30.0wheel from PyPI: the new test fails with theUnicodeEncodeErrorabove before the change, and passes with the twoencoding="utf-8"arguments applied to the installed copy.