Skip to content

Fix UnicodeEncodeError in write_calibration_table on Windows - #32707

Open
Raaif (Raaif-Yousuf) wants to merge 1 commit into
microsoft:mainfrom
Raaif-Yousuf:fix/write-calibration-table-windows-encoding
Open

Raaif (Raaif-Yousuf) wants to merge 1 commit into
microsoft:mainfrom
Raaif-Yousuf:fix/write-calibration-table-windows-encoding

Conversation

@Raaif-Yousuf

Copy link
Copy Markdown

Description

write_calibration_table() opens calibration.json and calibration.cache with a bare open(path, "w"), so Python picks the encoding from locale.getpreferredencoding(). On Windows that is the process ANSI code page (cp1252 on an English install), not UTF-8. Tensor names are written into calibration.cache as raw, unescaped text, so any tensor name containing a character outside that code page aborts the whole export with UnicodeEncodeError.

This pins both text writes to encoding="utf-8" and adds a regression test.

Minimal repro on Windows 11 against the released onnxruntime==1.30.0 wheel:

import tempfile
import numpy as np
from onnxruntime.quantization import write_calibration_table
from onnxruntime.quantization.calibrate import TensorData

cache = {
    "中文层名": TensorData(
        lowest=np.array(0.0, dtype=np.float32),
        highest=np.array(1.0, dtype=np.float32),
    )
}
write_calibration_table(cache, dir=tempfile.mkdtemp())
  File ".../onnxruntime/quantization/quant_utils.py", line 1055, in write_calibration_table
    file.write(value)
  File ".../encodings/cp1252.py", line 19, in encode
    return codecs.charmap_encode(input, self.errors, encoding_table)[0]
UnicodeEncodeError: 'charmap' codec can't encode characters in position 0-3: character maps to <undefined>

The failure leaves a truncated calibration.cache behind, and calibration.flatbuffers is never written, so the TensorRT INT8 calibration workflow just stops.

Motivation and Context

Tensor names come straight from the ONNX graph, so they are whatever the exporter produced and are not restricted to ASCII. The calibration table is a text format consumed by other tools, so its encoding should not depend on the machine that happens to write it; UTF-8 is the only sensible fixed choice, and it is a no-op for the ASCII names that make up the common case.

calibration.json does not crash today because json.dumps defaults to ensure_ascii=True and escapes non-ASCII to \uXXXX, but it is written with the same locale-dependent encoding and is fixed here for consistency.

Verified on Windows 11 with Python 3.12 against a clean onnxruntime==1.30.0 wheel from PyPI: the new test fails with the UnicodeEncodeError above before the change, and passes with the two encoding="utf-8" arguments applied to the installed copy.

write_calibration_table() opens calibration.json and calibration.cache
with the platform default text encoding. On Windows that default is the
process ANSI code page (e.g. cp1252), not UTF-8, so any tensor name
containing a character outside that code page raises UnicodeEncodeError
while writing calibration.cache, since tensor names are written as raw
text there (the JSON file escapes non-ASCII via ensure_ascii and so
doesn't hit this, but is fixed for consistency and defense in depth).

Tensor names come straight from the ONNX graph and can be non-ASCII
(CJK names, exporter-generated identifiers with special characters,
etc.), so this reliably breaks the documented TensorRT calibration
table workflow on Windows for such models.

Pin both writes to encoding="utf-8" and add a regression test that
calls write_calibration_table with a non-ASCII tensor name and checks
both output files.
Copilot AI balanced review requested due to automatic review settings September 20, 2026 19:22
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
There may be pipelines that require an authorized user to comment /azp run to run.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The focused portability fix is correct and includes adequate regression coverage.

Review effort: Balanced
Findings: None

What changed in this PR

Pins calibration text outputs to UTF-8, preventing Windows locale-dependent UnicodeEncodeError failures.

Changes:

  • Writes JSON and cache files using UTF-8.
  • Adds regression coverage for non-ASCII tensor names.
File Description
onnxruntime/​python/​tools/​quantization/​quant_utils.py Specifies UTF-8 for calibration text files.
onnxruntime/​test/​python/​quantization/​test_calibration.py Tests Unicode tensor-name export.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@Raaif-Yousuf

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants