Python 3.15 Makes UTF-8 the Default: What Analysts Should Check

Python 3.15 opens text files as UTF-8 by default. See who is affected, how to find risky open() calls and a scan for files that are not valid UTF-8.

Divya NairDivya NairAuthor9 October 20264 min read 1 views
Python 3.15 Makes UTF-8 the Default: What Analysts Should Check
In this article▾
  1. What changes
  2. Who is affected
  3. Check your code
  4. Find files that are not valid UTF-8
  5. A related trap: the byte order mark
  6. Other 3.15 changes you may meet
  7. When to move

The change in Python 3.15 that matters most to people who read files is not new syntax. It is that text files will be opened as UTF-8 unless you say otherwise, on every system. Per PEP 790, the 3.15.0 final release is scheduled for 9 October 2026, with release candidate 3 on 2 October. Python's download page still listed 3.14.8 as the latest release when I checked on 5 October, so confirm the final build on python.org. If your scripts read CSVs or text with plain open(), this is worth ten minutes.

What changes

PEP 686 makes UTF-8 mode the default. The 3.15 What's New page puts it this way: I/O operations without an explicit encoding, such as open('flying-circus.txt'), "will use UTF-8." It also gives the advice to follow: "ensure that an explicit encoding argument is always provided."

Until now, the default depended on the machine's locale. A script behaved one way on your laptop and another way on a colleague's. From 3.15 that default is fixed. PEP 686 says the change mainly affects Windows, since most Unix systems already use UTF-8 locales. You can switch the old behaviour back with PYTHONUTF8=0 or -X utf8=0.

The documentation says the new default applies only when you give no encoding argument. Anything where you have already named an encoding is unaffected.

Who is affected

You are, if your code does any of these without naming an encoding: open(), the csv module on a file handle, Path.read_text() or Path.write_text(). A file saved in a legacy code page such as Windows-1252 will now be read as UTF-8, and non-ASCII characters like "é" can raise a UnicodeDecodeError. Files you write on Windows will now be UTF-8, which older tools that assume a legacy code page may show as odd characters.

pandas is a special case. According to its documentation, pandas.read_csv already defaults to encoding='utf-8', so pd.read_csv("file.csv") does not change. It matters if you pass in a file you opened yourself with open().

Check your code

Python has had a switch for finding the risky calls since 3.10. Run a script with it and Python warns every time a text file is opened without an encoding:

python -X warn_default_encoding script.py

On Python 3.14 this raised EncodingWarning: 'encoding' argument not specified for a bare open(), which is the call you want to fix.

Then say what you mean:

import csv
from pathlib import Path

with open("sales.csv", encoding="utf-8", newline="") as f:    # explicit, same everywhere
    rows = list(csv.DictReader(f))

text = Path("old_export.csv").read_text(encoding="cp1252")    # a legacy Windows-1252 export

If you want to keep the old locale-based behaviour in a particular place, the documentation points to encoding='locale', supported since Python 3.10.

Find files that are not valid UTF-8

Before you upgrade, scan your data folder and list files that would fail. This short loop prints the ones that are not valid UTF-8:

from pathlib import Path

for p in Path("data").glob("*.csv"):
    try:
        p.read_text(encoding="utf-8")
    except UnicodeDecodeError as e:
        print(p.name, "is not valid UTF-8:", e.reason)

I tried it on one UTF-8 file and one Windows-1252 file containing "José". It passed the first and flagged the second with "invalid continuation byte". Convert flagged files once, then store them as UTF-8, rather than patching every script.

A common mistake is to "fix" the error by adding errors="ignore". That does not solve the problem. It silently drops the characters, and a name quietly loses a letter in your analysis.

Some programs write a UTF-8 file with a three-byte marker at the start, called a byte order mark. Read it as plain UTF-8 and the marker becomes an invisible character glued to your first column name. I checked this with a small file:

Path("bom.csv").read_text(encoding="utf-8").splitlines()[0]       # 'name,city'
Path("bom.csv").read_text(encoding="utf-8-sig").splitlines()[0]   # 'name,city'

That is why a column you can see on screen is "missing" in code: it is really named name. The utf-8-sig encoding strips the marker. This is not new in 3.15, but it is the next thing people hit once everything is UTF-8, so it is worth knowing before it costs you an hour.

Other 3.15 changes you may meet

The same page lists several additions. Three that analysts may notice:

  • Lazy imports (PEP 810). lazy import json delays loading a module until first use, which helps start-up time in scripts with heavy dependencies.

  • A sampling profiler (PEP 799). The new profiling.sampling tool, called Tachyon, can profile a running process at up to 1,000,000 Hz and write flame graphs.

  • A built-in frozendict (PEP 814). An immutable mapping you can use as a dictionary key or store in a set.

When to move

The What's New page I read was for release candidate 3 and still marked as a draft, so details may shift. My advice is to try 3.15 in a separate virtual environment now, run your test suite, then wait until the libraries you depend on, such as NumPy and pandas, publish builds for it. Production can wait a few months without any cost.

One more habit: add a single test row with a non-ASCII name, such as "José", to your CSV loader's test data. If an encoding change ever breaks reading or writing, that row fails loudly on the first run instead of surfacing as a strange name in a report.

The takeaway: add encoding="utf-8" to your file reads and writes this week. It is a ten-minute edit, it works on every Python version, and it makes your scripts behave the same on every machine.

Keep reading

More from Divya Nair

More from Divya Nair

More in Guides & Tutorials

Have a story of your own?

Publisha is free to start. Write with AI that keeps your voice, and publish in a click.