Skip to content

Avoid UTF-8 re-encoding in StringIO line reads - #15957

Merged
josevalim merged 1 commit into
elixir-lang:mainfrom
preciz:optimize-string-io-line-scanning
Oct 1, 2026
Merged

josevalim merged 1 commit into
elixir-lang:mainfrom
preciz:optimize-string-io-line-scanning

Conversation

@preciz

@preciz preciz commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Assisted-by: Codex CLI:GPT 6

Same optimization logic as in #15956 applied to lines.
Faster and less memory usage.

Mix.install([:benchee])

defmodule OldStringIO do
  def read_line(input, encoding) do
    case bytes_until_eol(input, encoding, 0) do
      {:split, 0} ->
        {:eof, ""}

      {:split, count} ->
        :erlang.split_binary(input, count)

      {:replace_split, count} ->
        {result, remainder} = :erlang.split_binary(input, count)
        result = binary_part(result, 0, byte_size(result) - 2) <> "\n"
        {result, remainder}

      :error ->
        {{:error, :collect_line}, input}
    end
  end

  defp bytes_until_eol("", _, count), do: {:split, count}
  defp bytes_until_eol(<<"\r\n"::binary, _::binary>>, _, count), do: {:replace_split, count + 2}
  defp bytes_until_eol(<<"\n"::binary, _::binary>>, _, count), do: {:split, count + 1}

  defp bytes_until_eol(<<head::utf8, tail::binary>>, :unicode, count) do
    bytes_until_eol(tail, :unicode, count + byte_size(<<head::utf8>>))
  end

  defp bytes_until_eol(<<_, tail::binary>>, :latin1, count) do
    bytes_until_eol(tail, :latin1, count + 1)
  end

  defp bytes_until_eol(<<_::binary>>, _, _), do: :error
end

defmodule NewStringIO do
  def read_line(input, encoding) do
    case bytes_until_eol(input, encoding, byte_size(input)) do
      {:split, 0} ->
        {:eof, ""}

      {:split, count} ->
        :erlang.split_binary(input, count)

      {:replace_split, count} ->
        {result, remainder} = :erlang.split_binary(input, count)
        result = binary_part(result, 0, byte_size(result) - 2) <> "\n"
        {result, remainder}

      :error ->
        {{:error, :collect_line}, input}
    end
  end

  defp bytes_until_eol("", _, size), do: {:split, size}

  defp bytes_until_eol(<<"\r\n"::binary, rest::binary>>, _, size),
    do: {:replace_split, size - byte_size(rest)}

  defp bytes_until_eol(<<"\n"::binary, rest::binary>>, _, size),
    do: {:split, size - byte_size(rest)}

  defp bytes_until_eol(<<_::utf8, rest::binary>>, :unicode, size) do
    bytes_until_eol(rest, :unicode, size)
  end

  defp bytes_until_eol(<<_, rest::binary>>, :latin1, size) do
    bytes_until_eol(rest, :latin1, size)
  end

  defp bytes_until_eol(<<_::binary>>, _, _), do: :error
end

defmodule StringIOLineBenchmark do
  def read_all(module, {input, encoding}) do
    read_lines(module, input, encoding, 0, 0)
  end

  defp read_lines(module, input, encoding, lines, bytes) do
    case module.read_line(input, encoding) do
      {:eof, ""} ->
        {lines, bytes}

      {{:error, reason}, _rest} ->
        raise "read failed: #{inspect(reason)}"

      {line, rest} ->
        read_lines(module, rest, encoding, lines + 1, bytes + byte_size(line))
    end
  end
end

inputs = %{
  "ASCII short lines" => {String.duplicate("hello world\n", 1000), :unicode},
  "ASCII long lines" => {String.duplicate(String.duplicate("a", 10_000) <> "\n", 100), :unicode},
  "UTF-8 short CRLF lines" => {String.duplicate("é🚀日\r\n", 1000), :unicode},
  "UTF-8 long lines" => {String.duplicate(String.duplicate("é🚀日", 3333) <> "\n", 100), :unicode},
  "Latin-1 short CRLF lines" => {String.duplicate(<<255, 128, ?\r, ?\n>>, 1000), :latin1},
  "Latin-1 long lines" =>
    {String.duplicate(String.duplicate(<<255>>, 10_000) <> "\n", 100), :latin1},
  "blank lines" => {String.duplicate("\n\r\n", 1000), :unicode}
}

Benchee.run(
  %{
    "old" => fn input -> StringIOLineBenchmark.read_all(OldStringIO, input) end,
    "new" => fn input -> StringIOLineBenchmark.read_all(NewStringIO, input) end
  },
  inputs: inputs,
  pre_check: :all_same,
  warmup: 1,
  time: 2,
  memory_time: 2
)

Results:

Operating System: Linux
CPU Information: AMD Ryzen 7 8845HS w
Number of Available Cores: 16
Available memory: 54.72 GB
Elixir 1.20.4
Erlang 29.0.5
JIT enabled: true

Benchmark suite executing with the following configuration:
warmup: 1 s
time: 2 s
memory time: 2 s
reduction time: 0 ns
parallel: 1
inputs: ASCII long lines, ASCII short lines, Latin-1 long lines, Latin-1 short CRLF lines, UTF-8 long lines, UTF-8 short CRLF lines, blank lines
Estimated total run time: 1 min 10 s
Excluding outliers: false

##### With input ASCII long lines #####
Name           ips        average  deviation         median         99th %
new         330.53        3.03 ms     ±3.87%        3.03 ms        3.25 ms
old          71.42       14.00 ms    ±26.97%       12.24 ms       24.54 ms

Comparison: 
new         330.53
old          71.42 - 4.63x slower +10.98 ms

Memory usage statistics:

Name    Memory usage
new        0.0158 MB
old         22.90 MB - 1450.95x memory usage +22.89 MB

**All measurements for memory usage were the same**

##### With input ASCII short lines #####
Name           ips        average  deviation         median         99th %
new        11.27 K       88.72 μs    ±18.06%       87.74 μs       99.90 μs
old         6.40 K      156.13 μs    ±11.99%      154.09 μs      175.05 μs

Comparison: 
new        11.27 K
old         6.40 K - 1.76x slower +67.41 μs

Memory usage statistics:

Name    Memory usage
new        156.20 KB
old        414.10 KB - 2.65x memory usage +257.91 KB

**All measurements for memory usage were the same**

##### With input Latin-1 long lines #####
Name           ips        average  deviation         median         99th %
new         190.21        5.26 ms     ±2.72%        5.24 ms        6.57 ms
old         153.72        6.51 ms     ±2.84%        6.47 ms        7.23 ms

Comparison: 
new         190.21
old         153.72 - 1.24x slower +1.25 ms

Memory usage statistics:

Name    Memory usage
new         16.16 KB
old         16.16 KB - 1.00x memory usage +0 KB

**All measurements for memory usage were the same**

##### With input Latin-1 short CRLF lines #####
Name           ips        average  deviation         median         99th %
new         6.76 K      148.02 μs     ±8.19%      147.13 μs      192.23 μs
old         6.22 K      160.67 μs    ±15.74%      155.87 μs      261.65 μs

Comparison: 
new         6.76 K
old         6.22 K - 1.09x slower +12.65 μs

Memory usage statistics:

Name    Memory usage
new        234.02 KB
old        234.02 KB - 1.00x memory usage +0 KB

**All measurements for memory usage were the same**

##### With input UTF-8 long lines #####
Name           ips        average  deviation         median         99th %
new         119.32        8.38 ms    ±18.91%        7.58 ms       11.97 ms
old          60.26       16.59 ms     ±2.23%       16.54 ms       17.28 ms

Comparison: 
new         119.32
old          60.26 - 1.98x slower +8.21 ms

Memory usage statistics:

Name    Memory usage
new        0.0158 MB
old         22.90 MB - 1450.80x memory usage +22.89 MB

**All measurements for memory usage were the same**

##### With input UTF-8 short CRLF lines #####
Name           ips        average  deviation         median         99th %
old         5.68 K      176.13 μs    ±22.17%      173.71 μs      190.68 μs
new         5.63 K      177.63 μs    ±16.99%      164.64 μs      287.07 μs

Comparison: 
old         5.68 K
new         5.63 K - 1.01x slower +1.50 μs

Memory usage statistics:

Name    Memory usage
old        311.78 KB
new        242.22 KB - 0.78x memory usage -69.56250 KB

**All measurements for memory usage were the same**

##### With input blank lines #####
Name           ips        average  deviation         median         99th %
new         5.10 K      196.23 μs     ±6.86%      194.54 μs      239.89 μs
old         5.02 K      199.27 μs    ±13.42%      196.23 μs      271.92 μs

Comparison: 
new         5.10 K
old         5.02 K - 1.02x slower +3.04 μs

Memory usage statistics:

Name         average  deviation         median         99th %
new        380.88 KB     ±0.00%      380.88 KB      380.88 KB
old        380.87 KB     ±0.00%      380.88 KB      380.88 KB

Comparison: 
new        380.88 KB
old        380.87 KB - 1.00x memory usage -0.00003 KB

Carry the original input byte size while scanning and derive the consumed size from the remaining suffix.

Assisted-by: Codex:GPT-6
@josevalim
josevalim merged commit 73b3325 into elixir-lang:main Oct 1, 2026
15 checks passed
@josevalim

Copy link
Copy Markdown
Member

💚 💙 💜 💛 ❤️

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

2 participants