Skip to content

perf: optimize left_right in datafusion-functions - #23762

Merged
comphead merged 1 commit into
apache:mainfrom
andygrove:auto-opt/left_right-datafusion-20260721-083541
Jul 22, 2026
Merged

comphead merged 1 commit into
apache:mainfrom
andygrove:auto-opt/left_right-datafusion-20260721-083541

Conversation

@andygrove

@andygrove andygrove commented Jul 21, 2026 •

Copy link
Copy Markdown
Member

Which issue does this PR close?

N/A

Rationale for this change

Optimize existing expression.

What changes are included in this PR?

Added an ASCII fast path to left_right_byte_length so left/right compute the n-th codepoint byte offset via O(1) arithmetic instead of a per-row char_indices() UTF-8 scan.

Are these changes tested?

Existing tests.

Benchmark (criterion):

  • string_view short_result: 4.848% faster (base 56784ns -> cand 54031ns)
  • string long_result: 38.098% faster (base 141926ns -> cand 87855ns)
  • string_view long_result: 44.817% faster (base 94346ns -> cand 52063ns)
  • string short_result: 13.184% faster (base 88532ns -> cand 76860ns)
  • string_view short_result: 15.165% faster (base 57136ns -> cand 48472ns)
  • string long_result: 39.278% faster (base 140969ns -> cand 85600ns)
  • string_view long_result: 40.29% faster (base 81808ns -> cand 48847ns)
  • string short_result: 14.369% faster (base 86809ns -> cand 74336ns)

Full criterion output:

left/string short_result
                        time:   [73.406 µs 73.726 µs 74.067 µs]
                        change: [−15.026% −14.369% −13.742%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
  1 (1.00%) low mild
left/string long_result time:   [85.227 µs 85.519 µs 85.827 µs]
                        change: [−39.588% −39.278% −38.969%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 3 outliers among 100 measurements (3.00%)
  1 (1.00%) low mild
  2 (2.00%) high mild
left/string_view short_result
                        time:   [48.206 µs 48.411 µs 48.613 µs]
                        change: [−17.549% −15.165% −13.033%] (p = 0.00 < 0.05)
                        Performance has improved.
left/string_view long_result
                        time:   [48.383 µs 48.691 µs 49.038 µs]
                        change: [−40.884% −40.290% −39.714%] (p = 0.00 < 0.05)
                        Performance has improved.

right/string short_result
                        time:   [76.804 µs 76.986 µs 77.159 µs]
                        change: [−13.955% −13.184% −12.465%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 6 outliers among 100 measurements (6.00%)
  3 (3.00%) low severe
  3 (3.00%) low mild
right/string long_result
                        time:   [88.219 µs 88.597 µs 88.938 µs]
                        change: [−38.351% −38.098% −37.853%] (p = 0.00 < 0.05)
                        Performance has improved.
right/string_view short_result
                        time:   [53.380 µs 53.545 µs 53.728 µs]
                        change: [−5.5061% −4.8482% −4.1872%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 3 outliers among 100 measurements (3.00%)
  3 (3.00%) high mild
right/string_view long_result
                        time:   [52.049 µs 52.332 µs 52.633 µs]
                        change: [−45.883% −44.817% −43.850%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 6 outliers among 100 measurements (6.00%)
  6 (6.00%) high mild

Are there any user-facing changes?

No

@github-actions github-actions Bot added the functions Changes to functions implementation label Jul 21, 2026
@andygrove
andygrove marked this pull request as ready for review July 21, 2026 15:08
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 80.71%. Comparing base (050046a) to head (e7f98fa).

Additional details and impacted files
@@            Coverage Diff             @@
##             main   #23762      +/-   ##
==========================================
- Coverage   80.71%   80.71%   -0.01%     
==========================================
  Files        1089     1089              
  Lines      368750   368753       +3     
  Branches   368750   368753       +3     
==========================================
- Hits       297631   297626       -5     
- Misses      53375    53380       +5     
- Partials    17744    17747       +3     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@getChan getChan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is_ascii() is very fast here thanks to std's SIMD/word-at-a-time checks. LGTM

@comphead comphead left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @andygrove makes sense to me!

@comphead
comphead added this pull request to the merge queue Jul 22, 2026
Merged via the queue into apache:main with commit e271f65 Jul 22, 2026
38 checks passed
kosiew pushed a commit to kosiew/datafusion that referenced this pull request Aug 12, 2026
## Which issue does this PR close?

N/A

## Rationale for this change

Optimize existing expression.

## What changes are included in this PR?

Added an ASCII fast path to left_right_byte_length so left/right compute
the n-th codepoint byte offset via O(1) arithmetic instead of a per-row
char_indices() UTF-8 scan.

## Are these changes tested?

Existing tests.

Benchmark (criterion):

- string_view short_result: 4.848% faster (base 56784ns -> cand 54031ns)
- string long_result: 38.098% faster (base 141926ns -> cand 87855ns)
- string_view long_result: 44.817% faster (base 94346ns -> cand 52063ns)
- string short_result: 13.184% faster (base 88532ns -> cand 76860ns)
- string_view short_result: 15.165% faster (base 57136ns -> cand
48472ns)
- string long_result: 39.278% faster (base 140969ns -> cand 85600ns)
- string_view long_result: 40.29% faster (base 81808ns -> cand 48847ns)
- string short_result: 14.369% faster (base 86809ns -> cand 74336ns)

Full criterion output:

```text
left/string short_result
                        time:   [73.406 µs 73.726 µs 74.067 µs]
                        change: [−15.026% −14.369% −13.742%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 1 outliers among 100 measurements (1.00%)
  1 (1.00%) low mild
left/string long_result time:   [85.227 µs 85.519 µs 85.827 µs]
                        change: [−39.588% −39.278% −38.969%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 3 outliers among 100 measurements (3.00%)
  1 (1.00%) low mild
  2 (2.00%) high mild
left/string_view short_result
                        time:   [48.206 µs 48.411 µs 48.613 µs]
                        change: [−17.549% −15.165% −13.033%] (p = 0.00 < 0.05)
                        Performance has improved.
left/string_view long_result
                        time:   [48.383 µs 48.691 µs 49.038 µs]
                        change: [−40.884% −40.290% −39.714%] (p = 0.00 < 0.05)
                        Performance has improved.

right/string short_result
                        time:   [76.804 µs 76.986 µs 77.159 µs]
                        change: [−13.955% −13.184% −12.465%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 6 outliers among 100 measurements (6.00%)
  3 (3.00%) low severe
  3 (3.00%) low mild
right/string long_result
                        time:   [88.219 µs 88.597 µs 88.938 µs]
                        change: [−38.351% −38.098% −37.853%] (p = 0.00 < 0.05)
                        Performance has improved.
right/string_view short_result
                        time:   [53.380 µs 53.545 µs 53.728 µs]
                        change: [−5.5061% −4.8482% −4.1872%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 3 outliers among 100 measurements (3.00%)
  3 (3.00%) high mild
right/string_view long_result
                        time:   [52.049 µs 52.332 µs 52.633 µs]
                        change: [−45.883% −44.817% −43.850%] (p = 0.00 < 0.05)
                        Performance has improved.
Found 6 outliers among 100 measurements (6.00%)
  6 (6.00%) high mild
```


## Are there any user-facing changes?

No

<!--
If there are user-facing changes then we may require documentation to be
updated before approving the PR.
-->

<!--
If there are any breaking changes to public APIs, please add the `api
change` label.
-->
neilconway added a commit to neilconway/datafusion that referenced this pull request Oct 5, 2026
apache#23762 added an ASCII fast path to `left_right_byte_length` that calls
`is_ascii()` on the whole string for every row. For short results from
longer strings, such as Utf8View data read from Parquet, that scan costs
more than the per-character scan it replaced.

ASCII bytes are never part of a multi-byte UTF-8 sequence, so if the `n`
bytes at the relevant end of the string are ASCII, they are exactly the
`n` characters at that end. Check only those bytes, and fall back to the
per-character scan otherwise.
joroKr21 pushed a commit to coralogix/arrow-datafusion that referenced this pull request Oct 5, 2026
## Which issue does this PR close?

- N/A

## Rationale for this change

apache#23762 added an ASCII fast-path for `left` and `right`. That improved
performance in many scenarios, but the implementation calls `is_ascii`
on every input string. That is expensive, particularly for the common
case that `left` and `right` are used to compute a small prefix/suffix
of a much longer string. That check is also overly conservative: for
example, we can take the ASCII fast-path for `left(s, k)` if the first
`k` bytes in the string are ASCII, even if there are multibyte
characters elsewhere in the string.

This PR implements two optimizations:

1. Only call `is_ascii` on the bytes necessary to determine if we can
take the fast-path, not the entire input string, as described above.
2. Benchmarking identified that for `Utf8View` inputs, the first
optimization regressed some benchmark cases (e.g., `negative_n`) on both
ARM and x86. Claude's theory is that `make_view` is out-of-line and does
an indirect jump on the length of the result string; it seems that after
implementing the first optimization, this jump was not well-handled by
the branch predictor. Instead, we add a helper `sub_view` that returns a
view that is a substring of an existing view. This can be inlined and
avoids the indirect jump incurred by `make_view`; it can also construct
the new view from the old view with bitwise ops, rather than building
the new view on the stack.

Benchmarks:

x86 (AMD EPYC Milan)
- left Utf8 long_result: 101.7 → 93.4 (−8%)
- left Utf8 n_exceeds_len: 91.7 → 93.8 (+2%)
- left Utf8 negative_n: 90.4 → 83.2 (−8%)
- left Utf8 per_row_n: 93.9 → 89.5 (−5%)
- left Utf8 short_result: 91.0 → 82.6 (−9%)
- left Utf8 short_result_long_input: 310.4 → 89.4 (−71%)
- left Utf8View long_result: 53.6 → 40.9 (−24%)
- left Utf8View n_exceeds_len: 55.9 → 50.7 (−9%)
- left Utf8View negative_n: 58.8 → 45.2 (−23%)
- left Utf8View per_row_n: 89.0 → 45.2 (−49%)
- left Utf8View short_result: 54.2 → 47.5 (−12%)
- left Utf8View short_result_long_input: 249.1 → 52.7 (−79%)
- right Utf8 long_result: 109.4 → 91.8 (−16%)
- right Utf8 n_exceeds_len: 97.9 → 98.5 (+1%)
- right Utf8 negative_n: 91.5 → 84.7 (−7%)
- right Utf8 per_row_n: 100.7 → 88.1 (−13%)
- right Utf8 short_result: 96.8 → 87.2 (−10%)
- right Utf8 short_result_long_input: 261.2 → 95.5 (−63%)
- right Utf8View long_result: 64.2 → 53.8 (−16%)
- right Utf8View n_exceeds_len: 63.9 → 56.0 (−12%)
- right Utf8View negative_n: 62.3 → 55.8 (−10%)
- right Utf8View per_row_n: 101.0 → 54.9 (−46%)
- right Utf8View short_result: 62.9 → 58.7 (−7%)
- right Utf8View short_result_long_input: 202.2 → 67.7 (−67%)

ARM (Apple M4 Max)
- left Utf8 long_result: 91.3 → 62.6 (−31%)
- left Utf8 n_exceeds_len: 69.4 → 71.4 (+3%)
- left Utf8 negative_n: 66.8 → 69.5 (+4%)
- left Utf8 per_row_n: 66.6 → 60.6 (−9%)
- left Utf8 short_result: 69.1 → 62.7 (−9%)
- left Utf8 short_result_long_input: 91.1 → 59.6 (−35%)
- left Utf8View long_result: 45.4 → 26.0 (−43%)
- left Utf8View n_exceeds_len: 37.6 → 34.4 (−9%)
- left Utf8View negative_n: 37.8 → 25.9 (−31%)
- left Utf8View per_row_n: 68.1 → 29.4 (−57%)
- left Utf8View short_result: 37.9 → 25.0 (−34%)
- left Utf8View short_result_long_input: 61.1 → 28.9 (−53%)
- right Utf8 long_result: 104.2 → 64.7 (−38%)
- right Utf8 n_exceeds_len: 80.1 → 79.1 (−1%)
- right Utf8 negative_n: 71.4 → 66.0 (−8%)
- right Utf8 per_row_n: 74.0 → 64.3 (−13%)
- right Utf8 short_result: 80.2 → 65.4 (−18%)
- right Utf8 short_result_long_input: 101.1 → 63.4 (−37%)
- right Utf8View long_result: 62.7 → 32.0 (−49%)
- right Utf8View n_exceeds_len: 42.0 → 38.5 (−8%)
- right Utf8View negative_n: 41.3 → 28.5 (−31%)
- right Utf8View per_row_n: 72.0 → 33.5 (−53%)
- right Utf8View short_result: 43.9 → 30.0 (−32%)
- right Utf8View short_result_long_input: 63.7 → 33.3 (−48%)

## What changes are included in this PR?

See above.

## What is the testing strategy for this PR?

Existing tests pass; no functional changes.

## Are there any user-facing changes?

No.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

functions Changes to functions implementation v55.0.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants