Skip to content

perf: Optimize left, right - #26039

Merged
neilconway merged 3 commits into
apache:mainfrom
neilconway:neilc/perf-left-right-ascii-prefix
Oct 5, 2026
Merged

neilconway merged 3 commits into
apache:mainfrom
neilconway:neilc/perf-left-right-ascii-prefix

Conversation

@neilconway

@neilconway neilconway commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Which issue does this PR close?

  • N/A

Rationale for this change

#23762 added an ASCII fast-path for left and right. That improved performance in many scenarios, but the implementation calls is_ascii on every input string. That is expensive, particularly for the common case that left and right are used to compute a small prefix/suffix of a much longer string. That check is also overly conservative: for example, we can take the ASCII fast-path for left(s, k) if the first k bytes in the string are ASCII, even if there are multibyte characters elsewhere in the string.

This PR implements two optimizations:

  1. Only call is_ascii on the bytes necessary to determine if we can take the fast-path, not the entire input string, as described above.
  2. Benchmarking identified that for Utf8View inputs, the first optimization regressed some benchmark cases (e.g., negative_n) on both ARM and x86. Claude's theory is that make_view is out-of-line and does an indirect jump on the length of the result string; it seems that after implementing the first optimization, this jump was not well-handled by the branch predictor. Instead, we add a helper sub_view that returns a view that is a substring of an existing view. This can be inlined and avoids the indirect jump incurred by make_view; it can also construct the new view from the old view with bitwise ops, rather than building the new view on the stack.

Benchmarks:

x86 (AMD EPYC Milan)

  • left Utf8 long_result: 101.7 → 93.4 (−8%)
  • left Utf8 n_exceeds_len: 91.7 → 93.8 (+2%)
  • left Utf8 negative_n: 90.4 → 83.2 (−8%)
  • left Utf8 per_row_n: 93.9 → 89.5 (−5%)
  • left Utf8 short_result: 91.0 → 82.6 (−9%)
  • left Utf8 short_result_long_input: 310.4 → 89.4 (−71%)
  • left Utf8View long_result: 53.6 → 40.9 (−24%)
  • left Utf8View n_exceeds_len: 55.9 → 50.7 (−9%)
  • left Utf8View negative_n: 58.8 → 45.2 (−23%)
  • left Utf8View per_row_n: 89.0 → 45.2 (−49%)
  • left Utf8View short_result: 54.2 → 47.5 (−12%)
  • left Utf8View short_result_long_input: 249.1 → 52.7 (−79%)
  • right Utf8 long_result: 109.4 → 91.8 (−16%)
  • right Utf8 n_exceeds_len: 97.9 → 98.5 (+1%)
  • right Utf8 negative_n: 91.5 → 84.7 (−7%)
  • right Utf8 per_row_n: 100.7 → 88.1 (−13%)
  • right Utf8 short_result: 96.8 → 87.2 (−10%)
  • right Utf8 short_result_long_input: 261.2 → 95.5 (−63%)
  • right Utf8View long_result: 64.2 → 53.8 (−16%)
  • right Utf8View n_exceeds_len: 63.9 → 56.0 (−12%)
  • right Utf8View negative_n: 62.3 → 55.8 (−10%)
  • right Utf8View per_row_n: 101.0 → 54.9 (−46%)
  • right Utf8View short_result: 62.9 → 58.7 (−7%)
  • right Utf8View short_result_long_input: 202.2 → 67.7 (−67%)

ARM (Apple M4 Max)

  • left Utf8 long_result: 91.3 → 62.6 (−31%)
  • left Utf8 n_exceeds_len: 69.4 → 71.4 (+3%)
  • left Utf8 negative_n: 66.8 → 69.5 (+4%)
  • left Utf8 per_row_n: 66.6 → 60.6 (−9%)
  • left Utf8 short_result: 69.1 → 62.7 (−9%)
  • left Utf8 short_result_long_input: 91.1 → 59.6 (−35%)
  • left Utf8View long_result: 45.4 → 26.0 (−43%)
  • left Utf8View n_exceeds_len: 37.6 → 34.4 (−9%)
  • left Utf8View negative_n: 37.8 → 25.9 (−31%)
  • left Utf8View per_row_n: 68.1 → 29.4 (−57%)
  • left Utf8View short_result: 37.9 → 25.0 (−34%)
  • left Utf8View short_result_long_input: 61.1 → 28.9 (−53%)
  • right Utf8 long_result: 104.2 → 64.7 (−38%)
  • right Utf8 n_exceeds_len: 80.1 → 79.1 (−1%)
  • right Utf8 negative_n: 71.4 → 66.0 (−8%)
  • right Utf8 per_row_n: 74.0 → 64.3 (−13%)
  • right Utf8 short_result: 80.2 → 65.4 (−18%)
  • right Utf8 short_result_long_input: 101.1 → 63.4 (−37%)
  • right Utf8View long_result: 62.7 → 32.0 (−49%)
  • right Utf8View n_exceeds_len: 42.0 → 38.5 (−8%)
  • right Utf8View negative_n: 41.3 → 28.5 (−31%)
  • right Utf8View per_row_n: 72.0 → 33.5 (−53%)
  • right Utf8View short_result: 43.9 → 30.0 (−32%)
  • right Utf8View short_result_long_input: 63.7 → 33.3 (−48%)

What changes are included in this PR?

See above.

What is the testing strategy for this PR?

Existing tests pass; no functional changes.

Are there any user-facing changes?

No.

@github-actions github-actions Bot added the functions Changes to functions implementation label Oct 4, 2026
@neilconway

Copy link
Copy Markdown
Contributor Author

FYI @andygrove @comphead

Comment on lines +1239 to +1271
pub(crate) fn sub_view(view: u128, source: &[u8], range: Range<usize>) -> u128 {
debug_assert!(range.start <= range.end && range.end <= source.len());

// The substring's first bytes, in the low-order bits. Any bits past the
// end of the substring are masked off by `inline_view`.
let leading_bytes = if source.len() <= MAX_INLINE_LEN {
// `source` is stored in `view` itself, after its 4-byte length.
(view >> 32) >> (8 * range.start)
} else {
// `source` has more than 12 bytes, so read the 12 bytes starting at
// `range.start`, or the last 12 bytes if that would run past the end,
// and skip any that come before `range.start`.
let window_start = range.start.min(source.len() - MAX_INLINE_LEN);
let window = source[window_start..window_start + MAX_INLINE_LEN]
.try_into()
.unwrap();
read_12_bytes(window) >> (8 * (range.start - window_start))
};

let len = range.len();
if len <= MAX_INLINE_LEN {
inline_view(leading_bytes, len)
} else {
let original = ByteView::from(view);
ByteView {
length: len as u32,
prefix: leading_bytes as u32,
offset: original.offset + range.start as u32,
..original
}
.as_u128()
}
}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This could potentially be used in other places (e.g., substr), but that will require more careful evaluation; I'll defer that for now.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The same logic also lives as a private substr_view in string/split_part.rs:509 and unicode/substrindex.rs:583. Neither calls append_view, so the follow-up can retire three definitions.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yep, makes sense!

@codecov-commenter

codecov-commenter commented Oct 4, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 82.67%. Comparing base (c3ef346) to head (1b6048d).
⚠️ Report is 11 commits behind head on main.

Additional details and impacted files
@@           Coverage Diff            @@
##             main   #26039    +/-   ##
========================================
  Coverage   82.66%   82.67%            
========================================
  Files        1147     1147            
  Lines      446357   446589   +232     
  Branches   446357   446589   +232     
========================================
+ Hits       368971   369205   +234     
+ Misses      54997    54981    -16     
- Partials    22389    22403    +14     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

apache#23762 added an ASCII fast path to `left_right_byte_length` that calls
`is_ascii()` on the whole string for every row. For short results from
longer strings, such as Utf8View data read from Parquet, that scan costs
more than the per-character scan it replaced.

ASCII bytes are never part of a multi-byte UTF-8 sequence, so if the `n`
bytes at the relevant end of the string are ASCII, they are exactly the
`n` characters at that end. Check only those bytes, and fall back to the
per-character scan otherwise.
`make_view` selects per-length copy code with a jump on the result
length. When result lengths vary from row to row, as with a negative
`n`, the CPU often mispredicts that jump. Checking only the result
bytes for ASCII removed a length-dependent loop that had been making
the jump predictable, so those cases got slower.

Add `sub_view`, which builds the view for a substring of an existing
view with shifts and masks: from the view itself when the source is
inlined, and otherwise from a 12-byte window of the source that is
always in bounds. Use it for Utf8View input in `left` and `right`.
@neilconway
neilconway force-pushed the neilc/perf-left-right-ascii-prefix branch from b3e5014 to 3a137c2 Compare October 5, 2026 01:57

@jayzhan211 jayzhan211 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @neilconway , overall LGTM

Comment thread datafusion/functions/src/strings.rs Outdated

@comphead comphead left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @neilconway makes a lot of sense for me

@comphead comphead left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking notes:

  • Bench coverage: left_right.rs builds its strings with arrow's alphanumeric generator, so every value is ASCII. The multibyte fallback and the "ASCII end, multibyte elsewhere" shortcut are not measured. Two small cases would cover both: a multibyte char inside the n-char window (fallback) and one only outside it (shortcut).
  • inline_view has a single caller and its debug_assert! repeats the len <= MAX_INLINE_LEN guard right above the call. Folding it into that arm would drop both.

// Byte offset of the `abs`-th character from the end.
Ordering::Less => {
let start = bytes.len().saturating_sub(abs);
if bytes[start..].is_ascii() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When abs >= bytes.len() no byte needs inspecting. The result is 0 here (bytes.len() in the Greater arm) whatever the content, since a string has at most bytes.len() chars. Today that case scans the whole string with is_ascii, and non-ASCII input then also pays a full char_indices pass to reach the same value. That matches n_exceeds_len being flat for Utf8 in the table (-1% to +3%) while short_result_long_input drops 35% to 71%.

An early exit per arm would skip it, for example start == 0 || bytes[start..].is_ascii() here and end == bytes.len() || bytes[..end].is_ascii() below. It adds a length-dependent branch, so short_result and per_row_n, where only some rows have abs >= len, are worth rechecking next to n_exceeds_len. I have not measured any of this.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree that that could be a win, but it merits separate evaluation / benchmarking. I'd rather land this and then take a look at that subsequent optimization in a followup.

Comment thread datafusion/functions/src/unicode/common.rs
@neilconway

Copy link
Copy Markdown
Contributor Author

On the other comments @comphead:

  • I agree that multi-byte benchmarks would be useful, although the same could be said of a lot of the string UDFs :) I'll defer that for now.
  • Personally I think leaving inline_view as a separate function makes the implementation more readable, so I'm inclined to leave that as-is.

@neilconway
neilconway enabled auto-merge October 5, 2026 16:50
@neilconway

Copy link
Copy Markdown
Contributor Author

Thanks for the reviews @jayzhan211 @comphead !

@neilconway
neilconway added this pull request to the merge queue Oct 5, 2026
Merged via the queue into apache:main with commit bd86190 Oct 5, 2026
42 checks passed
@neilconway
neilconway deleted the neilc/perf-left-right-ascii-prefix branch October 5, 2026 19:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

functions Changes to functions implementation v56.0.0

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants