Skip to content

core: add memchr-based single-ascii byte search to str_bytes - #161760

Open
pacak wants to merge 11 commits into
rust-lang:mainfrom
pacak:push-ptqpwloxpyru
Open

core: add memchr-based single-ascii byte search to str_bytes#161760
pacak wants to merge 11 commits into
rust-lang:mainfrom
pacak:push-ptqpwloxpyru

Conversation

@pacak

@pacak pacak commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

pacak added 7 commits August 25, 2026 08:13
Right now things are undertested and underspecified.

Some of the library code would get in a loop if searcher starts
returning empty rejects.

And there's no tests for backwards multi byte char matchers. Pull
request I'm reviving had a problem implementing that, so making sure
it's tested before the actual code lands.

Right now it is possible to break both tests (and user code) without
breaking anything else in the test suite I think.
Add a Haystack trait describing something that can be searched in and
make core::str::Pattern (and related types) generic on that trait.
This will allow Pattern to be used for types other than str (most
notably OsStr).

This somewhat follows the Pattern API 2.0 design.  While that design is
apparently abandoned (?), it is somewhat helpful when going for patterns
on OsStr, so I’m going with it unless someone tells me otherwise. ;)

For now leave Pattern, Haystack et al in core::str::pattern.  Since
they are no longer str-specific, I’ll move them to core::pattern in
future commit.  This one leaves them in place to make the diff
smaller.

@pacak: I moved some (or all new) of the `P: Pattern<&'a str>
constraints into where clause to keep things narrower:

```
pub fn foo<'a, P: Pattern<&'a str>>(&'a self, pat: P, ...) ...
```

to

```
pub fn replacen<'a, P>(&'a self, pat: P, ...) ...
     where
         P: Pattern<&'a str>,
```

Original code had indices in Haystack abstracted as an associated type
Cursor. Replaced with usize - Cursor adds noise with not much value.

Changed wording in 2-3 places - for example Searcher is generic over a
few types so it makes more sense to talk about split points in general
with utf8 split points as an example for `&str`.
Pattern is no longer str-specific, so move it from core::str::pattern
module to a new core::pattern module.  This introduces no changes in
behaviour or implementation.  Just moves stuff around and adjusts
documentation.
Introduce core::pattern::Split and core::pattern::SplitN internal types
which can be used to implement iterators splitting haystack into parts.
Convert str’s Split-family of iterators to use them.  In the future,
more haystacks will use those internal types.

Co-authored-by: Peter Jaszkowiak <p.jaszkow@gmail.com>

@pacak: Fixed some typos, added a few `#[inline]`. Since there's no
`H::Cursor` - I had to add `ctx: PhantomData<H>`.
This reverts commit 85cf233ced0d0fe02734c8a83b6d79ccc5432d06.

Gone for now, I'll reimplement it later in str_bytes.rs, will confirm
with the benchmarks included that the optimization still applies
Introduce core::pattern::EmptyNeedleSearcher internal type which
implements logic for matching an empty pattern against a haystack.
Convert core::str::pattern::StrSearcher to use it.  In future more
implementations will take advantage of it.

Also adapt and rework TwoWayStrategy into an internal SearchResult
trait  which abstracts differences between Searcher’s next, next_match
and next_rejects methods.  It makes it simpler to write a single generic
method implementing optimised versions of all those calls.


@pacak:
- Fixed a few typos.
- There's no H::Cursor parameter so code gets a bit simplified.
- Added a test to assert how TwoWaySearcher runs with
  EmptyNeedleSearcher
@pacak:
- made more things const fn
- there was a (copy-paste?) error in try_finish_byte_sequence so I
  added a test that checks try_next_code_point(_reverse) with some
  values, including invalid ones.
- reworded a few comments (passive voice, etc)

Also different comments: since former is public and later is private due
to historical reasons.

> This is different than [`next_code_point`] in that it doesn't assume

> This is different than `next_code_point_reverse` in that it doesn't assume
@rust-log-analyzer

This comment has been minimized.

@pacak
pacak force-pushed the push-ptqpwloxpyru branch from a348ca6 to 435783e Compare August 25, 2026 14:58
@rust-log-analyzer

This comment has been minimized.

pacak added 2 commits August 25, 2026 13:18
Introduce a new core::str_bytes module with types and functions which
handle string-like bytes slices.  String-like means that they code
treats UTF-8 byte sequences as characters within such slices but
doesn't assume that the slices are well-formed.

A `str` is trivially a bytes sequence that the module can handle but
so is OsStr (which is WTF-8 on Windows and unstructured bytes on
Unix).

Move bunch of code (most notably implementation of the two-way
string-matching algorithm) from core::str to core::str_bytes.

Note that this likely introduces regression in some of the str
function performance (since the new code cannot assume well-formed
UTF-8).  This is going to be rectified by following commit which will
make it again possible for the code to assume bytes format.  This is
not done in this commit to keep it smaller.

@pacak:
- Added a few comments
- tried to hide internal types from the diagnostic

And then there's two different bugs where it would report matched areas
as rejected. This broke str::trim_end_matches and who knows what else.
Caught it thanks to tests in the previous commit. And one underflow bug
on invalid input.
It works right now, but original implementation of the next
commit breaks them with none of existing tests catching this
regression.
pacak added 2 commits August 25, 2026 13:22
Since core::str_bytes module cannot assume byte slices it deals with
are well-formed UTF-8 (or even WTF-8), the code must be defensive and
accept invalid sequences.  This eliminates optimisations which would
be otherwise possible.

Introduce a `Flavour` trait which tags `Bytes` type with information
about the byte sequence.  For example, if a `Bytes` object is created
from `&str` it’s tagged with `Utf8` flavour which gives the code
freedom to assume data is well-formed UTF-8.

This brings back all the optimisations removed in previous commit.

@pacak:
- removed IS_WTF8 associated constant - unused
- fixed a bug related to multibyte reverse matching:
  `next_code_point_reverse` reads the input via Iterator::next_back,
  passing `bytes.iter().rev()` reverses it a second time. Not good.
I reverted `ByteNeedle` change earlier, time to add the same
functionality back.

`ByteSearcherState` is mostly copied from `CharSearcherState`, does
a single ascii byte search.

It is possible to do the dispatch inside of a CharSearcherState, but
that makes it a bit slower.
@pacak
pacak force-pushed the push-ptqpwloxpyru branch from 435783e to 59c8bca Compare August 25, 2026 18:39
@pacak

pacak commented Aug 25, 2026

Copy link
Copy Markdown
Contributor Author

r? @nia-e

@pacak
pacak marked this pull request as ready for review August 25, 2026 22:30
@rustbot

rustbot commented Aug 25, 2026

Copy link
Copy Markdown
Collaborator

stdarch is developed in its own repository. If possible, consider making this change to rust-lang/stdarch instead.

cc @Amanieu, @folkertdev, @sayantn

@rustbot rustbot added S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. and removed S-waiting-on-author Status: This is awaiting some action (such as code changes or more information) from the author. labels Aug 25, 2026
@rustbot

rustbot commented Aug 25, 2026

Copy link
Copy Markdown
Collaborator

⚠️ Warning ⚠️

  • The following commits have merge commits (commits with multiple parents) in your changes. We have a no merge policy so these commits will need to be removed for this pull request to be merged.

    You can start a rebase with the following commands:

    $ # rebase
    $ git pull --rebase https://github.com/rust-lang/rust.git main
    $ git push --force-with-lease
    

@rustbot rustbot added the has-merge-commits PR has merge commits, merge with caution. label Aug 25, 2026
@folkertdev

Copy link
Copy Markdown
Contributor

hey so we get pings for all of these, there are merge commits, and in general a stack this large is going to need endless rebases. Maybe just sit on these for a while?

@pacak

pacak commented Aug 25, 2026

Copy link
Copy Markdown
Contributor Author

Sorry for the pings, I was asked to split it into smaller pull requests and since I don't have write access - it looks like this. I'm aware of the conflict, it was there for me as a reminder that some stuff can be merged independently. I can remove it the next rebase. I don't mind doing rebases either. Original implementation of this pull request has been sat on until it was abandoned a few years ago so I'd rather get it merged.

@folkertdev

Copy link
Copy Markdown
Contributor

totally. it would help if you keep PRs as drafts until they are actually ready to be reviewed. Realistically that can't be the case for more than the top 2 or 3 PRs. That way pings are not actually sent before they are relevant. As it stands we'll now be pinged on all of these when

  • there are merge conflicts, which will happen as the top of the stack merges
  • rebases to a new main commit, which need to happen to resolve the conflicts

Which is all noise

@pacak

pacak commented Aug 25, 2026

Copy link
Copy Markdown
Contributor Author

I can change some of them back to drafts if it helps. Top is not going to be merged, only reviewed. To get to the situation when things are not worse than they currently are you need to merge PRs 3..13. Well, 2nd one can be merged (that's where conflict comes from).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

has-merge-commits PR has merge commits, merge with caution. S-waiting-on-review Status: Awaiting review from the assignee but also interested parties. T-compiler Relevant to the compiler team, which will review and decide on the PR/issue. T-libs Relevant to the library team, which will review and decide on the PR/issue.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants