Skip to content

Repository files navigation

fdf - High-Performance POSIX File Finder

CI (main)

fdf is a high-performance POSIX file finder written in Rust with extensive C FFI.

A faster alternative to tools such as fd/find/bfs, with a focus on speed, efficiency, and cross-platform compatibility. Benchmarks show fdf running roughly >=2x faster than comparable tools in most cases, achieved through low-level optimisation, SIMD/SWAR techniques, and direct syscalls where possible.

Note: This project will be renamed before 1.0, currently leaning towards 'frep' as the name. Windows support requires a significant rewrite and is planned for post-1.0.

Quick Installation:

cargo install fdf
cargo install --git https://github.com/alexcu2718/fdf
# cargo add fdf
# I don't recommend using as a library until 1.0, sorry!, the API will have warts and I may refactor the whole god damn thing as I get to windows.
# Primarily a CLI tool until 1.0!!!
## Additionally specify  --no-default-features to remove mimalloc dependency

Project Status

This is a performance-focused project that remains under active development towards a stable 1.0 release. The current name is temporary and will change before that release.

The CLI is already usable, but the internal library API is not yet stable. It is quite irritating to both expose a stable api and CLI (fd doesn't do it and walkdir/ignore are much simpler.)

Additional Note: If you use MacOS on native hardware and know your stuff, please contribute if you desire! I find it impossible to benchmark and improve(with certainty*) because I am virtualising x86_64 MacOS and I get so many baffling results from benchmarks due to the virtualisation overhead.

Platform Support

Fully Supported and CI Tested

(I don't use any inline assembly so other architectures ie Linux/FreeBSD aarch64 should work fine, will add to CI in future!)

  • Linux (x86_64, s390x (Big endian), Alpine( MUSL libc))
  • macOS (Intel and Apple Silicon)
  • FreeBSD (x86_64)
  • NetBSD (x86_64)
  • OpenBSD (x86_64)
  • Solaris/Illumos(x86_64)
  • Android (Termux) (aarch64)

Compiles with Limited Testing

Note: GitHub Actions does not yet provide Rust 2024 support for some of these platforms. Additional checks will be added when available.

  • 32-bit Linux

Not Yet Supported

  • Windows: Requires significant rewrite due to architectural differences with libc. Planned once the POSIX feature set is stable.

  • DragonflyBSD: Blocked on Rust 2024 support.

Probably Broken

  • MacOSx 32bit (granted the last 32bit OSx was released in 2009, I would possibly tick it off if I get very bored)

  • Other Niche Operating systems like Fuschia (granted, they probably won't support Rust 2024).

Testing

The project includes comprehensive testing with 100+ Rust (including doctests) tests and 15+ correctness benchmarks comparing against fd.

Miri is not practical here due to the extensive libc usage, so validation relies on intensive testing and Valgrind. See scripts/valgrind-test.sh.

  • Rust tests: Available here
  • Shell scripts clone the LLVM repository to provide an accurate testing environment
  • Tests run via GitHub Actions on all supported platforms

Running the Full Test Suite:

TMP_DIR="${TMP:-/tmp}"
git clone --depth 1 https://github.com/alexcu2718/fdf "$TMP_DIR/fdf_test"
cd "$TMP_DIR/fdf_test"

./scripts/run_benchmarks.sh

This runs the internal library tests, CLI tests, and benchmarks.

Performance Benchmarks

The benchmarks are repeatable using the testing code above and cover file type filtering, extension matching, file sizes, and several other scenarios. The following results were gathered on Linux against local directories and the LLVM repository and summarised from hyperfine output.

Easily repeatable examples found in './scripts' and './fd_benchmarks'

| Test Case                                                              | fdf Mean        | fd Mean         | Speedup   | Relative        |
| :----------                                                            | :--------:      | :-------:       | :-------: | :--------:      |
| cold-cache `.' '/home/alexc' -HI -d 4`                                 | 538.0 ± 52.4    | 889.9 ± 79.6    | 1.65x     | 1.65 ± 0.22     |
| cold-cache `.' '/tmp/llvm-project' -HI -d 2`                           | 14.3 ± 0.6      | 46.9 ± 0.9      | 3.28x     | 3.28 ± 0.15     |
| cold-cache `-HI --extension 'c' '' '/home/alexc`                       | 20.347 ± 1.666  | 47.247 ± 10.933 | 2.32x     | 2.32 ± 0.57     |
| cold-cache `-HI --extension 'c' '' '/tmp/llvm-project`                 | 27.2 ± 1.0      | 77.5 ± 1.8      | 2.85x     | 2.85 ± 0.12     |
| cold-cache `.' '/tmp/llvm-project' -HI`                                | 30.4 ± 1.2      | 77.3 ± 4.1      | 2.54x     | 2.54 ± 0.17     |
| cold-cache `.' '..' -HI`                                               | 32.1 ± 1.3      | 82.7 ± 2.9      | 2.58x     | 2.57 ± 0.14     |
| cold-cache `-HI --size -1mb '' '/home/alexc`                           | 31.754 ± 8.892  | 39.331 ± 14.878 | 1.24x     | 1.24 ± 0.58     |
| cold-cache `.' '/tmp/llvm-project' -HI --type d`                       | 27.9 ± 0.8      | 75.3 ± 1.5      | 2.70x     | 2.70 ± 0.09     |
| cold-cache `.' '/home/alexc' -HI --type e`                             | 21.994 ± 1.973  | 32.881 ± 12.948 | 1.49x     | 1.50 ± 0.60     |
| cold-cache `.' '/tmp/llvm-project' -HI --type e`                       | 53.1 ± 4.3      | 121.2 ± 3.1     | 2.28x     | 2.28 ± 0.19     |
| cold-cache `.' '/home/alexc' -HI --type x`                             | 21.332 ± 3.758  | 38.071 ± 24.323 | 1.78x     | 1.78 ± 1.18     |
| cold-cache `.' '/tmp/llvm-project' -HI --type x`                       | 40.6 ± 2.6      | 100.5 ± 2.2     | 2.48x     | 2.47 ± 0.17     |
| warm-cache `.' '/home/alexc' -HI -d 4`                                 | 36.9 ± 1.4      | 249.1 ± 10.3    | 6.75x     | 6.75 ± 0.38     |
| warm-cache `.' '/tmp/llvm-project' -HI -d 2`                           | 2.7 ± 0.4       | 5.5 ± 0.3       | 2.04x     | 2.06 ± 0.31     |
| warm-cache `-HI --extension 'c' '' '/home/alexc`                       | 466.3 ± 3.7     | 848.3 ± 4.3     | 1.82x     | 1.82 ± 0.02     |
| warm-cache `-HI --extension 'c' '' '/tmp/llvm-project`                 | 15.5 ± 0.5      | 30.2 ± 0.7      | 1.95x     | 1.95 ± 0.07     |
| warm-cache `.' '/home/alexc' -HI`                                      | 524.0 ± 2.9     | 954.0 ± 5.9     | 1.82x     | 1.82 ± 0.02     |
| warm-cache `.' '/tmp/llvm-project' -HI`                                | 17.6 ± 0.6      | 33.3 ± 0.8      | 1.89x     | 1.89 ± 0.08     |
| warm-cache `.' '..' -HI`                                               | 19.7 ± 0.6      | 35.2 ± 1.1      | 1.79x     | 1.79 ± 0.08     |
| warm-cache `-HI '.*[0-9].*(md\|\.c)$' '/home/alexc`                    | 492.8 ± 3.5     | 797.7 ± 3.4     | 1.62x     | 1.62 ± 0.01     |
| warm-cache `-HI '.*[0-9].*(md\|\.c)$' '/tmp/llvm-project`              | 17.0 ± 0.8      | 27.8 ± 0.8      | 1.64x     | 1.64 ± 0.09     |
| warm-cache `-HI --size +1mb '' '/home/alexc`                           | 755.6 ± 3.1     | 1781.2 ± 5.4    | 2.36x     | 2.36 ± 0.01     |
| warm-cache `-HI --size '+1mb' '' '/tmp/llvm-project`                   | 32.0 ± 1.7      | 70.5 ± 0.9      | 2.20x     | 2.21 ± 0.12     |
| warm-cache `-HI --size -1mb '' '/home/alexc`                           | 814.8 ± 5.0     | 1980.6 ± 19.4   | 2.43x     | 2.43 ± 0.03     |
| warm-cache `.' '/home/alexc' -HI --type d`                             | 489.6 ± 4.5     | 891.6 ± 18.5    | 1.82x     | 1.82 ± 0.04     |
| warm-cache `.' '/tmp/llvm-project' -HI --type d`                       | 15.1 ± 0.4      | 29.5 ± 1.1      | 1.95x     | 1.96 ± 0.09     |
| warm-cache `.' '/home/alexc' -HI --type e`                             | 1.091 ± 0.021   | 1.821 ± 0.008   | 1.67x     | 1.67 ± 0.03     |
| warm-cache `.' '/tmp/llvm-project' -HI --type e`                       | 35.1 ± 1.7      | 58.9 ± 1.1      | 1.68x     | 1.68 ± 0.09     |
| warm-cache `.' '/home/alexc' -HI --type x`                             | 699.9 ± 2.4     | 1233.2 ± 3.4    | 1.76x     | 1.76 ± 0.01     |
| warm-cache `.' '/tmp/llvm-project' -HI --type x`                       | 28.0 ± 1.7      | 47.3 ± 1.9      | 1.69x     | 1.69 ± 0.12     |

--Average Speedup: 2.20x--

Distinctions from fd/find

Symlink resolution in my method differs from fd and find. Although I generally advise against following symlinks, the option exists for completeness.

When following symlinks, behaviour will vary slightly. For example, fd can enter infinite loops with recursive symlinks (see recursive_symlink_fs_test.sh) Available here whereas my implementation prevents hangs.

To avoid issues, use --same-file-system when traversing symlinks. This ensures traversal terminates safely even in complex directories such as ~/.steam, ~/.wine, /sys, and /proc.

Extension matching is not using a regex, if you wish to search for regex ends, just use EG: '.tmp$' in your search regex

I do not deduplicate symlinks BECAUSE this would require basically calling a lot of unnecessary stat calls.

This is demonstrated here,

alt text

Technical Highlights

Key Optimisations

  • getdents64/getdents: Optimised the Linux/Android-specific/OpenBSD/NetBSD/Illumos/Solaris directory reading by significantly reducing the number of stat/statx/fstatat system calls

  • Reverse engineered MacOS syscalls(__getdirentries64) to exploit early EOF and no unnecessary stat/pthread_mutex calls at link here and this link for syscall implementation (Also works on FreeBSD)

  • A custom written crossbeam workstealing parallel traversal algorithm

Constant-Time Directory Strlen

The following function provides an elegant solution to avoid strlen on d_name but achieving it so with a rather evil trick, it's not all like this! I just found it cool. Check source code for further explanation in utils.rs, I have removed most of the explanatory comments**

This version is simplified from the actual implementation

// Computational complexity: O(1) - truly constant time (Well, cache effects are mostly negligible)
// SIMD within a register, so no architecture dependence
//http://www.icodeguru.com/Embedded/Hacker%27s-Delight/043.htm
#[cfg(any(target_os = "linux", target_os = "android", has_d_namlen))]
#[inline]
pub const unsafe fn dirent_const_time_strlen(drnt: *const dirent64) -> usize {
    #[cfg(has_d_namlen)] //Generated by cc build script.
    // SAFETY: `dirent` must be validated ( it was required to not give an invalid pointer)
    return unsafe { (*drnt).d_namlen as usize }; //trivial operation for systems with d_namlen field
    #[cfg(not(has_d_namlen))]
    // On these systems where we need a bit of 'black magic' (no d_namlen field)
    {
        use core::{mem::offset_of, num::NonZeroU64};
        // Offset from the start of the struct to the beginning of d_name.
        const DIRENT_HEADER_START: usize = offset_of!(dirent64, d_name);
        // Access the last field and then round up to find the minimum struct size
        const MIN_DIRENT_SIZE: usize = DIRENT_HEADER_START.next_multiple_of(8);
        const { assert!(MIN_DIRENT_SIZE == 24, "dirent min size must be 24!") };
        const LO_U64: u64 = 0x0101_0101_0101_0101;
        const HI_U64: u64 = 0x8080_8080_8080_8080;

        /*  SAFETY: `dirent` is valid by precondition */
        let reclen = unsafe { (*drnt).d_reclen } as usize;
        /*
          Read the last 8 bytes of the struct as a u64.
        This works because dirents are always 8-byte aligned. (it is guaranteed aligned by the kernel) */
        let mut last_word: u64 = unsafe { drnt.byte_add(reclen - 8).cast::<u64>().read() };

        #[cfg(target_endian = "little")]// MIN dirent size is 24, it is alwats a multiple of 8, 24 wraps, >=32 Doesn't
        let mask = (reclen as u64).wrapping_sub(25) >> 40;// Check implementation for explanation, HORRIBLE!
        // BE the bits in the correct position however we only want the first 3 bytes.
        #[cfg(target_endian = "big")]
        let mask = (reclen as u64).wrapping_sub(25) & 0xFFFF_FF00_0000_0000;

        debug_assert!( // handy debug test, explains what the above means!
            reclen == 24 && mask == u64::from_ne_bytes([0xFF, 0xFF, 0xFF, 0, 0, 0, 0, 0])
                || mask == 0 && reclen != 24,
            "Checking condition holds"
        );

        //Apply the mask to ignore non-name bytes while preserving name bytes.
        last_word |= mask;
        //SAFETY: The u64 can never be all 0's post-mask because the last word ALWAYS contains at least one NUL, which become 0x80
        #[cfg(target_endian = "little")]
        let masked_word = unsafe {
            NonZeroU64::new_unchecked(last_word.wrapping_sub(LO_U64) & !last_word & HI_U64)
        };
        #[cfg(target_endian = "big")]// //http://0x80.pl/notesen/2016-11-28-simd-strfind.html#algorithm-1-generic-simd
        //SAFETY: as in LE version.
        let masked_word = unsafe {
            NonZeroU64::new_unchecked(
                (!last_word & !HI_U64).wrapping_add(LO_U64) & (!last_word & HI_U64),
            )
        };

        // Find the position of the null terminator
        #[cfg(target_endian = "little")]
        let byte_pos = (masked_word.trailing_zeros() >> 3) as usize;
        #[cfg(target_endian = "big")]
        let byte_pos = (masked_word.leading_zeros() >> 3) as usize;

        reclen - DIRENT_HEADER_START + byte_pos - 8
    }
}

Why?

I started this project because I honestly got sick of windows being so slow, so naturally when transitioning to linux (funny story), I wanted something that found the stuff I needed instantly, find was slow, so I started to make it in rust, do it better. Then partway through I discover that fd existed, oh well, actually no, I got so offended by how slow searching is for random stuff that I just kept going. It taught me a lot, that's the main reason I continued, I only really look at this when I have a cool idea or I'm quite bored. So development is not guaranteed especially with a hectic life :@.

Performance Motivation

Rust's std::fs has inefficiencies for this workload, primarily by choice of excessive heap allocation/syscalls which I understand makes it much easier to write cross platform code, for better or worse (worse..) it is wiser to build a lower abstraction framework to achieve this. It's not very fun. Cross compatibility is awful at this level hence why I am procastinating windows (I should really write a rationale doc...)

Development Philosophy

Feature stability before breakage - I won't push breaking changes or advertise this anywhere until I've got a good baseline.

Open to contributions/Requests(mostly).

In short, this project explores performance, low-level programming, and practical tooling.

Acknowledgements/Disclaimers

I've directly taken code from fnmatch-regex, found at the link and modified it so I could convert globs to regex patterns trivially, this simplifies the string filtering model by delegating it to rust's extremely fast regex crate. Notably I modified it because it's quite old and has dependencies I was able to remove

(I have emailed and received approval from the author above)

I've also done so for some SWAR tricks from the standard library (see link) which is implemented at the following link I additionally emailed the author of memchr and got some nice tips, great guy!.

Future Plans

Feature Enhancements (Planned)

API cleanup, currently the CLI is the main focus but I'd like to fix that eventually!

POSIX Compliance: Mostly done, I don't expect to extend this beyond Linux/BSD/MacOS/Illumos/Solaris/Android (the other ones are embedded mostly, correct me if i'm wrong!), I have tentative work for other OS'es, it may support NuttX/few others but completely untested.

Platform Expansion

Windows Support: Post 1.0.

Installation and Usage

# Clone & build
git clone https://github.com/alexcu2718/fdf.git
cd fdf
cargo build --release

# Optional system install
cargo install --git https://github.com/alexcu2718/fdf


# Find all JPG files in the home directory (excluding hidden files)
fdf . ~ -e jpg

# Find all  Python files in /usr/local (including hidden files)
fdf . /usr/local -e py -H

# Null terminated all output instead of newlines, mainly for command passing to other functions
fdf -HI --print 0 . ~ | xargs -0 realpath




# Generate shell completions for Zsh/bash (also supports powershell/fish!)
# For Zsh
echo 'eval "$(fdf --generate zsh)"' >> ~/.zshrc

# For Bash
echo 'eval "$(fdf --generate bash)"' >> ~/.bashrc
Usage: fdf [OPTIONS] [PATTERN] [PATH]

Arguments:
  [PATTERN]
          Pattern to search for

  [PATH]
          Path to search (defaults to current working directory)

Options:
  -H, --hidden
          Shows hidden files eg .gitignore or .bashrc, defaults to off

  -S, --sort
          Sort the entries alphabetically (this has quite the performance cost)

  -s, --case-sensitive
          Enable case-sensitive matching, defaults to false

  -e, --extension <EXTENSION>
          An example command would be `fdf -HI -e  c '^str' /

  -j, --threads <THREAD_NUM>
          Number of threads to use, defaults to available threads available on your computer (or 4 on MacOS)

  -a, --absolute-path
          Starts with the directory entered being resolved to full

  -L, --follow
          Include symlinks in traversal,defaults to false

      --nocolour
          Disable colouring output when sending to terminal

  -g, --glob
          Use a glob pattern,defaults to off

  -n, --max-results <TOP_N>
          Retrieves the first eg 10 results, 'fdf  -n 10 '.cache' /

  -d, --depth <DEPTH>
          Retrieves only traverse to x depth

  -p, --full-path
          Use a full path for regex matching, default to false

  -F, --fixed-strings
          Use a fixed string not a regex, defaults to false

      --show-errors
          Show errors when traversing

      --same-file-system
          Only traverse the same filesystem as the starting directory

  -0, --print0
          Makes all output null terminated as opposed to newline terminated only applies to non-coloured output and redirected(useful for xargs)

  -I, --no-ignore
          Do not respect .gitignore rules during traversal

      --strip-cwd-prefix
          Strip the leading './' from results when searching the current directory

  -Q, --quoted
          Wrap printed file paths in double quotes

      --exec <CMD>...
          Execute a command once per search result.
          Use '{}' to insert the matched path into an argument; if '{}' is omitted, the path is appended as the final argument. This option should be the final CLI flag.
          Example: 'fdf 'junk.files' 'test_directory' -HI --exec rm -rf ' , delete all files meeting the criteria

      --ignore <PATTERN>
          Ignore paths that match this regex pattern (repeatable)

      --ignoreg <GLOB>
          Ignore paths that match this glob pattern (repeatable)

      --ignore-file <path>
          Add a custom ignore-file in '.gitignore' format. These files have a low precedence.

      --and <pattern>
          Add additional required search patterns, all of which must be matched.
          Multiple additional patterns can be specified. The patterns are regular expressions, unless '--glob' or '--fixed-strings' is used.

      --size <SIZE>
          Filter by file size

          PREFIXES:
            +SIZE    Find files larger than SIZE
            -SIZE    Find files smaller than SIZE
             SIZE     Find files exactly SIZE (default)

          UNITS:
            b        Bytes (default if no unit specified)
            k, kb    Kilobytes (1000 bytes)
            ki, kib  Kibibytes (1024 bytes)
            m, mb    Megabytes (1000^2 bytes)
            mi, mib  Mebibytes (1024^2 bytes)
            g, gb    Gigabytes (1000^3 bytes)
            gi, gib  Gibibytes (1024^3 bytes)
            t, tb    Terabytes (1000^4 bytes)
            ti, tib  Tebibytes (1024^4 bytes)

          EXAMPLES:
            --size 100         Files exactly 100 bytes
            --size +1k         Files larger than 1000 bytes
            --size -10mb       Files smaller than 10 megabytes
            --size +1gi        Files larger than 1 gibibyte
            --size 500ki       Files exactly 500 kibibytes

          Possible values:
          - 100:   exactly 100 bytes
          - 1k:    exactly 1 kilobyte (1000 bytes)
          - 1ki:   exactly 1 kibibyte (1024 bytes)
          - 10mb:  exactly 10 megabytes
          - 1gb:   exactly 1 gigabyte
          - +1m:   larger than 1MB
          - +10mb: larger than 10MB
          - +1gib: larger than 1GiB
          - -500k: smaller than 500KB
          - -10mb: smaller than 10MB
          - -1gib: smaller than 1GiB

  -T, --time-modified <TIME>
          Filter by file modification time

          PREFIXES:
            -TIME    Find files modified within the last TIME (newer)
            +TIME    Find files modified more than TIME ago (older)
             TIME    Same as -TIME (default)

          TIME RANGE:
            TIME..TIME   Find files modified between two times

          UNITS:
            s, sec, second, seconds     - Seconds
            m, min, minute, minutes     - Minutes
            h, hour, hours              - Hours
            d, day, days                - Days
            w, week, weeks              - Weeks
            y, year, years              - Years

          EXAMPLES:
            --time -1h        Files modified within the last hour
            --time +2d        Files modified more than 2 days ago
            --time 1d..2h     Files modified between 1 day and 2 hours ago
            --time -30m       Files modified within the last 30 minutes

          Possible values:
          - -1h:    modified within the last hour
          - -30m:   modified within the last 30 minutes
          - -1d:    modified within the last day
          - +2d:    modified more than 2 days ago
          - +1w:    modified more than 1 week ago
          - 1d..2h: modified between 1 day and 2 hours ago

  -t, --type <TYPE_OF>
          Filter by file type

          Possible values:
          - d: Directory
          - u: Unknown type
          - l: Symbolic link
          - f: Regular file
          - p: Pipe/FIFO
          - c: Character device
          - b: Block device
          - s: Socket
          - e: Empty file
          - x: Executable file

      --generate <GENERATE>

              Generate shell completions for bash/zsh/fish/powershell
              To use: eval "$(fdf --generate SHELL)"
              Example:
              # Add to shell config for permanent use
              echo 'eval "$(fdf --generate zsh)"' >> ~/.zshrc && source ~/.zshrc

          [possible values: bash, elvish, fish, powershell, zsh]

  -h, --help
          Print help (see a summary with '-h')

  -V, --version
          Print version

Potential Future Enhancements

1. io_uring System Call Batching

  • Investigate batching of stat and similar operations.
  • Key challenges:
    • No native getdents support in io_uring.
    • Would require async runtime integration (e.g. Tokio).
    • Conflicts with the project’s minimal-dependency design.
    • Linux-only feature, making it a low-priority and high-effort addition. I will likely NOT do this

2. Allocation-Optimised Iterator Adaptor

  • Implement a filtering mechanism that avoids unnecessary directory allocations See link for reference

3. Additional Ideas

  • Implement features such as ownership tracking.
  • Maybe use NO_ATIME to avoid disk writes, this has a lot of drawbacks however.
  • Write my sorting algorithm with parallelism, currently it's single threaded, not a big issue for small results but if you want to sort your entire filesystem then you're much better off just piping the results to sort (hence why it's not been optimised much)
  • Do some testing of statx vs fstat--- NO BAD IDEA, it's slower than fstat, only marginally and makes the code a lot worse to write for no real advantage, well except if you want 64bit st_ntime

4. macOS optimisations

  • Unfortunately, due to lacking real Apple hardware (and incentive, I'll be blunt...), testing/benchmarking is done via QEMU (on x64 - you can't emulate ARM, unfortunately).
  • This makes it extremely hard to benchmark (surprisingly, benchmarks hold up well on emulated FreeBSD, etc.).
  • Benchmarking on a VM is FULL of inconsistent, silly things that makes a lot of results basically useless, I need to test on real hardware
  • I use OSX-KVM to emulate it if you're interested.
  • There's some interesting stuff in pfind, but I am suspect of some of it's claims.
  • I need to properly search for all the relevant links/references - the interesting part is how, somehow, lower thread counts on Apple hardware lead to better performance. It's really bizarre.
  • There is ****-all documentation for APFS also, which makes everything folklore based on old ass forum posts... I AM LIVING IN HELL.

About

Fd Faster

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages