r/rust • • 16h ago

Article from Daniel Lemire: How many strings can you create per second?

https://lemire.me/blog/2026/09/25/how-many-strings-can-you-create-per-second/

This is a fairly new article from Daniel Lemire (the SIMD expert, author of simdjson).

It seems to_string() need to use the same trick.

C++ wins by a wide margin at 5.4 ns per string. The trick is the small string optimization: a std::string stores short strings, directly inside the object. Our strings have at most eight digits, so C++ never calls the memory allocator.

50 Upvotes

20 comments sorted by

63

u/vdrnm 15h ago edited 14h ago

That would be a huge braking change and I doubt it will even be considered.

Small string optimizations do not have universal benefit (branch misses and additional instructions mean that if smallstring ends up being heap-allocated, it will be slower then std String).

When you do need it, there are hundreds of crates for them. Just type "string" in lib.rs or crates.io

8

u/nightcracker 8h ago

The part about it being a breaking change is true. However:

branch misses and additional instructions mean that if smallstring ends up being heap-allocated, it will be slower then std String

There are no branch misses. Selecting the base pointer based on the length can be done entirely branch-free. There are some extra instructions though.

3

u/vdrnm 5h ago

That's a good point, and cpp implementation does not have them.

Looking at the code of smallvec when derefing to a slice I'd expect rustc to replace the branch with cmov, but unfortunately it does not.

1

u/CocktailPerson 1h ago

cmov is often a pessimization, actually. There ain't no such thing as a free lunch, and cmov can have very negative effects on CPU pipelining. You have to work pretty hard to get a LLVM to generate a cmov these days, because so many benchmarks have shown that it's just not worth it.

1

u/scook0 2h ago

Selecting the base pointer based on the length can be done entirely branch-free. There are some extra instructions though.

In a vacuum, it's unclear to me whether selecting a pointer with cmov is actually an improvement over branching. The extra instructions are still an optimization barrier, and I would expect branch prediction to do a pretty good job here.

Also note that in C++ it's possible for dereferencing an inline string to be completely unconditional, by having the string pointer point directly to its inline storage. That approach is impractical in Rust because moving the string would invalidate the pointer. Though I don't know offhand whether commonly-used C++ string implementations actually do this.

36

u/masklinn 15h ago edited 15h ago

It seems to_string() need to use the same trick.

It can not, because String::as_mut_vec exists and Vec guarantees that it doesn't do SSO.

Furthermore SSO is not necessarily beneficial, and Rust is less likely to fall on the beneficial side due to no copy constructor, so less copy, so less gain from a stack-allocated string, so unless you're cloning a lot the branching from SSO can create more costs than it saves (anecdotally every time I tried SSO strings they behaved worse than good ol String for the stuff I was doing).

And finally there's a whole stable of SSO strings available which while not trivially usable as String can commonly be drop-in thanks to the broad usage of &str.

And here since

Our strings have at most eight digits

I'd use a stack allocated string directly. Here's using heapless, on an M1 pro (so an older and slower CPU than Daniel's):

i.to_string()              18.81 ns/string      53.2 M/s
itoa + to_owned()          18.54 ns/string      53.9 M/s
heapless + write!()        10.89 ns/string      91.8 M/s
heapless + itoa()           4.88 ns/string     204.8 M/s

The code is not identical to the original because it's missing the nice to_string/to_owned, and heapless::String::push_str is more try-ish, but it's not exactly monstrous:

let mut s = heapless::String::new();
_ = write!(s, "{}", i);
buf[(i & 1023) as usize] = s;

let mut s = heapless::String::new();
_ = s.push_str(b.format(i));
buf[(i & 1023) as usize] = s;

13

u/CryZe92 14h ago edited 14h ago

You don’t even need heapless, both itoa and even std have buffers specifically for this purpose: https://doc.rust-lang.org/stable/std/primitive.i32.html#method.format_into

2

u/matthieum [he/him] 13h ago

I just realized that unfortunately NumBuffer cannot quite act as a string, since it doesn't remember how many bytes of its buffer were used :'(

24

u/MvKal 15h ago

Kinda a weird thing to measure ngl, basically just benchmarking mem allocation speed. Also you can do small string optimization in rust as well if you do end up needing those nanoseconds for whatever reason.

17

u/matthieum [he/him] 13h ago

Honestly, meh?

This is so artificial a benchmark.

I mean, if you're formatting such short strings, you're most likely holding the tool wrong. For example, I've regularly seen newbies creating long string using catenation:

"Hello, " + name + "! Would you like " + std::to_string(n) + " apples?"

It's easy, but don't. Instead, you'd want to use formatting:

std::format("Hello, {}! Would you like {} apples?", name, n)

And lo and behold, no short-lived small string remains.

13

u/TDplay 13h ago

We create new strings all the time.

I consider this assertion dubious.

If your program creates so many strings that it is a performance issue, you should probably put some thought into how you handle strings, instead of just accepting the programming language's default. Well-implemented string interning will likely outperform any general-purpose string type.

It seems to_string() need to use the same trick.

It does not need to do this, and in fact, it cannot do this. Rust has made a stable guarantee that neither String nor Vec will ever implement this.

Even if it were possible, small-string optimisation can actually be detrimental to performance when you don't have enough small strings to justify it.

If you need the small string optimisation, then find a crate that implements it:

  • smallvec stores a user-specified number of elements inline. Larger vectors spill onto the heap.
  • smol_str stores up to 23-byte strings inline. Larger strings are reference-counted.

12

u/kibwen 12h ago

Having SSO was the right default for C++, because best practice there is to err on the side of caution by defensively copying string buffers, making string copies tremendously more common. Famous case study: "std::string is responsible for almost half of all allocations in the Chrome browser process; please be careful how you use it! In the course of optimizing SyzyASan performance, the Syzygy team discovered that nearly 25000 (!!) allocations are made for every keystroke in the Omnibox." https://www.reddit.com/r/cpp/comments/2od5l0/stdstring_is_responsible_for_almost_half_of_all/

Conversely, having SSO would have been the wrong default for Rust, because 1) the existence of the borrow checker and a proper string view from the beginning mean that string APIs can confidently pass around pointers without needing to resort to defensive copying, which eliminates the vast majority of copies relative to an analogous C++ codebase; 2) SSO isn't an unambiguous good, because being clever with the layout would prevent zero-cost coercion from String to &str, which Rust benefits greatly from (SSO also introduces a branch on access, but that's a less important problem). Furthermore, note that Rust does still guarantee that empty strings don't perform any allocation in the first place.

Of course, sometimes SSO is the right choice for a specific application, in which case you have the power to implement that type yourself (or use a crate).

16

u/atlasgorn 15h ago

Another win on cpp design, right next to vector<bool> being a bitvec because that's more performant of course

13

u/ART1SANNN 14h ago

vector<bool> is hated by more senior c++ devs because it breaks the typical container rules and type expectations. IRL perf also don’t really yield speed improvements either, in fact it
might be worse in some cases

8

u/masklinn 13h ago

Yeah AFAIK vector<bool> access is slower in most cases, having more cache misses on a vector<uint8_t> might save the former if the vector is very large and the accesses are so random the prefetcher can't figure it out but that's million-element or above. And because vector<bool> has to contort itself via proxies and can only fit the std::vector interface, a dedicated bitset is almost certainly significantly faster.

1

u/-Redstoneboi- 6h ago

precisely

1

u/barsoap 41m ago

Premature optimisation is the root of all evil.

9

u/thermiter36 12h ago

Seems like the others don't get the joke, but this gave me a chuckle

3

u/JoshTriplett rust · lang · libs · cargo 10h ago

Would be interested to see this benchmark include the fastest of the small-string-optimization crates in the Rust ecosystem.

1

u/BigHandLittleSlap 6h ago

Slight side topic, something I've noticed in these articles comparing programming languages is that Java and C# don't exist in the minds of many developers.

There's definitely "cliques" of developers where some will simply never consider languages used by other cliques, even when comparing programming languages.