r/rust • • 8h ago

Embedded std?

34 Upvotes

I was wondering if there were any projects like micro python to port std in to make a barebones os designed for microcontrollers? I understand that the std relies on OS level syscalls, so can we do something like implement a basic fat32 file system to accommodate certain write or read requests. In general this would be used to run std in projects without a OS? Does this already exist, and how hard would it be to make a basic version?


r/rust • • 8h ago

"Tick Tock: Maintaining Time" | RustConf 2026

Thumbnail youtu.be
12 Upvotes

r/rust • • 16h ago

Compiling the Linux kernel with gccrs | RustConf 2026

Thumbnail youtu.be
47 Upvotes

r/rust • • 1d ago

📡 official blog Generic Const Args and You | Inside Rust Blog

Thumbnail blog.rust-lang.org
222 Upvotes

r/rust • • 20h ago

🛠️ project ArcColdString: A 1-word (8-byte) atomically reference-counted SSO string that saves up to 32 bytes over Arc<str>

Thumbnail github.com
61 Upvotes

Disclaimer: Re-post approved by u/matthieum, as previous post was erroneously removed.

I’ve been working on a specialized string type called ColdString. The goal is to create the most memory-efficient string representation possible.

  • Size: Exactly 1 usize (8 bytes on 64-bit).
  • Inline Capacity: Up to 8 bytes (Small String Optimization).
  • Niche Optimization: Option<ColdString> has no memory overhead
  • Heap Overhead: Only 1–9 bytes (VarInt length header) instead of the standard 16-byte (pointer, length) pair.

(Since my last post, the 8th inlineable byte and null-niche optimization were suggestions from the community!)

ArcColdString

I'm presenting ArcColdString, a reference counted ColdString with 8 bytes overhead (inspired by arcstr).

  • Same size and inlining rules as ColdString. Inlined strings are copied instead of reference counted.
  • Heap Overhead: 8 bytes for the AtomicUsize reference count (in addition to the VarInt header).
  • Smaller Reference Counts: ArcColdString32, ArcColdString16, and ArcColdString8 use AtomicU32, AtomicU16, and AtomicU8 counts respectively.

Usage

Available in https://crates.io/crates/cold-string/0.4.0

use cold_string::ArcColdString;

let first = ArcColdString::new("a string longer than one machine word");
let second = first.clone();
assert_eq!(first, second);

Memory Comparisons

Theoretical overhead on a 64-bit target, excluding the UTF-8 payload:

Type 8 bytes 128 bytes 512 bytes
Arc<str> 32 32 32
arcstr::ArcStr 24 24 24
ArcColdString 0 (inline) 17 18
ArcColdString32 0 (inline) 13 14

RSS bytes per unique string in a pre-sized Vec:

Type 8 bytes 128 bytes 512 bytes
Arc<str> 47.1 175.4 560.0
arcstr::ArcStr 39.1 167.5 552.2
string_cache::DefaultAtom 71.5 199.5 586.0
ArcColdString 8.0 167.5 552.1
ArcColdString32 8.0 151.4 536.8

(If I'm not mistaken, string_cache is not a 1:1 comparison since it has to hash string contents, which is 40 bytes overhead?).

Reference Counting Performance

ArcColdString's primary goal is memory and portability. Second to that, my goal is to have ArcColdString's referencing counting performance comparable to Arc<str> and arcstr::ArcStr. While drop and clone speed is comparable now, there's still room for improvement.

Below are single threaded clone and drop measurements, done using criterion on AMD Ryzen 9 5900X 12-Core Processor (3.70 GHz):

Clone (ns) 16 bytes 128 bytes 512 bytes
Arc<str> 3.86 3.82 3.82
arcstr::ArcStr 3.39 3.41 3.43
ArcColdString 3.44 3.45 3.43
Drop (ns) 16 bytes 128 bytes 512 bytes
Arc<str> 2.43 2.49 2.43
arcstr::ArcStr 2.52 2.53 2.52
ArcColdString 2.17 2.18 2.19

Below benchmarks measured 4 threads cloning and dropping strings in a shared size pool of 1024, 16, and 1 string(s), using criterion on Intel(R) Core(TM) i7-9750H CPU @ 2.60GHz (2.59 GHz). The more strings in the pool, the less contention:

4 Threads Clone + Drop (ns) 1024 Strings 16 Strings 1 String
Arc<str> 26.55 54.69 144.81
arcstr::ArcStr 29.29 119.39 214.31
cold_string::ArcColdString32 25.69 90.53 262.83

You can read more benchmarks and implementation details in https://github.com/tomtomwombat/cold-string


r/rust • • 21h ago

Article from Daniel Lemire: How many strings can you create per second?

Thumbnail lemire.me
59 Upvotes

This is a fairly new article from Daniel Lemire (the SIMD expert, author of simdjson).

It seems to_string() need to use the same trick.

C++ wins by a wide margin at 5.4 ns per string. The trick is the small string optimization: a std::string stores short strings, directly inside the object. Our strings have at most eight digits, so C++ never calls the memory allocator.


r/rust • • 1d ago

📡 official blog Demoting i686 Windows targets to std-only

Thumbnail blog.rust-lang.org
130 Upvotes

r/rust • • 1d ago

🎙️ discussion Tyler Mandry: "Beyond the &: A Future for Native Smart Pointers in Rust" | RustConf 2026

Thumbnail youtu.be
51 Upvotes

r/rust • • 22h ago

🎙️ discussion API design question: should a 407 from a proxy check be Ok or Err?

9 Upvotes

I'm working on a Rust SDK that has a small check() method for testing a proxy connection, and I'm not sure I'm using Result in the least surprising way here.

If the proxy replies with 407 Proxy Authentication Required, I currently return Ok(Check), not Err.

My reasoning is basically: the check itself worked) I connected to the proxy and got a real response back. For this method, the status/reason/headers are useful information.

Err is only for cases where I couldn't get a usable response at all - timeout, DNS failure, connection failure, malformed response, etc.

Roughly:

match proxy.check(connect)? {
    check if check.status == 200 => {
        // accepted
    }
    check if check.status == 407 => {
        // proxy replied, auth rejected
    }
    check => {
        // some other resp
    }
}

The part that feels a bit weird is this:

let check = proxy.check(connect)?;

That kind of reads like "the proxy is fine", while what it really means is only "the proxy gave a usable response".

Would you expect 407 to be an Err here, or does treating a protocol-level refusal as a successful diagnostic result make sense? 🙊

Edit: Thanks a lot for all the replies. I think the main question is less about 407 itself and more about what check() is supposed to mean: “did I get a valid proxy response” vs “is this proxy ready to use”

The naming and ? ergonomics points were especially useful. I’m probably going to make that distinction more explicit in the API instead of making callers infer it from the status code


r/rust • • 1d ago

📡 official blog Rust 1.99.0 is out

Thumbnail blog.rust-lang.org
703 Upvotes

r/rust • • 1d ago

Robert Seacord: "Unsafe Rust" | RustConf 2026

Thumbnail youtube.com
63 Upvotes

r/rust • • 1d ago

🎙️ discussion Rust in the kernel? What about Rust without the kernel!

Thumbnail kerkour.com
167 Upvotes

Bare-metal Rust may have a brighter future than micro-kernels.


r/rust • • 1d ago

🛠️ project Rust Glancer 0.3: new trait solver and other goodies

Thumbnail rust-glancer.github.io
100 Upvotes

r/rust • • 1d ago

🧠 educational Discovering the language: labeled block - implement early return style control flow w/o sparate function

6 Upvotes

There are situations when decisions must be made based on many variables, and in some points in evalution process, additional costly operations must be performed in order to decide. These decisions are difficult to implement in the form of a single logical expression, and even if it is possible, it is unreadable and hard to modify.

Usually the best pattern for this is to create a separate function which implements the decision, where we can use early returns: the trivial and easy cases that can be decided based on a simple condition (especially exceptions) are evaluated first, we return with the result as soon as possible, so as we're going foward, we reduce the complexity of the remaining cases, and the last case is often a simple condition.

Sometimes a separate function is not an option, because too many parameters would have to be passed (and returned, but Rust have tuples for it), or it simply just does not feel right to split a single decision into two functions, I think, the single-responsibility principle (SRP) must work this way, too. It's even more true for not too complex but nested cases (certain part of the condition is consist of more sub-conditions).

In C and C++, I usually implement this pattern by creating a do..while(false) loop, from which I break at several points, skipping the rest of the evalutation. I store the result in a variable declared just before the block, usually initialized to the default value.

It may just be my fault that I haven't studied the Rust textbook enough, but I've only found the functionally equivalent syntax for this today, and I am very happy with it, with its flexibility and elegance: labeled block.

Let's see an example!

The decision is about whether close (do_close_old flag) the old time window and/or replace with a new one (do_replace_window flag), also log some info about the decision (replace_reason enum), see example (close to actual code):

let (do_close_old, do_replace_window, replace_reason) = 'switch: {

    if just_created {
        break 'switch (false, false, Reason::Create);
    }

    if window.is_retired {
        break 'switch (false, true, Reason::Replace);
    }

    if window.is_expired(timestamp, lifetime_duration) {
        break 'switch (true, true, Reason::Timeout);
    }

    // complicated condition, fake
    let compli = if (x && y) | (a && !b);
    let cation = if z > w {
      countries.find("USA") && languages.find("English")
    } else {
      false
    };
    if compli && cation {
        break 'switch (true, true, Reason::ComplicationHappened);
    }

    (false, false, Reason::None)
};

(Please, don't review my code, I know, I know, somehow the result tuple should be replaced with some named thing.)

I think, it's pretty well readable, even if you haven't met with labeled blocks before. Just as me, until today.

AI disclaimer:

  • helped in translation (my English is not suitable for publications)
  • the pattern was also suggested by AI, I often ask it to refactor short code snippets

r/rust • • 14h ago

🙋 seeking help & advice Which GUI crate would you choose for this project?

0 Upvotes

Hi everyone! I asked this earlier in the week, but I don't think I explained my question very well (sorry for the translation; I'm a Spanish speaker).

I want to learn how to develop graphical user interfaces (GUIs).

I'm generally new to programming, but I really like Rust; I don't know any other languages—just Rust.

I want to build a project for the company I work for. We mostly use Excel for our data; our "database" (an Excel file) is 500k rows by 25 columns (though there are several files covering different areas, so the total exceeds a million data points).

I want to automate this. That million data points is spread across years; the most critical process involves each of us reviewing 100 items in 30/40 minutes—classifying, validating, sorting, etc.

I need to create something like dashboards featuring:

charts and tables that are sortable and filterable.

These two components are absolutely critical.

For the rest, I'm thinking of a data grid (something like Excel, but not a full spreadsheet—just small tables with editable cells).

We need this because our values ​​aren't always 100% accurate; there's a lot of human error involved in our work.

For example: Our system registers a client who has a discount percentage, but that wasn't included in the initial registration request. So, the discount doesn't apply the first time. Later, we receive the correction, and the discount applies to the second sale—at which point we also correct the first sale.

This happens across all areas, so I need something editable. Given our limited review time, I need the interface to be fast and intuitive.

I don't need mobile support—though it would be a nice-to-have, it's not essential. I need a recommendation focused 100% on desktop.

It also needs to properly support Windows, Mac, and Linux.

Which GUI crate would you choose? (I know it's difficult for me, but I want to learn a lot)


r/rust • • 2d ago

🎙️ discussion Google is doing "Large Scale Codebase Migrations and Optimizations" of C/C++ to Rust

Thumbnail blog.google
670 Upvotes

r/rust • • 1d ago

Proficient in both Rust and C, when to pick C?

46 Upvotes

What would be your criteria for a new project (or parts of it) to be better fit for plain C in 2026?


r/rust • • 1d ago

The Second Golden Spike: Memory Safety Across the Valen/Rust Boundary

Thumbnail verdagon.dev
30 Upvotes

r/rust • • 19h ago

🛠️ project ratatui-hypertile 0.4.2: palette controls, stable tab IDs, core without Crossterm, 10.000 downloads and a reality check on AI

0 Upvotes

Hello, Rustaceans!

I released 0.4.2 yesterday and wanted to introduce people who might be interested in the project, talk about the changes in the release, where I plan on taking the project in the future and a quick rant on the AI policy the project will have going forward.

First of all, I want to thank @ Seuros, TheAiteb and tom-lubenow for their contributions in the past few weeks and months. This is my first ever open source project and I'm glad that people didn't only find a use case for it but also actively helped improve it!

What has changed in v0.4.2:

  • More control over the palette. Choose which plugin types it offers, open and render it from your app, and receive the selection to handle yourself. Useful when choosing a plugin should focus an existing pane or follow your own application logic.
  • Stable tab IDs. Each workspace tab now has a TabId, with helpers to check which tabs remain open. This makes tracking state such as notifications per tab easier.
  • The core no longer pulls in Crossterm. It now builds for Wasm targets, including wasm32-wasip2. Thanks again to Tom for this.
  • Fixed a bug causing keyboard actions to fire twice on Windows (already fixed in 0.4.1 but I didn't make a post about it)

AI Policy

It's undeniable that LLMs today are not the ones from last year, they are more than capable of producing quality code, finding bugs and improving performance. Hypertile DOES have AI generated code in it, even though it started out as my little learning project. That being said, as of yesterday, the project will adopt ripgreps AI policy as well as an AGENT.md linking to .rules, this one being adopted from Zed. While I do not have any plans on trying to create an artisan or sophisticated guide for LLMs, I am not delusional enough to believe that people contributing won't use LLMs to generate their code. I have also adopted policies from the Linux Kernel regarding the use of coding assistants of any kind.

TLDR: You can use AI for code, you must disclose what you used, you may not communicate in PRs, Issues or anywhere using LLMs. You own the code you publish regarding of how it was created.

FAQ:

I did not think I'd actually need a FAQ question but one thing that constantly keeps popping up is:

Why not just use tmux or ZelliJ?

They are not mutually exclusive! The point of Hypertile is allowing people to write their CLI tools in a way that allows them to be modified at runtime instead of the look and feel being baked in at compile time as most TUIs do. You can use your ratatui-hypertile app in ZelliJ just fine.

As always, any and all feedback is welcome! :)

Repository - Core crate - Extras crate


r/rust • • 1d ago

🛠️ project Arm64 emulator in Rust that compiles into wasm and allows running Alpine Linux in a web browser

18 Upvotes

Hi all,

Last couple of months I've been working on the largest Rust project I've ever built (also because everything I built before were small tools), happy to share it now.

It started with me getting annoyed that to try a 2 MB TUI tool I have to install it, while my browser happily runs a 40 MB landing page with React doing massive rerenderings on every move. So I thought it MUST be possible to run a thin Linux with TUI app in the browser, taking into consideration the power of CPUs we have these days.

Many months of grinding resulted in Arm64JS. It's an interpreter runs a small VM in web browser and emulates a minimal hardware stack, so VM thinks it runs on real hardware. The whole emulator is Rust compiled to wasm32, one module instantiated per web worker, all sharing one SharedArrayBuffer as the VM's "RAM".

CPU speed is of course lower than native — but works fine for shells, TUI apps, Python/PHP and small servers. Typing lag 8 ms in the shell, 37 ms in vim, booting from a snapshot a bit more than a second (plus download if data are remote).

The code of the interpreter is not OSS at the moment, but I've made a small JS SDK to make it easy for others to create their own VMs. It's free for non-commercial use.

Feedback is welcome.

You can read more about its story here: https://arm64js.com/blog/tui-apps-in-the-browser/

And here's the link to the JS SDK I mentioned: https://github.com/kooler/arm64js-sdk


r/rust • • 2d ago

📸 media Three years migrating latency-sensitive services to Rust and one Tokio failure

Thumbnail image
391 Upvotes

tl;dr Over three years, we moved our latency-sensitive services to Rust. LB (load balancer) latency dropped from 600ms to 101ms, publish API latency from ~350µs to ~50µs, and a later Presence API redesign cut peak memory about 6x. We also let an unbounded Tokio workload turn a 100 MiB pod into a 3.7 GiB pod.

Two years ago, I wrote about moving one data-pipeline service from Python to Rust in 120ms to 30ms: Python 🐍 to Rust 🦀🚀. Since then, we have moved the rest of our performance-sensitive stack.

Disclosure: I worked on the original C systems years ago. Now a new team rebuilt these systems in Rust.

Our core services were a mix of Python, Go, JVM, and C. We migrated the critical ones one service and one region at a time, running old and new implementations side by side until we had enough data to move the remaining traffic.

The rewrites cut latency and resource use. Memory behavior became easier to reason about, and the team prefers working in the new codebases.

Once latency became more stable, we could see details the old runtime noise had hidden. A 100µs excursion now sticks out and we can easily investigate changes of 30µs.

Most panels show a baseline and cutover. Panels 8 and 9 compare Rust with Rust. Panel 14 shows a rolling replacement; panels 4 and 13 show steady state only.

Reddit post gets only one image, so I tiled fifteen numbered panels together in the order discussed below.

The nginx replacement

Our load balancer decides where each message goes next. Its average time had sat at 600ms for so long that we had stopped questioning it.

We replaced the nginx-based balancer with Pingora, Cloudflare's Rust proxy framework. We kept the same box, traffic, and routing decisions.

Panel 1: Balancer time, 600ms to 101ms.

The drop starts around 09:40 as the Rust balancer takes over. We shifted traffic in two increments, which produced the brief ledge at ~300ms. Once the cutover finished, latency held at 101ms on the same hardware and traffic.

The right edge of panel 1 includes broker latency. That internal hop stays at ~350µs during the cutover. We had already upgraded the broker to Rust, so the balancer accounted for nearly all of the 600ms.

PubSub: lower publish latency and variance

Publish was already under a millisecond, so I didn't expect much room for improvement.

Panel 2: Publish latency, from a 400-500µs band to roughly 40-60µs.

The three series, average_publish, average_internal_publish, and average_signal, sat in the 400 / 500µs band, with a ~1ms excursion around 14:55.

A smaller bump reaches ~600µs at 15:10 as the cutover starts. Over the next few minutes, the band drops to roughly 40-60µs and stays there for the remaining twenty minutes. We watched the service for weeks before calling the migration complete. Neither the spikes nor the unexplained 1ms tail in customer p99 dashboards returned.

The lower variance mattered more to our latency guarantees than the lower average.

Measuring server hops in microseconds

The migration changed how I measure server-side latency. Microseconds are now a useful unit for these internal hops.

Panel 3 shows the same publish path over a day and a half.

Panel 3: Publish latency over 36 hours, ~350µs to ~50µs.

Green (average_internal_publish) fluctuates between 350µs and 460µs on the left, while orange and yellow (average_signal, average_publish) sit around 300 / 350µs. Around 10-22 08:00, green joins them near 300µs. Just before 10:00, all three fall into a 40 / 60µs band and stay there for the next 24 hours.

The hottest path went from ~350µs to ~50µs, with less variance. The entire chart stays under 500µs, and small regressions are easier to spot.

On the right, green rises to ~100 / 185µs around 10-23 00:00 / 04:00 while orange and yellow stay near 40-60µs. The bump affects internal publish only. A 100µs excursion disappeared inside the old stack's normal variance; here it is clear enough to investigate.

The replication layer runs at 27µs to 31µs.

Panel 4: Replication average latency, 27-31µs.

This steady-state panel shows average replication latency for two nodes, average_21 and average_23, across a vertical range of 27µs to 31µs.

Green moves between 28.5µs and 30.5µs. Yellow sits at 27.5 / 29.5µs and follows the traffic cycle. The ~1.5µs gap between the nodes is visible. Three years ago, GC pauses hid it.

GC did not account for all of the old latency. We owned the code and could have kept improving it. The first migration gave us enough evidence that a Rust rewrite would repay its cost.

We can now investigate whether the NIC, the scheduler, or our code makes one node trail another by 1.5µs. A 5µs improvement in this 30µs hop will show up.

At 3 trillion API calls a month, even single-digit-microsecond changes are measurable.

Presence: latency and memory

Presence tracks who is online and which channel they occupy across regions. It holds state for millions of concurrent occupants while heartbeats arrive, expire, and reconcile across the network. Presence had caused more incidents than any other service, so this was our riskiest migration.

Subscribe join latency, the presence online status event

Panel 5: Presence subscribe join latency (median), az1 from 1-3s to 200-250ms.

Green (az1) is on the old stack, moving between 1s and 3s with a 2.9s median spike at ~13:58. Orange and blue (az2, az4) are already on Rust and hold near ~350ms while az1 spikes.

At 14:03, we cut az1 over. It drops to ~200 / 250ms, below az2 and az4 at around 300ms because az1 received a later build. The other regions have since moved to that version and now match it.

Web heartbeat join latency

The web/HTTP path drops from 700ms to 200ms.

Panel 6: Presence web heartbeat join latency, 700ms to 200ms.

The blue refresh series fluctuates between 600ms and 800ms before the 10:30 cutover, then holds near ~200ms. join_announce (green) stays around 100 / 150ms before moving toward 200ms at the right edge. join_interval (yellow) was near zero and stops reporting around 10:25 because we removed that metric during the cutover.

The internal heartbeat fanout

Panel 7: Subscribe internal heartbeat latency, 1.5-2s to 500ms.

Three regions sit in the 1.4s / 1.8s range because they share an overloaded path. One spikes to 2.1s during the cutover. Green follows the same rhythm at a lower ~1s baseline. After the ~17:35 cutover, all four settle into a ~500 / 600ms band with less variance, about three times faster than before.

Presence memory

Panel 8: Presence pod memory, Rust vs Rust, a 3.4 GiB peak down to a 256-512 MiB band.

Panel 8 shows memory per pod in the presence namespace. Both halves are presence-rust deployments from two ReplicaSets, 6cdfbf75d5 and 6f78d4f459. This is a comparison between Rust designs, separate from the language migration.

In the old ReplicaSet, the top pod peaks at 3.4 GiB while the rest of the fleet ranges from 512 MiB to 2.5 GiB. Memory declines over more than a day as the service trims accumulated occupancy state. That build still carried substantial per-occupant overhead.

After the 04/23 15:00 / 18:00 gap, the new ReplicaSet holds within a ~256 / 512 MiB band. The growth curve disappears, and pods no longer approach an OOMKill during peak traffic.

Removing the GC made the profile easier to reason about, but we still had to fix the data layout. The first Rust build beat the Python service it replaced; the redesign cut peak memory by another 6x.

Push notifications: FCM

Push is a fan-out-and-wait workload. Apple and Google account for much of the latency, so our part needs to stay small and consistent.

Panel 9: FCM Rust average time, Rust vs Rust, from 135-250ms to roughly 135ms.

The single series, fcm rust avg, is the Rust path on both sides of the chart. The left side ranges from 135ms to 250ms and spikes at 08:45. Just after 09:00, it settles at ~135ms for the rest of the window.

At 09:00, we enabled connection pooling and reuse against FCM. The low runtime noise let us isolate the remaining variance and trace it to our code rather than Google's service.

That flat line is my favorite kind of graph.

Event Processing: fewer warnings and retries

Panel 10 counts requests that hit a degraded or retry-worthy path on internal_heartbeat, a high-volume route through event processing.

Panel 10: API warnings on the internal_heartbeat route, 1.5-4.9K down to under 250.

Green (az2) sustains 1.55K to 4.9K warnings throughout the afternoon. Yellow (az4), already migrated, usually stays between 100 and 500 with occasional excursions to ~860.

At 17:03, az2 drops into the same range as az4. Both spike at 17:20, when green reaches 1.02K and yellow ~740, then touch 400 / 600 a few times through 17:45. After ~17:50, both usually stay under ~250 with occasional spikes to ~400.

With roughly an order of magnitude fewer warnings and retries, alerts stand out. We no longer tune thresholds to ignore the service's normal behavior.

Memory and CPU

The migrations reduced memory and CPU use as well.

Memory, per shard

Panel 11: Shard 5 max memory, 95% down to a 20-40% band.

This panel covers shard 5 across all regions for two days. On the left, red sits between 90% and 97%, above the critical threshold, while yellow falls from ~90% to 85%. One traffic spike could have exhausted the remaining memory.

The cutover runs from 10-06 22:30 to after 10-07 00:00, when every series lands in the 30 / 40% band. Peak traffic on 10-07 between 12:00 and 16:00 pushes usage to ~69%. Overnight it falls to 12 / 25%, then returns to 20 / 37% with daytime traffic.

On the same boxes, peak memory moved from the critical range to below the 75% warning threshold.

Memory, per region

Panel 12: Max memory usage by region, with the step down at cutover.

The regional view shows the same step down on 06/22. The tooltip values come from after the cutover. Before it, the lines seesaw as the runtime allocates and collects memory.

After the cutover, the lines become thinner and smoother. iad remains highest at ~72% because this fleet still included unmigrated capacity; the other four regions had completed the rollout. We moved each region only after comparing it with capacity still running the old implementation.

(This iad fleet is now fully migrated. Other services there moved earlier; panels 8 and 10 both show iad clusters. We rolled out each service on its own schedule.)

CPU

Panel 13: Max CPU usage across regions, holding in the 0-40% band.

Across five days, regional CPU stays in the 0 / 40% band with occasional peaks at 60 / 70%. It never reaches the 75% warning or 90% critical lines. The sawtooth pattern follows normal daily traffic.

CPU per pod

Panel 14: CPU usage by pod, 0.5 cores down to 0.22 during a rolling replacement.

This panel shows per-pod CPU during a rolling replacement. The old pods use 0.4 / 0.6 cores before they drain and terminate.

The new pods stabilize at ~0.22 cores, about half the CPU for the same work. We used the savings to run fewer pods with more headroom.

Delivery semantics

We required the new system to match or improve our delivery guarantees. The delivery changes mattered more to us than the latency graphs.

For years, we replicated messages over a hand-rolled TCP protocol. It moved trillions of messages and supported retry and redelivery. It also left us responsible for framing, backpressure, reconnect logic, sequence numbers, and versioning between old and new nodes. Only a few engineers knew all of its failure modes.

We replaced the protocol with gRPC and a store-and-forward model.

Store and forward: persist first, stream second.

We replicate an incoming event to multiple physical nodes while beginning to stream it. We return delivery status only after the durable write. If the stream breaks or the downstream service is rolling, we query the replicas for messages that still need redelivery.

Streaming first leaves an unrecoverable gap if the process crashes between "sent" and "recorded." Persisting first may produce a duplicate delivery attempt. The record remains on the replica nodes, and an idempotency key handles the duplicate. The wire guarantee is at least once; the idempotency key makes delivery effectively once at the application layer. We do not claim exactly-once delivery over an unreliable network. For acknowledged messages, the durable replicas preserve a redelivery path, and the idempotency key prevents a repeated attempt from becoming an observable duplicate.

gRPC's automatic retry policy, configured at the service level.

The gRPC service config defines retryable status codes, attempt limits, and backoff. The channel handles retries below the application call, so many brief network failures no longer reach application code.

Previously, each service implemented its own retry policy. Conservative policies risked dropped messages; aggressive ones could overload a service during a brief failure.

gRPC absorbs transient failures such as momentary TCP resets. Store and forward handles failures that outlast the retry policy.

Rust's compile-time concurrency checks gave us more confidence in the new replication model. We ran both implementations side by side and watched the dashboards shown here.

Why Rust fit this workload

Rust suited this workload: a latency-sensitive message bus that holds millions of concurrent occupants in memory. These results do not mean every CRUD service needs a rewrite.

We spent years profiling and tuning the existing services. Managed runtimes can go far, but the collector still decides when to pause the program. The 1ms publish excursion and 2s heartbeat oscillation show the latency cost; the regional memory chart shows the allocation churn. The Rust services no longer incurred GC pauses.

We use Go for many API services. For this part of the message bus, we could not meet our latency and memory targets without removing GC pauses and reducing per-object overhead.

Our C codebase was 14 years old and needed a major update for current libraries and compilers. Since we faced a substantial rewrite either way, we chose Rust. We could have made some of the same performance improvements in C, and our measurements put Rust within the noise of C. The practical difference was that more engineers felt comfortable changing the concurrent Rust code, while the compiler caught ownership and data-race mistakes before deployment.

The first six months with the borrow checker were rough. Experienced engineers got frustrated, PRs stalled, and people questioned the migration. By month eight, review cycles had shortened and we spent less time debugging ownership mistakes after compilation.

cargo helped too. A shared toolchain, formatter, test runner, and build command save time in infrastructure spread across four languages.

Our Tokio mistake: unbounded in-flight work

Panel 15 shows the costliest mistake in the migration.

Panel 15: Memory usage by pod, from a ~100 MiB baseline to a 3.7 GiB burst.

Bursts arrived faster than we could complete them. We spawned a task per request, Tokio queued the tasks, and each one held buffers and state while waiting on I/O. Leak detection found nothing because the memory was live. Without an explicit cap, in-flight work grew with each burst. Waiting tasks use little CPU, but their state still consumes memory.

We first tuned the HPA with lower scale-up thresholds and faster reaction windows. It reacted sooner, but scaling out gave the unbounded backlog more places to accumulate.

We added backpressure at the ingest edge. A cap on in-flight work sends the burst to a bounded queue instead of letting tasks accumulate on the heap. Excess work waits or gets shed, so memory no longer grows with the arrival rate. We deployed the change in every region, and the same pods now hold their baseline under comparable bursts instead of climbing into GiB. I don't yet have a matching after panel for a comparable burst.

Two years ago, my Python-to-Rust post included mpsc::channel(100) and recommended Tokio MPSC for ingest paths. We failed to follow that advice here. Async can schedule large numbers of in-flight tasks at low CPU cost, but every task still holds state. Without a limit somewhere in the path, bursts consume the available memory.

Advice after three years

  • Instrument before you start. Every dashboard in this post existed before the migration, so we can support claims such as "600ms to 101ms."
  • Budget for the learning curve. This took three years, and the first six months were the most expensive.
  • Optimize for predictable latency. Customers notice the spikes more than the average.
  • Treat the first Rust release as a baseline. Panels 8 and 9 compare Rust with Rust: Presence peak memory fell about 6x, and connection pooling removed most FCM variance.
  • Work on durability alongside performance. We moved from our TCP protocol to gRPC and store and forward to improve correctness.

This is a follow-up to my 2024 Python-to-Rust post. I could not separate the benefits of Rust from the benefits of rewriting the architecture in that first migration. Panels 8 and 9 provide useful counterexamples: both compare Rust with Rust, and both show that data layout and connection reuse still mattered after the language migration.

Across these services, we now use less memory and about half the CPU per pod, and we see an order of magnitude fewer warnings on the route shown in panel 10. We also retired a wire protocol that only a few engineers understood. Those results justified the three-year migration for this workload.


r/rust • • 2d ago

🗞️ news Rust Berlin Talks (30/09/2026) - Livestream recording

Thumbnail youtube.com
22 Upvotes

Lineup:

  • Egor Lebedev - Welcome and a short note from the RustRover team
  • Iván Ovejero - Compiling Rust Ideas into Go: Learnings from Lisette
  • Orhun Parmaksiz - Debugging Rust in 2026

Event page: https://www.meetup.com/rust-berlin/events/316661690/


r/rust • • 2d ago

📅 this week in rust This Week in Rust #671

Thumbnail this-week-in-rust.org
41 Upvotes

r/rust • • 1d ago

RV Bare-metal?

4 Upvotes

I'm a beginner in OS dev and have been tinkering with creating a RISC-V operating system in Rust. There are a lot of great resources out there, but most of the posts I found are slightly outdated, so I wanted to write a bit about my own journey. My goal is to use Rust's features as much as possible!

First, we'll go over creating a bare-metal binary and finish with a small executable that immediately shuts down the machine when it runs in QEMU. Hope you find this interesting, as I did!

https://ltungv.com/note/rv-bare-metal/


r/rust • • 1d ago

🙋 seeking help & advice Does probe-rs support external flashing?

0 Upvotes

If yes is there a tutorial or a guide somewhere? I am using the STM32H7R3Z8J6 dev board from WeAct studio.