r/rust • • 13d ago

📢 announcement No More Code Dumps

1.6k Upvotes

TL;DR

Code Dumps, ie links to code repositories, are no longer welcome on r/rust.

If you are excited about a project you built, and just wish to share it, consider posting a casual top-level comment in the weekly What's Everyone Working on this Week? thread.

History

When r/rust was created, in December 2010, Rust was a nascent programming language. It was very different, it was also fast-moving, and it was unclear whether it would amount to much, or not.

In these early days, new projects being written in Rust were both a proof that the language was suitable for a variety of purposes, and a cause for celebration in seeing the language being adopted. Code dumps, then, were both cause to rejoice, and an opportunity to see how well (or not) the language was suited to the task. They were even a good opportunity to learn the latest developing idioms of this fast-changing language.

Fast forward 15 years, and code dumps have become boring. Utterly uninteresting. There's nothing exciting, or newsworthy, about a 10th hypervisor project, a 100th webserver project, or a 1000th TUI project. And for the most there's no exciting idiom, or pattern, to be found in their code either... even if some poor soul bothered to look.

And mass AI generation appeared on the scene, and for the past year(s), the already dwindling appeal of code dumps has just completely vanished, buried in AI slop.

Purpose

r/rust is to discuss all things Rust. The language, the ecosystem, the community.

Of late, most code dumps barely foster any discussion. And when they do, apart from AI witch hunts, those discussions are mostly about the domain of the project, rather than about Rust. Depending on the domain, the discussion may or may not be of interest to r/rust users... but whether it is or not is a matter of chance: they did not come here seeking such a topic. Those discussions are tolerated on posts which are on-topic, but if a post only fosters "tolerated" discussions, then that is a sign that the post itself may simply not really be on-topic for r/rust.

And that is not to say that code dumps are, by nature, low-effort. Surely, a project is created for a purpose, to solve a specific problem that is unsolved, or solve it with different trade-offs than the currently available solutions. None of that is elucidated in a code dump.

New rules

Code dumps are no longer welcome on r/rust.

Code dumps of applications or libraries written in Rust are Low-Effort unless otherwise newsworthy. Newsworthy means that the very existence of the application or library is important, for example consider applications or libraries breaking new ground, such as the first hypervisor, first OS, first XXX standard safety-critical application, first Rust application in a major company, etc...

Announcements of new libraries which are not otherwise newsworthy are only welcome if they tick ALL the following boxes:

  • The library has been worked on regularly for at least 4 months.
  • The library has already been used, or built upon, thereby validating its design to a degree.
  • The announcement is a full-fledged text post, or a link to a full-fledged article. A README, or manual, is not an article.
  • The announcement should motivate the relevance of this library to the wider r/rust community.
  • The announcement should articulate the trade-offs differentiating this library from similar popular libraries in the field.

(It goes without saying, but the announcement should NOT be AI-generated, not even in part, as AI-generated posts & articles violate the Low-Effort rule)

Release notes of libraries are still welcome for popular enough libraries, such as libraries with reverse-dependencies on crates.io pointing to sufficient adoption by the wider community to justify notifying said community of a new release.

Of course, text posts, or links to articles, centered on Rust, with an application or library as context are always welcome. For example, said posts or articles could present gnarly problems the author encountered, and how they solved, or which features of the language or a library helped or hindered them in the making of their project.

(It goes without saying, but said text posts & articles should NOT be AI-generated, not even in part, as AI-generated posts & articles violate the Low-Effort rule)

As a reminder, if an author just wishes to share their joy of building something in Rust, they are welcome to use the What's Everyone Working on This Week? thread. And as per the title, it need not even be a complete/mature project, or an open-source project for that matter as links to code are strictly optional there, it's all about casually chatting about your work on Rust projects.

Q & A

See pinned comment below.


r/rust • • 4d ago

🙋 questions megathread Hey Rustaceans! Got a question? Ask here (40/2026)!

11 Upvotes

Mystified about strings? Borrow checker has you in a headlock? Seek help here! There are no stupid questions, only docs that haven't been written yet. Please note that if you include code examples to e.g. show a compiler error or surprising result, linking a playground with the code will improve your chances of getting help quickly.

If you have a StackOverflow account, consider asking it there instead! StackOverflow shows up much higher in search results, so ahaving your question there also helps future Rust users (be sure to give it the "Rust" tag for maximum visibility). Note that this site is very interested in question quality. I've been asked to read a RFC I authored once. If you want your code reviewed or review other's code, there's a codereview stackexchange, too. If you need to test your code, maybe the Rust playground is for you.

Here are some other venues where help may be found:

/r/learnrust is a subreddit to share your questions and epiphanies learning Rust programming.

The official Rust user forums: https://users.rust-lang.org/.

The unofficial Rust community Discord: https://bit.ly/rust-community

Also check out last week's thread with many good questions and answers. And if you believe your question to be either very complex or worthy of larger dissemination, feel free to create a text post.

Also if you want to be mentored by experienced Rustaceans, tell us the area of expertise that you seek. Finally, if you are looking for Rust jobs, the most recent thread is here.


r/rust • • 10h ago

Compiling the Linux kernel with gccrs | RustConf 2026

Thumbnail youtu.be
34 Upvotes

r/rust • • 2h ago

Embedded std?

8 Upvotes

I was wondering if there were any projects like micro python to port std in to make a barebones os designed for microcontrollers? I understand that the std relies on OS level syscalls, so can we do something like implement a basic fat32 file system to accommodate certain write or read requests. In general this would be used to run std in projects without a OS? Does this already exist, and how hard would it be to make a basic version?


r/rust • • 20h ago

📡 official blog Generic Const Args and You | Inside Rust Blog

Thumbnail blog.rust-lang.org
210 Upvotes

r/rust • • 7h ago

🛠️ project Enki – Write GPU compute kernels in pure stable Rust

20 Upvotes

Hey guys!

I've been building a GPU compute platform on pure stable-rust to let you write compute kernels directly in standard rust, Built entirely on Vulkan 1.3 (Compute) for the host api runtime & MLIR/LLVM for the backend JIT compiler.

It can make you write iters-enums-generics-etc.. in your function, and run it across GPU silicon, because the syntax is literally rust, It can execute on native CPU.

Here is a simple example:

```rust use enki::*; use glam::Vec2;

const COUNT: usize = 5;

// Declare the compute kernel with #[nam]

[nam]

fn scale_vectors(_space: &Space, input: &Vec2, output: &mut Vec2, factor: f32) { *output = *input * factor; }

fn main() { // Initialize the headless GPU runtime let enki = Enki::init();

// Allocate physical data directly in GPU VRAM
let in_gpu = gpu_vec![
    Vec2::new(1.0, 2.0),
    Vec2::new(3.0, 4.0),
    Vec2::new(5.0, 6.0),
    Vec2::new(7.0, 8.0),
    Vec2::new(9.0, 10.0),
];
let mut out_gpu = gpu_vec![Vec2::ZERO; COUNT];

let factor = 2.5f32;

// Record and dispatch directly to GPU silicon
enki.flow(|_| {
    scale_vectors.run(
        &Space::gpu_x(COUNT),
        &in_gpu,
        &mut out_gpu,
        GpuParam::new(factor),
    );
});

// Dual Execution: Run the exact same function on CPU native rust
let in_cpu = vec![
    Vec2::new(1.0, 2.0),
    Vec2::new(3.0, 4.0),
    Vec2::new(5.0, 6.0),
    Vec2::new(7.0, 8.0),
    Vec2::new(9.0, 10.0),
];
let mut out_cpu = vec![Vec2::ZERO; COUNT];

for i in 0..COUNT {
    scale_vectors(&Space::cpu_x(i, COUNT), &in_cpu[i], &mut out_cpu[i], factor);
}

// Verify bit-for-bit equivalence
assert_eq!(&out_cpu[..], &out_gpu.to_vec()[..]);
println!("\nGPU Results: {:?}", out_gpu);
println!("\nCPU Results: {:?}", out_cpu);
println!("\nCPU and GPU outputs match.");

} ```

Result:

``text enki_matmul on  master [✘!+] is 📦 v0.1.0 via 🦀 v1.98.1 ❯ cargo run Compiling enki_matmul v0.1.0 (/home/mohiman/Projects/enki_demo_pack/enki_matmul) Finisheddevprofile [optimized + debuginfo] target(s) in 1.86s Runningtarget/debug/enki_matmul`

GPU Results: [Vec2(2.5, 5.0), Vec2(7.5, 10.0), Vec2(12.5, 15.0), Vec2(17.5, 20.0), Vec2(22.5, 25.0)]

CPU Results: [Vec2(2.5, 5.0), Vec2(7.5, 10.0), Vec2(12.5, 15.0), Vec2(17.5, 20.0), Vec2(22.5, 25.0)]

CPU and GPU outputs match.

enki_matmul on  master [✘!+] is 📦 v0.1.0 via 🦀 v1.98.1 took 2s
❯ ```

I've verified it on Linux (Docker Container Testing) & Windows (Wine) & colab T4, and my intel UHD laptop, every thing was great, i'd really love your help testing across more diverse hardware (AMD, Nvidia, Intel)

You can test SDF showcase in 3 commands:

bash git clone https://github.com/enkiruntime/enki_sdf.git cd enki_sdf cargo run --release

(pressing space will toggle execution from GPU (Enki) to CPU (rust-rayon))

Enki is currently in Public Alpha (v0.2.x) release, I would really appreciate any feedback, architectural thoughts, or bug reports.


r/rust • • 14h ago

🛠️ project ArcColdString: A 1-word (8-byte) atomically reference-counted SSO string that saves up to 32 bytes over Arc<str>

Thumbnail github.com
53 Upvotes

Disclaimer: Re-post approved by u/matthieum, as previous post was erroneously removed.

I’ve been working on a specialized string type called ColdString. The goal is to create the most memory-efficient string representation possible.

  • Size: Exactly 1 usize (8 bytes on 64-bit).
  • Inline Capacity: Up to 8 bytes (Small String Optimization).
  • Niche Optimization: Option<ColdString> has no memory overhead
  • Heap Overhead: Only 1–9 bytes (VarInt length header) instead of the standard 16-byte (pointer, length) pair.

(Since my last post, the 8th inlineable byte and null-niche optimization were suggestions from the community!)

ArcColdString

I'm presenting ArcColdString, a reference counted ColdString with 8 bytes overhead (inspired by arcstr).

  • Same size and inlining rules as ColdString. Inlined strings are copied instead of reference counted.
  • Heap Overhead: 8 bytes for the AtomicUsize reference count (in addition to the VarInt header).
  • Smaller Reference Counts: ArcColdString32, ArcColdString16, and ArcColdString8 use AtomicU32, AtomicU16, and AtomicU8 counts respectively.

Usage

Available in https://crates.io/crates/cold-string/0.4.0

use cold_string::ArcColdString;

let first = ArcColdString::new("a string longer than one machine word");
let second = first.clone();
assert_eq!(first, second);

Memory Comparisons

Theoretical overhead on a 64-bit target, excluding the UTF-8 payload:

Type 8 bytes 128 bytes 512 bytes
Arc<str> 32 32 32
arcstr::ArcStr 24 24 24
ArcColdString 0 (inline) 17 18
ArcColdString32 0 (inline) 13 14

RSS bytes per unique string in a pre-sized Vec:

Type 8 bytes 128 bytes 512 bytes
Arc<str> 47.1 175.4 560.0
arcstr::ArcStr 39.1 167.5 552.2
string_cache::DefaultAtom 71.5 199.5 586.0
ArcColdString 8.0 167.5 552.1
ArcColdString32 8.0 151.4 536.8

(If I'm not mistaken, string_cache is not a 1:1 comparison since it has to hash string contents, which is 40 bytes overhead?).

Reference Counting Performance

ArcColdString's primary goal is memory and portability. Second to that, my goal is to have ArcColdString's referencing counting performance comparable to Arc<str> and arcstr::ArcStr. While drop and clone speed is comparable now, there's still room for improvement.

Below are single threaded clone and drop measurements, done using criterion on AMD Ryzen 9 5900X 12-Core Processor (3.70 GHz):

Clone (ns) 16 bytes 128 bytes 512 bytes
Arc<str> 3.86 3.82 3.82
arcstr::ArcStr 3.39 3.41 3.43
ArcColdString 3.44 3.45 3.43
Drop (ns) 16 bytes 128 bytes 512 bytes
Arc<str> 2.43 2.49 2.43
arcstr::ArcStr 2.52 2.53 2.52
ArcColdString 2.17 2.18 2.19

Below benchmarks measured 4 threads cloning and dropping strings in a shared size pool of 1024, 16, and 1 string(s), using criterion on Intel(R) Core(TM) i7-9750H CPU @ 2.60GHz (2.59 GHz). The more strings in the pool, the less contention:

4 Threads Clone + Drop (ns) 1024 Strings 16 Strings 1 String
Arc<str> 26.55 54.69 144.81
arcstr::ArcStr 29.29 119.39 214.31
cold_string::ArcColdString32 25.69 90.53 262.83

You can read more benchmarks and implementation details in https://github.com/tomtomwombat/cold-string


r/rust • • 20h ago

📡 official blog Demoting i686 Windows targets to std-only

Thumbnail blog.rust-lang.org
122 Upvotes

r/rust • • 15h ago

Article from Daniel Lemire: How many strings can you create per second?

Thumbnail lemire.me
52 Upvotes

This is a fairly new article from Daniel Lemire (the SIMD expert, author of simdjson).

It seems to_string() need to use the same trick.

C++ wins by a wide margin at 5.4 ns per string. The trick is the small string optimization: a std::string stores short strings, directly inside the object. Our strings have at most eight digits, so C++ never calls the memory allocator.


r/rust • • 2h ago

"Tick Tock: Maintaining Time" | RustConf 2026

Thumbnail youtu.be
4 Upvotes

r/rust • • 23h ago

🎙️ discussion Tyler Mandry: "Beyond the &: A Future for Native Smart Pointers in Rust" | RustConf 2026

Thumbnail youtu.be
48 Upvotes

r/rust • • 1d ago

📡 official blog Rust 1.99.0 is out

Thumbnail blog.rust-lang.org
697 Upvotes

r/rust • • 16h ago

🎙️ discussion API design question: should a 407 from a proxy check be Ok or Err?

5 Upvotes

I'm working on a Rust SDK that has a small check() method for testing a proxy connection, and I'm not sure I'm using Result in the least surprising way here.

If the proxy replies with 407 Proxy Authentication Required, I currently return Ok(Check), not Err.

My reasoning is basically: the check itself worked) I connected to the proxy and got a real response back. For this method, the status/reason/headers are useful information.

Err is only for cases where I couldn't get a usable response at all - timeout, DNS failure, connection failure, malformed response, etc.

Roughly:

match proxy.check(connect)? {
    check if check.status == 200 => {
        // accepted
    }
    check if check.status == 407 => {
        // proxy replied, auth rejected
    }
    check => {
        // some other resp
    }
}

The part that feels a bit weird is this:

let check = proxy.check(connect)?;

That kind of reads like "the proxy is fine", while what it really means is only "the proxy gave a usable response".

Would you expect 407 to be an Err here, or does treating a protocol-level refusal as a successful diagnostic result make sense? 🙊


r/rust • • 1d ago

Robert Seacord: "Unsafe Rust" | RustConf 2026

Thumbnail youtube.com
62 Upvotes

r/rust • • 1d ago

🎙️ discussion Rust in the kernel? What about Rust without the kernel!

Thumbnail kerkour.com
164 Upvotes

Bare-metal Rust may have a brighter future than micro-kernels.


r/rust • • 1d ago

🛠️ project Rust Glancer 0.3: new trait solver and other goodies

Thumbnail rust-glancer.github.io
95 Upvotes

r/rust • • 23h ago

🧠 educational Discovering the language: labeled block - implement early return style control flow w/o sparate function

6 Upvotes

There are situations when decisions must be made based on many variables, and in some points in evalution process, additional costly operations must be performed in order to decide. These decisions are difficult to implement in the form of a single logical expression, and even if it is possible, it is unreadable and hard to modify.

Usually the best pattern for this is to create a separate function which implements the decision, where we can use early returns: the trivial and easy cases that can be decided based on a simple condition (especially exceptions) are evaluated first, we return with the result as soon as possible, so as we're going foward, we reduce the complexity of the remaining cases, and the last case is often a simple condition.

Sometimes a separate function is not an option, because too many parameters would have to be passed (and returned, but Rust have tuples for it), or it simply just does not feel right to split a single decision into two functions, I think, the single-responsibility principle (SRP) must work this way, too. It's even more true for not too complex but nested cases (certain part of the condition is consist of more sub-conditions).

In C and C++, I usually implement this pattern by creating a do..while(false) loop, from which I break at several points, skipping the rest of the evalutation. I store the result in a variable declared just before the block, usually initialized to the default value.

It may just be my fault that I haven't studied the Rust textbook enough, but I've only found the functionally equivalent syntax for this today, and I am very happy with it, with its flexibility and elegance: labeled block.

Let's see an example!

The decision is about whether close (do_close_old flag) the old time window and/or replace with a new one (do_replace_window flag), also log some info about the decision (replace_reason enum), see example (close to actual code):

let (do_close_old, do_replace_window, replace_reason) = 'switch: {

    if just_created {
        break 'switch (false, false, Reason::Create);
    }

    if window.is_retired {
        break 'switch (false, true, Reason::Replace);
    }

    if window.is_expired(timestamp, lifetime_duration) {
        break 'switch (true, true, Reason::Timeout);
    }

    // complicated condition, fake
    let compli = if (x && y) | (a && !b);
    let cation = if z > w {
      countries.find("USA") && languages.find("English")
    } else {
      false
    };
    if compli && cation {
        break 'switch (true, true, Reason::ComplicationHappened);
    }

    (false, false, Reason::None)
};

(Please, don't review my code, I know, I know, somehow the result tuple should be replaced with some named thing.)

I think, it's pretty well readable, even if you haven't met with labeled blocks before. Just as me, until today.

AI disclaimer:

  • helped in translation (my English is not suitable for publications)
  • the pattern was also suggested by AI, I often ask it to refactor short code snippets

r/rust • • 8h ago

🙋 seeking help & advice Which GUI crate would you choose for this project?

0 Upvotes

Hi everyone! I asked this earlier in the week, but I don't think I explained my question very well (sorry for the translation; I'm a Spanish speaker).

I want to learn how to develop graphical user interfaces (GUIs).

I'm generally new to programming, but I really like Rust; I don't know any other languages—just Rust.

I want to build a project for the company I work for. We mostly use Excel for our data; our "database" (an Excel file) is 500k rows by 25 columns (though there are several files covering different areas, so the total exceeds a million data points).

I want to automate this. That million data points is spread across years; the most critical process involves each of us reviewing 100 items in 30/40 minutes—classifying, validating, sorting, etc.

I need to create something like dashboards featuring:

charts and tables that are sortable and filterable.

These two components are absolutely critical.

For the rest, I'm thinking of a data grid (something like Excel, but not a full spreadsheet—just small tables with editable cells).

We need this because our values ​​aren't always 100% accurate; there's a lot of human error involved in our work.

For example: Our system registers a client who has a discount percentage, but that wasn't included in the initial registration request. So, the discount doesn't apply the first time. Later, we receive the correction, and the discount applies to the second sale—at which point we also correct the first sale.

This happens across all areas, so I need something editable. Given our limited review time, I need the interface to be fast and intuitive.

I don't need mobile support—though it would be a nice-to-have, it's not essential. I need a recommendation focused 100% on desktop.

It also needs to properly support Windows, Mac, and Linux.

Which GUI crate would you choose? (I know it's difficult for me, but I want to learn a lot)


r/rust • • 14h ago

🛠️ project ratatui-hypertile 0.4.2: palette controls, stable tab IDs, core without Crossterm, 10.000 downloads and a reality check on AI

0 Upvotes

Hello, Rustaceans!

I released 0.4.2 yesterday and wanted to introduce people who might be interested in the project, talk about the changes in the release, where I plan on taking the project in the future and a quick rant on the AI policy the project will have going forward.

First of all, I want to thank @ Seuros, TheAiteb and tom-lubenow for their contributions in the past few weeks and months. This is my first ever open source project and I'm glad that people didn't only find a use case for it but also actively helped improve it!

What has changed in v0.4.2:

  • More control over the palette. Choose which plugin types it offers, open and render it from your app, and receive the selection to handle yourself. Useful when choosing a plugin should focus an existing pane or follow your own application logic.
  • Stable tab IDs. Each workspace tab now has a TabId, with helpers to check which tabs remain open. This makes tracking state such as notifications per tab easier.
  • The core no longer pulls in Crossterm. It now builds for Wasm targets, including wasm32-wasip2. Thanks again to Tom for this.
  • Fixed a bug causing keyboard actions to fire twice on Windows (already fixed in 0.4.1 but I didn't make a post about it)

AI Policy

It's undeniable that LLMs today are not the ones from last year, they are more than capable of producing quality code, finding bugs and improving performance. Hypertile DOES have AI generated code in it, even though it started out as my little learning project. That being said, as of yesterday, the project will adopt ripgreps AI policy as well as an AGENT.md linking to .rules, this one being adopted from Zed. While I do not have any plans on trying to create an artisan or sophisticated guide for LLMs, I am not delusional enough to believe that people contributing won't use LLMs to generate their code. I have also adopted policies from the Linux Kernel regarding the use of coding assistants of any kind.

TLDR: You can use AI for code, you must disclose what you used, you may not communicate in PRs, Issues or anywhere using LLMs. You own the code you publish regarding of how it was created.

FAQ:

I did not think I'd actually need a FAQ question but one thing that constantly keeps popping up is:

Why not just use tmux or ZelliJ?

They are not mutually exclusive! The point of Hypertile is allowing people to write their CLI tools in a way that allows them to be modified at runtime instead of the look and feel being baked in at compile time as most TUIs do. You can use your ratatui-hypertile app in ZelliJ just fine.

As always, any and all feedback is welcome! :)

Repository - Core crate - Extras crate


r/rust • • 1d ago

Proficient in both Rust and C, when to pick C?

48 Upvotes

What would be your criteria for a new project (or parts of it) to be better fit for plain C in 2026?


r/rust • • 2d ago

🎙️ discussion Google is doing "Large Scale Codebase Migrations and Optimizations" of C/C++ to Rust

Thumbnail blog.google
669 Upvotes

r/rust • • 1d ago

The Second Golden Spike: Memory Safety Across the Valen/Rust Boundary

Thumbnail verdagon.dev
29 Upvotes

r/rust • • 1d ago

🛠️ project Arm64 emulator in Rust that compiles into wasm and allows running Alpine Linux in a web browser

17 Upvotes

Hi all,

Last couple of months I've been working on the largest Rust project I've ever built (also because everything I built before were small tools), happy to share it now.

It started with me getting annoyed that to try a 2 MB TUI tool I have to install it, while my browser happily runs a 40 MB landing page with React doing massive rerenderings on every move. So I thought it MUST be possible to run a thin Linux with TUI app in the browser, taking into consideration the power of CPUs we have these days.

Many months of grinding resulted in Arm64JS. It's an interpreter runs a small VM in web browser and emulates a minimal hardware stack, so VM thinks it runs on real hardware. The whole emulator is Rust compiled to wasm32, one module instantiated per web worker, all sharing one SharedArrayBuffer as the VM's "RAM".

CPU speed is of course lower than native — but works fine for shells, TUI apps, Python/PHP and small servers. Typing lag 8 ms in the shell, 37 ms in vim, booting from a snapshot a bit more than a second (plus download if data are remote).

The code of the interpreter is not OSS at the moment, but I've made a small JS SDK to make it easy for others to create their own VMs. It's free for non-commercial use.

Feedback is welcome.

You can read more about its story here: https://arm64js.com/blog/tui-apps-in-the-browser/

And here's the link to the JS SDK I mentioned: https://github.com/kooler/arm64js-sdk


r/rust • • 12h ago

🎙️ discussion Rust in Model Research and Development

0 Upvotes

Hi everyone,

I'm Abinash. For the past few weeks, I have been working on building a neural network in Rust to test the idea of using Rust in model development and research.

Rust is being extensively used in LLM infrastructure and is growing in the inference space as we get official libraries to run CUDA kernels from Rust binaries.

But if you see the model design, research or development space, Rust is non-existent.

So I thought I'd give it a try. I tried to build a fairly easy neural network to solve the MNIST handwritten digit identification problem.

I know it's not a huge milestone. But I tried to give it a try to test my idea how easy it is to develop neural networks in Rust.

The neural network I'm developing has 2 hidden layers with 512 and 128 neurons, respectively. The goal is to achieve 95% accuracy on my local CPU-only system.

I got the input system right; it can now take training and testing datasets and prepare the matrices for the training and testing stages.

After that, I plan to start developing a real transformer-based LLM from scratch in Rust.

I'm not sure how it will play out, but I will give it a try.

I'd love to have your thoughts on it. If you're a senior ML engineer or MTS at a frontier lab, I'd love to have your thoughts on it.


r/rust • • 2d ago

📸 media Three years migrating latency-sensitive services to Rust and one Tokio failure

Thumbnail image
377 Upvotes

tl;dr Over three years, we moved our latency-sensitive services to Rust. LB (load balancer) latency dropped from 600ms to 101ms, publish API latency from ~350µs to ~50µs, and a later Presence API redesign cut peak memory about 6x. We also let an unbounded Tokio workload turn a 100 MiB pod into a 3.7 GiB pod.

Two years ago, I wrote about moving one data-pipeline service from Python to Rust in 120ms to 30ms: Python 🐍 to Rust 🦀🚀. Since then, we have moved the rest of our performance-sensitive stack.

Disclosure: I worked on the original C systems years ago. Now a new team rebuilt these systems in Rust.

Our core services were a mix of Python, Go, JVM, and C. We migrated the critical ones one service and one region at a time, running old and new implementations side by side until we had enough data to move the remaining traffic.

The rewrites cut latency and resource use. Memory behavior became easier to reason about, and the team prefers working in the new codebases.

Once latency became more stable, we could see details the old runtime noise had hidden. A 100µs excursion now sticks out and we can easily investigate changes of 30µs.

Most panels show a baseline and cutover. Panels 8 and 9 compare Rust with Rust. Panel 14 shows a rolling replacement; panels 4 and 13 show steady state only.

Reddit post gets only one image, so I tiled fifteen numbered panels together in the order discussed below.

The nginx replacement

Our load balancer decides where each message goes next. Its average time had sat at 600ms for so long that we had stopped questioning it.

We replaced the nginx-based balancer with Pingora, Cloudflare's Rust proxy framework. We kept the same box, traffic, and routing decisions.

Panel 1: Balancer time, 600ms to 101ms.

The drop starts around 09:40 as the Rust balancer takes over. We shifted traffic in two increments, which produced the brief ledge at ~300ms. Once the cutover finished, latency held at 101ms on the same hardware and traffic.

The right edge of panel 1 includes broker latency. That internal hop stays at ~350µs during the cutover. We had already upgraded the broker to Rust, so the balancer accounted for nearly all of the 600ms.

PubSub: lower publish latency and variance

Publish was already under a millisecond, so I didn't expect much room for improvement.

Panel 2: Publish latency, from a 400-500µs band to roughly 40-60µs.

The three series, average_publish, average_internal_publish, and average_signal, sat in the 400 / 500µs band, with a ~1ms excursion around 14:55.

A smaller bump reaches ~600µs at 15:10 as the cutover starts. Over the next few minutes, the band drops to roughly 40-60µs and stays there for the remaining twenty minutes. We watched the service for weeks before calling the migration complete. Neither the spikes nor the unexplained 1ms tail in customer p99 dashboards returned.

The lower variance mattered more to our latency guarantees than the lower average.

Measuring server hops in microseconds

The migration changed how I measure server-side latency. Microseconds are now a useful unit for these internal hops.

Panel 3 shows the same publish path over a day and a half.

Panel 3: Publish latency over 36 hours, ~350µs to ~50µs.

Green (average_internal_publish) fluctuates between 350µs and 460µs on the left, while orange and yellow (average_signal, average_publish) sit around 300 / 350µs. Around 10-22 08:00, green joins them near 300µs. Just before 10:00, all three fall into a 40 / 60µs band and stay there for the next 24 hours.

The hottest path went from ~350µs to ~50µs, with less variance. The entire chart stays under 500µs, and small regressions are easier to spot.

On the right, green rises to ~100 / 185µs around 10-23 00:00 / 04:00 while orange and yellow stay near 40-60µs. The bump affects internal publish only. A 100µs excursion disappeared inside the old stack's normal variance; here it is clear enough to investigate.

The replication layer runs at 27µs to 31µs.

Panel 4: Replication average latency, 27-31µs.

This steady-state panel shows average replication latency for two nodes, average_21 and average_23, across a vertical range of 27µs to 31µs.

Green moves between 28.5µs and 30.5µs. Yellow sits at 27.5 / 29.5µs and follows the traffic cycle. The ~1.5µs gap between the nodes is visible. Three years ago, GC pauses hid it.

GC did not account for all of the old latency. We owned the code and could have kept improving it. The first migration gave us enough evidence that a Rust rewrite would repay its cost.

We can now investigate whether the NIC, the scheduler, or our code makes one node trail another by 1.5µs. A 5µs improvement in this 30µs hop will show up.

At 3 trillion API calls a month, even single-digit-microsecond changes are measurable.

Presence: latency and memory

Presence tracks who is online and which channel they occupy across regions. It holds state for millions of concurrent occupants while heartbeats arrive, expire, and reconcile across the network. Presence had caused more incidents than any other service, so this was our riskiest migration.

Subscribe join latency, the presence online status event

Panel 5: Presence subscribe join latency (median), az1 from 1-3s to 200-250ms.

Green (az1) is on the old stack, moving between 1s and 3s with a 2.9s median spike at ~13:58. Orange and blue (az2, az4) are already on Rust and hold near ~350ms while az1 spikes.

At 14:03, we cut az1 over. It drops to ~200 / 250ms, below az2 and az4 at around 300ms because az1 received a later build. The other regions have since moved to that version and now match it.

Web heartbeat join latency

The web/HTTP path drops from 700ms to 200ms.

Panel 6: Presence web heartbeat join latency, 700ms to 200ms.

The blue refresh series fluctuates between 600ms and 800ms before the 10:30 cutover, then holds near ~200ms. join_announce (green) stays around 100 / 150ms before moving toward 200ms at the right edge. join_interval (yellow) was near zero and stops reporting around 10:25 because we removed that metric during the cutover.

The internal heartbeat fanout

Panel 7: Subscribe internal heartbeat latency, 1.5-2s to 500ms.

Three regions sit in the 1.4s / 1.8s range because they share an overloaded path. One spikes to 2.1s during the cutover. Green follows the same rhythm at a lower ~1s baseline. After the ~17:35 cutover, all four settle into a ~500 / 600ms band with less variance, about three times faster than before.

Presence memory

Panel 8: Presence pod memory, Rust vs Rust, a 3.4 GiB peak down to a 256-512 MiB band.

Panel 8 shows memory per pod in the presence namespace. Both halves are presence-rust deployments from two ReplicaSets, 6cdfbf75d5 and 6f78d4f459. This is a comparison between Rust designs, separate from the language migration.

In the old ReplicaSet, the top pod peaks at 3.4 GiB while the rest of the fleet ranges from 512 MiB to 2.5 GiB. Memory declines over more than a day as the service trims accumulated occupancy state. That build still carried substantial per-occupant overhead.

After the 04/23 15:00 / 18:00 gap, the new ReplicaSet holds within a ~256 / 512 MiB band. The growth curve disappears, and pods no longer approach an OOMKill during peak traffic.

Removing the GC made the profile easier to reason about, but we still had to fix the data layout. The first Rust build beat the Python service it replaced; the redesign cut peak memory by another 6x.

Push notifications: FCM

Push is a fan-out-and-wait workload. Apple and Google account for much of the latency, so our part needs to stay small and consistent.

Panel 9: FCM Rust average time, Rust vs Rust, from 135-250ms to roughly 135ms.

The single series, fcm rust avg, is the Rust path on both sides of the chart. The left side ranges from 135ms to 250ms and spikes at 08:45. Just after 09:00, it settles at ~135ms for the rest of the window.

At 09:00, we enabled connection pooling and reuse against FCM. The low runtime noise let us isolate the remaining variance and trace it to our code rather than Google's service.

That flat line is my favorite kind of graph.

Event Processing: fewer warnings and retries

Panel 10 counts requests that hit a degraded or retry-worthy path on internal_heartbeat, a high-volume route through event processing.

Panel 10: API warnings on the internal_heartbeat route, 1.5-4.9K down to under 250.

Green (az2) sustains 1.55K to 4.9K warnings throughout the afternoon. Yellow (az4), already migrated, usually stays between 100 and 500 with occasional excursions to ~860.

At 17:03, az2 drops into the same range as az4. Both spike at 17:20, when green reaches 1.02K and yellow ~740, then touch 400 / 600 a few times through 17:45. After ~17:50, both usually stay under ~250 with occasional spikes to ~400.

With roughly an order of magnitude fewer warnings and retries, alerts stand out. We no longer tune thresholds to ignore the service's normal behavior.

Memory and CPU

The migrations reduced memory and CPU use as well.

Memory, per shard

Panel 11: Shard 5 max memory, 95% down to a 20-40% band.

This panel covers shard 5 across all regions for two days. On the left, red sits between 90% and 97%, above the critical threshold, while yellow falls from ~90% to 85%. One traffic spike could have exhausted the remaining memory.

The cutover runs from 10-06 22:30 to after 10-07 00:00, when every series lands in the 30 / 40% band. Peak traffic on 10-07 between 12:00 and 16:00 pushes usage to ~69%. Overnight it falls to 12 / 25%, then returns to 20 / 37% with daytime traffic.

On the same boxes, peak memory moved from the critical range to below the 75% warning threshold.

Memory, per region

Panel 12: Max memory usage by region, with the step down at cutover.

The regional view shows the same step down on 06/22. The tooltip values come from after the cutover. Before it, the lines seesaw as the runtime allocates and collects memory.

After the cutover, the lines become thinner and smoother. iad remains highest at ~72% because this fleet still included unmigrated capacity; the other four regions had completed the rollout. We moved each region only after comparing it with capacity still running the old implementation.

(This iad fleet is now fully migrated. Other services there moved earlier; panels 8 and 10 both show iad clusters. We rolled out each service on its own schedule.)

CPU

Panel 13: Max CPU usage across regions, holding in the 0-40% band.

Across five days, regional CPU stays in the 0 / 40% band with occasional peaks at 60 / 70%. It never reaches the 75% warning or 90% critical lines. The sawtooth pattern follows normal daily traffic.

CPU per pod

Panel 14: CPU usage by pod, 0.5 cores down to 0.22 during a rolling replacement.

This panel shows per-pod CPU during a rolling replacement. The old pods use 0.4 / 0.6 cores before they drain and terminate.

The new pods stabilize at ~0.22 cores, about half the CPU for the same work. We used the savings to run fewer pods with more headroom.

Delivery semantics

We required the new system to match or improve our delivery guarantees. The delivery changes mattered more to us than the latency graphs.

For years, we replicated messages over a hand-rolled TCP protocol. It moved trillions of messages and supported retry and redelivery. It also left us responsible for framing, backpressure, reconnect logic, sequence numbers, and versioning between old and new nodes. Only a few engineers knew all of its failure modes.

We replaced the protocol with gRPC and a store-and-forward model.

Store and forward: persist first, stream second.

We replicate an incoming event to multiple physical nodes while beginning to stream it. We return delivery status only after the durable write. If the stream breaks or the downstream service is rolling, we query the replicas for messages that still need redelivery.

Streaming first leaves an unrecoverable gap if the process crashes between "sent" and "recorded." Persisting first may produce a duplicate delivery attempt. The record remains on the replica nodes, and an idempotency key handles the duplicate. The wire guarantee is at least once; the idempotency key makes delivery effectively once at the application layer. We do not claim exactly-once delivery over an unreliable network. For acknowledged messages, the durable replicas preserve a redelivery path, and the idempotency key prevents a repeated attempt from becoming an observable duplicate.

gRPC's automatic retry policy, configured at the service level.

The gRPC service config defines retryable status codes, attempt limits, and backoff. The channel handles retries below the application call, so many brief network failures no longer reach application code.

Previously, each service implemented its own retry policy. Conservative policies risked dropped messages; aggressive ones could overload a service during a brief failure.

gRPC absorbs transient failures such as momentary TCP resets. Store and forward handles failures that outlast the retry policy.

Rust's compile-time concurrency checks gave us more confidence in the new replication model. We ran both implementations side by side and watched the dashboards shown here.

Why Rust fit this workload

Rust suited this workload: a latency-sensitive message bus that holds millions of concurrent occupants in memory. These results do not mean every CRUD service needs a rewrite.

We spent years profiling and tuning the existing services. Managed runtimes can go far, but the collector still decides when to pause the program. The 1ms publish excursion and 2s heartbeat oscillation show the latency cost; the regional memory chart shows the allocation churn. The Rust services no longer incurred GC pauses.

We use Go for many API services. For this part of the message bus, we could not meet our latency and memory targets without removing GC pauses and reducing per-object overhead.

Our C codebase was 14 years old and needed a major update for current libraries and compilers. Since we faced a substantial rewrite either way, we chose Rust. We could have made some of the same performance improvements in C, and our measurements put Rust within the noise of C. The practical difference was that more engineers felt comfortable changing the concurrent Rust code, while the compiler caught ownership and data-race mistakes before deployment.

The first six months with the borrow checker were rough. Experienced engineers got frustrated, PRs stalled, and people questioned the migration. By month eight, review cycles had shortened and we spent less time debugging ownership mistakes after compilation.

cargo helped too. A shared toolchain, formatter, test runner, and build command save time in infrastructure spread across four languages.

Our Tokio mistake: unbounded in-flight work

Panel 15 shows the costliest mistake in the migration.

Panel 15: Memory usage by pod, from a ~100 MiB baseline to a 3.7 GiB burst.

Bursts arrived faster than we could complete them. We spawned a task per request, Tokio queued the tasks, and each one held buffers and state while waiting on I/O. Leak detection found nothing because the memory was live. Without an explicit cap, in-flight work grew with each burst. Waiting tasks use little CPU, but their state still consumes memory.

We first tuned the HPA with lower scale-up thresholds and faster reaction windows. It reacted sooner, but scaling out gave the unbounded backlog more places to accumulate.

We added backpressure at the ingest edge. A cap on in-flight work sends the burst to a bounded queue instead of letting tasks accumulate on the heap. Excess work waits or gets shed, so memory no longer grows with the arrival rate. We deployed the change in every region, and the same pods now hold their baseline under comparable bursts instead of climbing into GiB. I don't yet have a matching after panel for a comparable burst.

Two years ago, my Python-to-Rust post included mpsc::channel(100) and recommended Tokio MPSC for ingest paths. We failed to follow that advice here. Async can schedule large numbers of in-flight tasks at low CPU cost, but every task still holds state. Without a limit somewhere in the path, bursts consume the available memory.

Advice after three years

  • Instrument before you start. Every dashboard in this post existed before the migration, so we can support claims such as "600ms to 101ms."
  • Budget for the learning curve. This took three years, and the first six months were the most expensive.
  • Optimize for predictable latency. Customers notice the spikes more than the average.
  • Treat the first Rust release as a baseline. Panels 8 and 9 compare Rust with Rust: Presence peak memory fell about 6x, and connection pooling removed most FCM variance.
  • Work on durability alongside performance. We moved from our TCP protocol to gRPC and store and forward to improve correctness.

This is a follow-up to my 2024 Python-to-Rust post. I could not separate the benefits of Rust from the benefits of rewriting the architecture in that first migration. Panels 8 and 9 provide useful counterexamples: both compare Rust with Rust, and both show that data layout and connection reuse still mattered after the language migration.

Across these services, we now use less memory and about half the CPU per pod, and we see an order of magnitude fewer warnings on the route shown in panel 10. We also retired a wire protocol that only a few engineers understood. Those results justified the three-year migration for this workload.