r/rust • • 2d ago

🎙️ discussion Google is doing "Large Scale Codebase Migrations and Optimizations" of C/C++ to Rust

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
662 Upvotes

128 comments sorted by

223

u/Sirisian 2d ago

The relevant sections:

Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.

For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.

63

u/-p-e-w- 1d ago

If you had asked me 5 years ago when AI would be ready to do massive-scale unsupervised coding work on heavy-duty software at a company like Google, I would have guessed somewhere between the years 2070 and 2100.

44

u/Budget-Minimum6040 1d ago

massive-scale unsupervised coding work on heavy-duty software at a company like Google

"undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production."

19

u/Nothing_from_void 1d ago

It's a fancy way of saying "we're tasking a bunch of engineers to use our claude code-like tool to re-write some stuff to rust". I've been doing this kind of stuff at work for a while now, when you already have the code base and expected behavior it's one of the most straightforward use cases for AI

0

u/poelzi 13h ago

Yes and no. You have a known good test target, but your porting pipeline could easy endup with c-ism styled code just written in rust instead of rusty code. I'm building specialized nix flake experts in tix.im that know the domain and basically fuse classical expert systems with llms into deterministic workflows composed by nix. Works better then expected tbh

2

u/Nothing_from_void 12h ago

LLMs just prefer dynamic, procedural code in general, like cramming everything into bytes or strings and parsing out what's needed, in large procedures that run the algorithm step by step. I doubt re-writing from C++ to Rust is going to change that

0

u/poelzi 8h ago

Exactly the opposite. They prefer hard typed or side effect free. rust or nix they are quite good at. You need to create the proper harness / constitution for enforcing traits, new types etc. Works like a charm

22

u/-p-e-w- 1d ago

That review will mostly use AI as well (if indeed it happens at all). Do you think Google will allocate human engineers to manually review a re-implementation of 800k LoC in another language?

Just look at what major companies now do on GitHub. 25k LoC pull request, merged 20 minutes later with a single comment “LGTM!”, which is the “manual review”.

3

u/Nothing_from_void 1d ago

they are going to use antigravity CLI to write the code and say AI did the whole thing while the engineers are doing everything that keeps it on the rails

7

u/mss-anixe 1d ago

5y ago yes but 3 years ago it already looked likely.

That being said, while massive-scale unsupervised coding work is possible, I consider people who go this way very brave, lightly speaking.

Also, it seems that Google is not (yet) that insane:

such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.

1

u/exater 1d ago

So they write standard safe rust and avoid manual simd implementation through autovectorization?

How do they confirm its simd? They check the assembly that was compiled?

How come they dont just rewrite the C to be autovectorized?

14

u/Gaolaowai 1d ago

Because the C compiler won’t check for buffer overflows, lifetimes, use after frees, etc.

106

u/geo-ant 2d ago

Does that mean the whole Carbon thing is dead?

68

u/skeptic11 2d ago

Still experimental. Still getting commits.

How much C++ does Google have left?

59

u/geo-ant 2d ago

Good question, probably unfathomable amounts to my feeble mind. Also of course google is a megacorp with probably Byzantine structures so just because one part of the org does one thing doesn’t mean “google does x”. Still, it makes me wonder how useful the idea of Carbon still seems to google. I just don’t know.

17

u/2AMMetro 2d ago

Last I remember (> 1 year ago) it was a fucking lot. Like 80% if I had to throw a number out there. Carbon was nowhere close to being adopted for any major projects.

There are a lot of internal apps written in go, python, or java. But pretty much anything consumer facing and performance critical (ads, search, youtube, etc.) is in c++.

1

u/moltonel 20h ago

No matter how much C++ they have, the right question is do they prefer to * Maintain in C++ * Migrate to Carbon * Rewrite in Rust

The decision will vary by projects and teams, but Carbon's own docs state "use Rust/Go/Python if you can". Teams that can afford to do a large rewrite are not going to choose Carbon. It used to be that rewriting 100K lines was a no-go, but GenAI is changing that, making the RIIR option very appealing for a place like Google.

Carbon (and the various safer C++ endeavors) are arriving far too late.

15

u/arka2947 1d ago

To me it seems Rust is eating the future of Cpp. The safety improvements that are still debated in Cpp circles, are already done in rust. If Google needs that today, what else can you do than use Rust?

14

u/TDplay 1d ago

Not really. There are a lot of large C++ codebases. They still need maintaining, and they still offer utility for new code.

Rust has the benefit of being new. It could introduce all these safety features without worrying about compatibility with existing code. C++ does not have the luxury. If you introduce a safety feature that doesn't integrate well with existing code, then you would effectively split C++ into two languages: one with the safety feature, and one with all the legacy code. Effectively, C++ would be unchanged, and the safety feature would exist only in a dead-on-arrival language that would just be "Rust but worse".

That is why the safety improvements are so hotly debated in C++. They must be useful for legacy code, or else they are useless.

Rust has eaten C++'s future for greenfield projects. But C++ still has a long future ahead of it for maintenance of existing projects, as well as for new projects with extensive reliance upon existing C++ code.

15

u/steveklabnik1 rust 1d ago

This thread is about one of the largest users of C++ porting their brownfield code to Rust.

3

u/TDplay 1d ago

But as they say,

such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production

In effect, most of the work is still being done. I am sceptical that this process can practically be done for all the major C++ codebases, including bringing those rewrites up to the quality of the original codebases, within the next decade, so I believe C++ still has a future.

3

u/zxyzyxz 1d ago

I doubt there is really any actual manual auditing beyond reading some files and stamping LGTM on the PR. The automated testing is what's gonna actually catch bugs.

1

u/TDplay 1d ago

But can that alone assure the quality to the same standard as the existing mature codebase?

The best automated tests are those written in response to specific bugs, to assure that those bugs never return. But a rewrite will introduce a whole host of never-before-seen bugs, for which there will be no tests.

We won't know if this is truly successful until the rewrites are released to the public. There will undoubtedly be a massive influx in bugs, as there is for any other rewrite. With the lack of people familiar with the new codebase, fixing those bugs may prove unexpectedly difficult.

4

u/valarauca14 1d ago

But can that alone assure the quality to the same standard as the existing mature codebase?

Not to be a full fledge AI bro, but this has never mattered in a large enterprise environment. Shipping features before the end of the quarter has always been the benchmarks.

1

u/zxyzyxz 1d ago

Of course not but that's what will happen, same as the bun rewrite

6

u/matthieum [he/him] 1d ago

Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel.

I would imagine that the 800K+ lines Fuchsia Zircon kernel (C++) uses established C++ practices: callbacks, aliasing pointers, etc...

If this can be rewritten to idiomatic Rust, I can't think of much that couldn't.

3

u/Nothing_from_void 1d ago

llama.cpp, pytorch, pretty much all highly optimized AI code is written in C++ and has no plans to migrate to Rust. And also LLVM. Google has always been obsessed with safe code more than most companies, but there's still tons of C and C++ out there and people that like using it. They mentioned an 800k LoC codebase; that's not a huge amount honestly.

-3

u/pjmlp 22h ago

Well, both Rust compilers still depend on C++ for their own development.

That should be the very first place to cut ties with C++ dependencies in the Rust world.

Also there are lots of industries where C and C++ rule, there are standards using those languages, and Rust is hardly anywhere to be seen.

Gamedev, VFX, HPC, HFT, LLVM and GCC, GPGPU,....

Even if everything could be ported via AI, you still have the humans that might not want to use Rust anyway.

184

u/Toiling-Donkey 2d ago

I just love how they aren’t stubbornly migrating everything to Go…

91

u/SpoonLord57 2d ago edited 2d ago

isn’t Go meant to be more of a backend server language? they’ve always treated it more like a java replacement than a C/++ replacement edit: see below

77

u/DragonSnooz 2d ago

There was an interview where it was brought up that Go was created to fill a niche C/C++ nor Java were filling adequately. It was never intended to compete with either.

14

u/QuaternionsRoll 1d ago edited 1d ago

I mean, the first line of original design notes was literally

>Starting point: C, fix some obvious flaws, remove crud, add a few missing features

and they weren’t particularly eager to reject the notion of Go being “garbage-collected C” until long after it had found its niche at Google.

Arrays are an unmitigated disaster in true C fashion, I’ll give them that.

10

u/SpoonLord57 2d ago

That’s what I must be remembering, thanks for the correction

18

u/Equivalent_Head_4803 2d ago

Why would they do that?

27

u/projct 2d ago

because it's in-house

28

u/Ok-Opportunity-9731 2d ago

More of googles code base is c++ than go

14

u/klowny 2d ago

More of Google's codebase is Java than Go too. Go's niche at Google is mostly to replace their usage of Python.

5

u/Budget-Minimum6040 1d ago

The main goal of Go was to create a language for college graduates that can start creating output immediatly.

Rob Pike (main creator of Go):

The key point here is our programmers are Googlers, they’re not researchers. They’re typically, fairly young, fresh out of school, probably learned Java, maybe learned C or C++, probably learned Python. They’re not capable of understanding a brilliant language but we want to use them to build good software. So, the language that we give them has to be easy for them to understand and easy to adopt.

1

u/zxyzyxz 1d ago

I wonder what the point of Go is anymore when it's now AI writing the code and name people are now AI writing Rust without ever having actually learned Rust.

1

u/pjmlp 20h ago

That applies to any programming language, eventually AI will be good enough to generate machine code directly.

That is already a reality when using low code tooling like Boomi, Workato, Power Platform and co.

Agentic workflows, connecting prompts and MCP tools on diagrams.

1

u/Equivalent_Head_4803 1d ago

It’s incredibly fast, easy to read, and has amazing concurrency. In a world where someone craps out 5k lines in 20 mins, it’s pretty important that our code be readable. Every GO codebase is the same. AI isn’t going to invent a million different ways to do something when there’s pretty much one way or another few ways to do it. It’s a great language for AI and super easy to review IMO.

1

u/pjmlp 20h ago edited 20h ago

The main goal of Go was for their creators, two UNIX/Plan9/Inferno key figures, and one Oberon scholar, not having to deal with C++ at their new employer.

That came later, see Less is exponentially more regarding the adoption failure among C++ devs at Google.

2

u/Ok-Opportunity-9731 1d ago

Yes. C++ is about 40% Java/kotlin is 35% Go is 5%

1

u/Equivalent_Head_4803 1d ago

That logic doesn’t really track for me, but I know what you mean. GO is also open sourced and Google shepards it I guess? TBH I don’t really understand it’s open source situation.

It’s just solving a completely different problem than something like Rust, so I was wondering why the two would be considered against one another.

1

u/projct 7h ago

well yes, I don't agree with it either. it's just what the top level commenter's statement implied.

MS did a lot of this stubbornly migrating shit, same for oracle etc. NIH syndrome

114

u/TheRealMasonMac 2d ago

In my experience, I’ve found that agents produce code that works but is frankly dumb from a human perspective. Especially for larger tasks like this, you kind of have to rewrite it all. For example, I had a generalized design document for soft-wrapping explicitly support plain text and markdown, and used GPT-6-Astra to implement it. I used an ungodly amount of review agent rounds and interjections by me. When I actually go to review it, I wanted to cry. It silently pivoted towards implementing its own entirely separate wrapping mechanism rather than refactor the existing mechanism for plain text. I reviewed its code and found it to be trash, and I am now in the process of doing it myself… Never again. The short-term thrill of “having something that works” was outweighed by the fact that time was effectively wasted.

I mean, if the alternative was it never getting done, I guess agents are still better than nothing.

56

u/kabocha_ 2d ago edited 2d ago

Yep -- in my experience:

  • If you just want some throw-away thing to experiment on some idea, it's fine to let it go nuts on its own.
  • If you want to be able to maintain the results, you gotta check its work after every step or two, and you'll probably have to give it lots of feedback even then. Otherwise you're likely to end up with either a giant spaghetti mess, or FizzBuzz Enterprise.

No clue why it does that, I guess maybe a large portion of the training set is newbies dumping their first projects onto Github and super old legacy codebases, and not much in the middle, lol. They haven't learned to have good style taste, yet.

43

u/PM_ME_UR_BRAINSTORMS 2d ago

The problem I have with the latter is that it sucks lol. I don't enjoy micromanaging AI and in my experience it's not any faster than just doing it myself.

Every time I try using agents I almost immediately end up in a situation where it does something weird and just typing the code myself is faster than typing the prompt. And then I think "well I'm already here it will take me 2 seconds to type this other code" and so on and so on until I've written the whole thing.

Maybe it's just my ADHD but I don't understand how any one could work like that. Feels like pair programming with an intern which sucked to do but you did it because the point was to teach them.

24

u/notgreat 2d ago

I've found recent LLMs to be useful for finding things. You give it a vague description of the bug, and a couple minutes later it's got a 95% chance of having pinpointed the exact problem. Tell it to review a pull request, and often it finds real bugs. Not all its findings are good, but "review this list of eight possible problems" is a lot easier than "review these 500 changed lines of code for problems" which often just turns into "looks good to me".

They're also good at writing things with a lot of boilerplate gruntwork, though you still have to review their output. Like, you can describe a test and it'll usually do a good job implementing it. Just don't ask too vaguely, a general "test this area" will get terrible results as it locks in implementation details and leaves critical behavior untested.

2

u/PM_ME_UR_BRAINSTORMS 2d ago

Yeah that's how I use it most of the time and it's great. And those are all things no developer wanted to do anyway. No one was like "Yeah my favorite part of the job is hunting down an obscure error message on a decade old stackoverflow thread" and that's not some sort of skilled task where the effort produced better software. It's just grunt work.

Which is really going to fuck the industry in a few years because that means a lot less work for junior developers.

12

u/kabocha_ 2d ago

I agree 100%. Work really wants me to use AI, so typically I end up doing the latter, but it's not fun. I don't really use AI all that much for personal projects because I'd rather just write it myself.

11

u/PM_ME_UR_BRAINSTORMS 2d ago

It's not fun and the quality is just not good.

I use a reasonable amount of AI on my personal projects. I just treat it like an intern at my beck and call 24/7 and give it the kind of work I would trust an intern with. Which is something like maybe 10-20% of the codebase give or take. Basic boiler plate stuff, plus a ton of research tasks and code review and some help with debugging. All stuff that developers didn't really want to do anyway and all stuff that saves me a ton of time.

Idk why that's not enough for all these companies. I'm still way more productive, but I produce way higher quality code and I can do it with the $20/mo claude subscription. It seems like a win-win-win to me.

But it's clear that the companies just want to replace us. So it needs to be writing 100% of the code and doing 100% of the thinking, whether it's capable or not.

3

u/zxyzyxz 1d ago

AI is much faster than I can write by hand if only for the simple fact that they work in parallel not sequential as I can work on multiple tickets at projects at once. I literally have my personal laptop and work laptop on the desk and they're churning away.

2

u/forgot_semicolon 2d ago

Exactly. I have to deal with so much of it, and people don't appreciate that I can literally type the code faster than I can prompt it to write the code and then fix the inevitable mess -- or worse, re-prompt it to be how I would have written it in the first place

1

u/Zde-G 1d ago

or worse, re-prompt it to be how I would have written it in the first place

I think that's you problem. You want some certain code to be written — and try to make AI do that it. It simply doesn't work.

Just a question: how often have you been driving projects and asked other humans to write code for your from vague description?

There you would face the exact same issue, one human doesn't write like another human — and if you try to force it… hoo boy, are you in a world of pain.

Yes, AI works at less than human level now, but it's pretty close to “ask someone with entirely different habits to code something for you”, at this point.

Making it write code that works is not trivial, but doable, making it to write code that would like like a code written by you… don't even try.

2

u/IWasGettingThePaper 1d ago

Er, no. Some other humans can write good code. I've still never seen good code produced by an LLM without serious hand-holding. And yes, this is the case even on the latest models.

For small targeted coding tasks and mechanical boilerplate, it's great. For tracing through large codebases and summarising how some code path works - that I can't be bothered to do manually because jump to definition gets stuck at a function pointer, or some macro magic, or untyped python meta-nonsense, and I would have to start literally grepping - it's a real time saver.

The main trap with LLM code gen is you initially _think_ you're making progress at lightning speed. This is until you actually look at what it's produced in detail. Then comes the inevitable 12 hour prompt loop where you're trying to coax it into tidying up the slop its produced; rework the architecture to make structural sense; use meaningful variable/fn/struct names; remove the code it's added that doesn't do anything; rework the error handling so it makes actual sense; remove the unit tests it added for unit tests; rewrite the unit tests that don't test anything or rely too much on implementation detail; rewrite the impenetrable docs and comments it creates; remove the superfluous helper functions; refactor actually duplicated code blocks into functions; and on and on. Only then it dawns on you that it would've actually been quicker to write it by hand.

1

u/Zde-G 1d ago

Some other humans can write good code.

Yes, but these are rare and expensive. Average developers produce garbage similar to what LLM produces on the first try. Then they run tests and read the guides and adjust the code. You have to do the same with coding agent, for that to be apples-to-apples comparison.

And yes, this is the case even on the latest models.

The model is much less important than the quality of your documentation that explains precisely what you expect to see and precisely how to test what model creates.

LLMs rely on tests and formal documentation much more severely then even novice developer, that's not changing any time soon, you just have to accept that.

rework the architecture to make structural sense

If you allow LLM to develop architecture in code then it's precisely as expected. LLMs can not produce more than 100-200 line of anything without becoming weird.

That means that if you are writing something large enough to warrant an architecture then you have to ask LLM to first write plan - 100-200 lines of it, not more. Then implement said plan as steps that are 100-200 lines. If plan is too big to fit into 100-200 lines then first make it write an outline. And so on.

Don't expect model to produce large chunk and hope it'll be good, LLMs don't work like that.

remove the code it's added that doesn't do anything

Instruct it to use VCS (usually git) instead.

Again: the goal is to never allow LLM to produce more than 100-200 lines of code without verifying said code, somehow. Even if verifier is another LLM that looks on the plan or outline that you've made.

Only then it dawns on you that it would've actually been quicker to write it by hand.

Well, duh. You are trying to use rope as a pole… it doesn't work like that.

0

u/forgot_semicolon 1d ago

It's your last sentence that was the whole point of my comment. Getting it to write code that works (well, is readable, etc) takes longer than it does to just write the code properly the first time and not have to worry about its integrity later. Not saying no one makes small mistakes, but LLMs do some weird stuff that a human wouldn't.

I feel really silly any time I spend the same amount of keystrokes prompting for something than it would to write code. I mean, am I a dev or a pm? I feel like PMs wish they could just write features the way they want it, themselves -- I can! Why should I give that up AND pay for the privilege to do so?

1

u/Zde-G 1d ago

Not saying no one makes small mistakes, but LLMs do some weird stuff that a human wouldn't.

Sure. But then it runs tests and fixes them.

Most developers do that, some more than the others, LLMs need tests to even write very simple things, while humans may not need to do them.

Why should I give that up AND pay for the privilege to do so?

Because you are expected to deliver features 10x times faster, now. You either deliver them or are replaced, it's as simple as that.

If you have the luxury of writing code by hand — cherish it, pretty soon jobs where one may do that would disappear.

2

u/stumblinbear 2d ago

I do the latter with 2-3 agents at once while I'm away from my PC doing whatever the heck I want. Checking in on them to review and correct them every 20-30 minutes or so is fine

Something they do not tell you: you cannot just trust that the agent understands how to write good code. You need to lay ground rules. Just have it search up best practices and turn it into a skill that it must load before editing, reviewing, or writing code. You'll immediately start getting significantly better results. It's pretty much night and day

Current generation of agents are incredibly good instruction followers. If you've got it in a skill and written up properly, they do a pretty good job. Still needs corrections, just much less

0

u/PM_ME_UR_BRAINSTORMS 2d ago

Yeah idk I've tried tons of different skills and tools and guardrails and it still produces mostly slop for me outside of extremely basic boiler plate tasks that most 1st year CS students could probably handle. Which is still useful don't get me wrong. But not really worth having it touch like the majority of my codebase.

3

u/stumblinbear 2d ago

I usually use Opus 5.5 to draft up plans (I genuinely give it very little to go off of, maybe a few sentences), then correct a few assumptions or bad decisions it makes, it splits it up into smaller tasks, then it uses a sub agent to write the implementation. It runs a review, triages issues, beings design issues to me for making a decision on, then does the loop again. It typically takes a couple rounds

That said, that's mostly for working within an established codebase. For new feature work I usually have a longer discussion, let it explore different implementations in worktrees, then I pick and choose which parts I like out of them. Either that, or I already know exactly how I want it done, and telling it how to do it and letting it go off to the races works fine

It took me a couple of months to fine tune my workflow, skills, agent definitions, etc, but I can often go a couple of hours without checking in on them and usually I only need to make a few minor corrections. Shit, I had an agent going for the last 7 hours on a pretty large (though pretty mechanical) refactor and I didn't have to correct it at all

I was often getting trash output until I gave it permission to not worry about churn, to consider the correct fix not just the quick fix, and told it to surface issues it runs into instead of working around them

1

u/PM_ME_UR_BRAINSTORMS 2d ago

See this makes no sense to me. To me the plan and the implementation aren't like mutually exclusive things? Like what is software architecture if not the specific shape of the codebase.

Regardless of how good AI gets it can't read my mind. For the instructions I give it to be simpler and easier than just writing the code myself, it's going to have to make tons of assumptions for anything more complex than boiler plate. And that's where I always catch it producing slop.

And if I give it enough detail to for it to generate exactly what I want, well that's basically just a DSL and I might as well just have written the code.

Another user in another thread pointed out it basic just becomes this

2

u/stumblinbear 1d ago

I have a side project game in Bevy that I've been working on off-and-on for about 5-6 years for fun. I came across an article for how to better organize it, sent it to Claude and had it draft up a plan for reorganizing things under it. It asked me some questions regarding implementation decisions, then split it up into 7 steps, ordered by risk and in a way that helps untangle some longstanding dependency issues between features. This took it about 20 minutes while I fucked off doing whatever.

I looked over the plan, corrected one thing, then told it to go, and it went. It got through about 2 steps before it needed assistance on a design question that the plan didn't account for, so we spent about 4 turns and 5-10 minute discussing it before it got resolved and it continued. It's actively doing this work now as we speak. It would have taken me at least a week or two to do the reorganization myself. It's looking like it'll be done in a day and a half with only an hour or two worth of input from me.

At work, I came across a public API I can use to fetch some player data (I work at a gaming adjacent company) which we can use as a fallback when we're missing some info. I looked into the API myself, and it seemed like exactly what we needed, and I had a good idea of how to integrate it, but I essentially just told it "we want to use <link> as a fallback method for fetching player data when we're missing data for a player".

It looked into the API, looked at our codebase, and brought up that it actually has a full database available for download along with publishing hourly patches for the db. I had completely missed this, as it's not well-advertized on the site. It implemented support for this in about an hour.

Let me tell you, ain't no way I'm able to write a script to scrape this once per hour, along with setting up the data models and queries for the DB, along with hooking it into our own database to fill in the data, along with setting up the hourly trigger in Terraform, along with testing various scenarios... In a couple of hours. Hell naw. It was 2-3 hours from first idea to a pretty good implementation while I had other agents doing other things at the same time.

Giving it a precise spec is part of the problem. Modern LLMs are surprisingly smart, you just need to give them permission to have opinions and surface what they think is right or wrong while doing the implementation, otherwise they hack around issues instead of behaving like a proper engineer. Yes, they usually fuck something up, but those fixes are usually pretty minor or are style related and can be fixed in review. If not, it's because it made an assumption it shouldn't have, and that's... Frankly usually just a matter of updating the skill or agent definitions to prevent those issues happening in the future

1

u/PM_ME_UR_BRAINSTORMS 1d ago

Idk reorganizing some existing code and writing a script to hit an api and throw some data in a db doesn't sound all that impressive to me 🤷‍♀️ the latter especially seems like something a first year cs student could handle.

But to each their own man. If it works for you it works for you. I'm just sharing my experience.

2

u/stumblinbear 1d ago

You really overestimate first year cs students

But yes, it's relatively simple work. But it's work I don't want to waste time doing. I can handle other things while it does the grunt labor.

Besides, if you break down a problem enough, every problem is a set of simple changes that even a Junior could handle. Claude is capable of splitting up work into those smaller reviewable pieces, meaning it's very capable at doing large complex features with a bit of guidance and review. I just gave you two examples that were at the top of my mind because I literally did them that day

Small, reviewable pieces is how I got Claude to write a distributed rate limit coordinator for services that need to use a shared API key without blowing their rate limits. I didn't write a single line. It has been running strong for months without issue

2

u/mdedetrich 1d ago

No clue why it does that, I guess maybe a large portion of the training set is newbies dumping their first projects onto Github and super old legacy codebases, and not much in the middle, lol. They haven't learned to have good style taste, yet.

This is largely vibes based assessment (pun intended), but I noticed that the larger the models working set memory the higher changes of these things happening are (I use Claude myself).

For small and maybe medium problems, the AI is often very good. But once you start doing larger overarching things, its better to split the work into chunks, not just code chunks but also creating design documents.

The more the AI has to come up with myself the more it does weird stuff.

8

u/nicoburns 2d ago

I’ve found that agents produce code that works but is frankly dumb from a human perspective.

Agree that this can often be the case. For the specific case of translating between programming languages, the situation is often a lot better though: the original design is already available, the LLM just has to reproduce it.

I used an ungodly amount of review agent rounds and interjections by me.

This can actually make things significantly worse in my experience. Often the original version isn't quite right, but is relatively simple. And it adds all sort of of crap in response to review.

2

u/addmoreice 2d ago

I've had a lot of benefit from creating lots of llm generate unit tests and boundary tests based on the code structure and with emulation tests with comparisons to the original code.

I'm basically working on a binary -> rust conversion system right now as an experiment and doing this kind of thing has actually been really helpful. Combine that with what I'm calling 'root first and leaf first' conversion, where you can figure things out from reading from the entry points of exe's/dlls where we often have clear understood structure as well as from leaf functions that call into 3rd party code where we clearly understand the API, it's going very well.

I'm no where near running a big conversion experiment, but the small scale experiments I've done so far while developing the automatic workflow has been surprisingly effective. My eventual hope is to be able to automatically (90% or so) convert software where we have the binary, but no longer have the source code. This kind of thing would be great for preservation. I have absolutely no interest in subverting things like obfuscation so that's an actual non-feature I'm avoiding.

1

u/nicoburns 1d ago

FWIW, my experience of having LLMs write unit tests has not been good. They tend not to be good at generating assertions. They're good is the external environment provides constraints.

1

u/addmoreice 1d ago edited 1d ago

I think that part of the reason its been working so well for me is that everything I've been working on has been very local, very focused, and has lots of context 'around' the specific part I've focused the LLM on. Combine that with an incremental context expansion system and lots of 'here are the third-party API details that connect to this stuff, work backwards' and it has worked well. It's been a fun experiment so far, but I'll be soundly surprised if this keep scaling up as my experiment gets larger. I mean, I *want* it to work, and the research that exists is suggestive, but I'm still skeptical. Still, my LLM is hardly being fully used as is, so there is no reason not to throw lots of experimental hobby research at it and see what it pulls off. <shrug>

4

u/sreekanth850 2d ago edited 2d ago

Had opposite experience. But i never give agents full autonomy. Use compiler mcp and graft. and then implementation based on task ledger with small slices. Porting is one of the best use case for agents as you have clear reference and tests already written and have an upstream for references.

2

u/jl2352 20h ago

Google has done a number of papers on how they get AI working. The tl;dr is it has very little to do with the model. You don’t need a frontier model to have agents writing whole projects.

It’s all down to the harness, and how the AI is run. Google has invested heavily into this area, and that’s why it’s effective.

I’ve had great results myself using AI in very constrained ways to write code. That means breaking down a technical document into tickets, lots of skills, looking at the actions they took and changing skills to better optimise them, and having many guardrails.

1

u/ApokatastasisPanton 5h ago

Google has done a number of papers on how they get AI working.

do you have references for these?

5

u/heybart 2d ago

Yeah this is how I feel. I know I'm a pretty mid coder but I can tell good, elegant, well structured code when I see it, and LLM code ain't it. But for little apps that do a job and I don't need to ever touch it, it's great. Or when I just don't have time, if it works and passes the tests, I'm going to see no evil hear no evil it, because if I look at the code, chances are I'm going to feel not so good.

The incredible thing about LLMs is they're proficient or familiar with just about every language, framework, SDK, library, platform. They're superhuman in this respect. Yet they still sometimes do things that make you think "do you know how to code at all?"

What's going to happen when the models have sucked up all the good code and start training on other LLM's code? Not a pretty thought

4

u/stumblinbear 2d ago

Thing is, they don't know how to code at all. Not in your project. You have to tell them what style to follow, or they just do the quickest thing possible. Have them throw together a code architecture skill that lays out best practices, and you'll immediately start getting better results. You even specifically have to tell them to follow the style of the project, or they won't. Current models are great instruction followers, but also almost a blank slate

0

u/sosdoc 2d ago

If you look at all the “misalignment incidents” (huggingface and others) it’s pretty clear that the current models are heavily skewed towards reaching a goal by any means necessary.

It’s kind of a monkey’s paw situation, you have to be extremely specific about your wish, otherwise you get to suffer.

2

u/6d26d3af 2d ago

If you think of a scale where open ended "agentic" code is 0 and human handcrafted code as 100, not enough people think that you have actually control where you want it to end up. You can push it all the way to 80-90 by writing all the AST linters and checks and evals you can ever think of. It's still writing smelly code? Ban that pattern. Imagine today you can even have semantic checking for symbols (which is the bane of programmers) and weird heuristics unheard of. What a chance to be creative right?

You write these things once, reuse wherever, and never see the issues again. I don't understand why most people aren't yet picking up on this

15

u/nonotan 2d ago

You realize you're still allowed to just... write code, right? It's faster, cheaper, easier, more fun, and you'll get a better result. By the time you've written 20k lines of linters and checks that will still let garbage through because "automatically detect bad code" turns out to be a slightly harder problem than the actual problems you wanted to tackle with your software to begin with, I've already finished handwriting several entire libraries that Just Work. And actually had fun and learned something that will stick with me in the process.

Even in a best-case scenario, the end result of working endlessly to make an agent produce perfect, immaculate code for me going forward is... all I get to do for the rest of my career is typing orders in a chat box and reviewing the output. That's if they don't notice I miraculously managed to make it so good it barely needs reviews, and fire me because they don't need me anymore. Yeah, I'm going to not do that, thank you.

2

u/6d26d3af 1d ago

You're proving my point by thinking in binary scenarios. Who claimed immaculate code? Humans don't even write immaculate code as a baseline, what the hell is this false dichotomy

If you're still at the stage where you would rather write all the code yourself, nothing I say will change your mind and you'll have to wait until it clicks for you. If you read what I said carefully the ideal scenario is you reach about 80-90 on the aforementioned scale to eliminate almost all "bad code" while gaining extreme upside on speed. How do you close the gap to 100 then? Sure you could write more validators if you want, but you can also just let the LLM do it for you because the reality is they actually ARE good. The problem is preventing that drift vs your own ideals. Better future models will inch their way to close that gap objectively without any further context, but today this harness is a necessity at scale

3

u/funforgiven 1d ago

Writing the code yourself is definitely more fun, and you'll often learn more and probably end up with a result you understand better.

Unfortunately, I just don't think it's actually faster, cheaper, or easier anymore. It usually takes much longer, your time is probably the most expensive part of the process, and with an agent you don't have to personally think through every implementation detail.

That's kind of the depressing part. The thing that's more enjoyable and satisfying is increasingly not the most efficient way to get the work done.

2

u/ixid 1d ago

I think a lot of people, even coders, are still in denial about what you can make with LLMs.

1

u/pjmlp 20h ago

It suffices to be better than what I usually review in offshoring projects.

1

u/Professional_Top8485 2d ago

They do ugly code. They can also surprise.

I had project where i told to use design pattern and actually the result was better than i did. They just lack esthetic eye, but they can incorporate instructions how to do things.

I think they're getting better as well. Before llm broke the code base because it wasn't able to work on it and it always tried to fix (brake) thing I told not to touch. Now, with better model, no problem anymore.

Usually their coding style can be very rude and blunt but it's not always bad thing to have simple coding style. It doesn't abstract problems away.

1

u/yes_u_suckk 1d ago

I agree with you, but there are tools nowadays that try to minimize this, like spec driven development.

It's not a silver bullet, but if you review the spec before asking the agent to start working, the chances of drifting to something unexpected or bizarre are much lower than simply entering the prompt "convert everything to Rust, make no mistake".

1

u/ixid 1d ago

You can use the existing code as a byte identical output oracle requirement for your LLM port.

Do that, then a style and idiomatic pass and it's OK.

-6

u/teerre 2d ago

Agents don't come up with stuff on the spot, they do what you tell them to do. If you give an open ended problem, you'll get an open ended solution. This is compounded by the overwhelming return to the average these models are bound to and OAI / Anthropic insert layers and layers of indirection in their commercial offerings

It's a skill of its own to be able to design a collection of prompts, oracles and guidelines that achieves the result. Is that faster than doing it manually? As many things, it depends

7

u/TheRealMasonMac 2d ago edited 2d ago

In this case, the solution was defined. I had already solved the problem for it in math and essentially needed the model to translate it into code. I can only guess since reasoning isn’t exposed, but I suspect that they:

- “shadow-box with their own demons” as I like to call it by contemplating scenarios that can never happen and never validate the existence of. When reviews come in, they don’t completely backtrack from the idea and instead try to twist their original overly defensive idea to fit the problem.

- they really suck at thinking outside the box which I suppose is the inherent limitation of LLMs as they currently exist; they have to attend to their context per their training data

- “good code” is contextual; and that kind of training data is relatively scarce. But, one could also argue that models generally struggle with this even when the data is available. Code comments, for instance, often violate the cardinal “don’t repeat the code” rule (though they’re getting better).

I am also reminded of this article: https://joinhandshake.com/research/ai/deepswe-reward-hacking/

5

u/sparky8251 2d ago

LLMs also have no conception of time/consequences. Thats the biggest one... They rush to solve the IMMEDIATE problem, regardless of costs that it incurs later, and each pass does it again and again and it compounds over time to crazy bad code that no human would ever make...

Until LLMs can have this, I cant see them replacing humans any time soon... And my understanding is this is not an easy problem for LLMs specifically. Unsolvable is too strong a claim, but...

And then like, how do you even automatically test this stuff? These things need SO MUCH data to train off, need to run SO many times because they are so "dumb" in how they learn, how can you even score this stuff in a way that wont make it optimize off into insanity?

4

u/dark_bits 2d ago

The only way to describe business logic + make sure your prompt is not open ended is to actually pass the business context along with exactly the shape of the code you're looking to produce (and by shape I mean every method, class, namespace collection, file names, structure, etc). This would take way more time than actually writing it manually a lot of times.

1

u/Professional-You4950 1d ago

I completely disagree, even with extremely clear LOCALIZED problems and directions there are issues. They stem from context, scope, hallucinations, misunderstandings, and more. The longer you let it spin the more likely something catastrophic happens.

-1

u/teerre 1d ago

Prompts are just a beginning. Like I said, you need an oracle, a goal checker, an objective metric etc

1

u/oceantume_ 1d ago

In my experience it's possible to do serious engineering and refactoring but you need to be careful never to ask too much at once and you must be able to completely scrap branches and may even have to ask it eg to ignore that other branch entirely so it doesn't load it in context and use it as a reference.

It can be quite painful, but I'm managing to do very interesting projects with it rather quickly and it's hard to stop using it

1

u/arbv 2d ago

I don't know why you are being downvoted. You pretty much described my experience as well.

🤷‍♂️

-1

u/theanointedduck 2d ago

Bro! Im crying right now. Was supposed to ship 2 weeks ago, and after seeing the convoluted garbage that “passed all tests” i knew maintenance and cloud costs would be a nightmare.

It doesnt matter what model from Luna to Astra. It doesnt matter what guard rails, skills, etc. its lipstick on a pig

Im learning how complexity rears its ugly head and it’s so hard to define and for the model to stick with the plan.

0

u/Future_Natural_853 1d ago

The agent is merely supposed to spit the code out. You need a solid PRD, then you ask it to write technical plans that you review carefully, then you implement them one by one, then you review them.

0

u/BenchEmbarrassed7316 1d ago

In my experience, I’ve found that agents produce code that works but is frankly dumb from a human perspective

Something like this:

https://www.reddit.com/r/rustjerk/comments/1wusj30/i_think_this_guy_is_lying_ai_isnt_capable_of/

?)

0

u/mss-anixe 1d ago

A LOT depends on many factors. The model, the project, personal experience and luck.

5.6 Sol before lobotomy produced for me an amazingly good piece of rust code that I would probably never invented on my own (not only did I fail to think of that solution, but even if I had been capable of doing so, I wouldn't have invested enough time to figure it out).

5.6 Luna produced small-to-smalish pieces of code that made a lot of sense.

Current gpt models, including (nerfed) 5.6 and 6.x, need again prompt engineering or multiple attempts to produce good quality solutions but are still viable for some kind of brainstorming.

That being said, so far I never managed to create a big project (say 5000+ lines) with ai that I would not be afraid to use for anything else than a prototype which is later re-written.

-20

u/mguerrette 2d ago

Code is no longer for human eyes. Did you benchmark the generated code? Or was it just trash in your opinion (style, architecture, etc)? I think you are missing the reality change where code is no longer for humans.

9

u/TheRealMasonMac 2d ago edited 2d ago

It was egregiously horrible stylistically, architecturally, maintenance-wise, and in performance. The stupidest one that came to mind was it allocating strings for each markdown element. That's like &str 101. I wish I could tell you there was a good reason for this, but there wasn't. So, it ended up doubling memory usage in string allocations alone. There were also a bunch of other areas where it turned linear time functions into exponential ones.

Also: agents do need to read the code as well. Ignoring context rot, the matter of garbage in -> garbage out remains.

2

u/sparky8251 2d ago

Ive had it make a parser that took something like 4 minutes to parse a few kb of files because it refused to surface a bad assumption I made in the design that a library forced on us. It was also a degenerate but rather common case of scope drastically, drastically ballooning parse time so all its tests showed up as fast, because they were small scoped.

Its fun for exploring stuff with, but MAN would i not trust it...

And I DID use all kinds of clippy features to ban functions, macros, etc with reasons as to why they were banned to try and make it do better. Its just too concerned with the now and quickly doing what it thinks it was asked, no matter how it has to contort itself or the code to make that happen...

2

u/stumblinbear 2d ago

Oddly enough, if you just tell it to consider the long term consequences as a matter of course, it improves its decision making. It's a blank slate, it doesn't understand a project's priorities if you don't tell it what the priorities are

2

u/arbv 1d ago

Yeah, I am accumulating failure modes to write them down for the models concisely, that is one of them.

1

u/arbv 2d ago

Did not you see it:

  • Write code
  • Write tests
  • Realise the code is wrong
  • Fail to fix it
  • Rewrite the tests because "they are wrong"

?

I did and I didn't like it.

3

u/sparky8251 1d ago edited 1d ago

Ive had it do that too, yes. Had to be watchful to make sure it didnt make the tests pass to claim victory... It misunderstands the point of tests pathologically, because its an optimizer optimized to solve whatever it decides is its task, and that often means test fixing is a very well rewarded shortcut to its goal.

-6

u/mguerrette 2d ago

Maybe it’s an issue with Rust maturity then. I can say for C and C++ there is a rich set of tools for constraints both stylistic and to conform to best practices. Agents generally can use these and generate very solid code for those languages in both perf and style. I find myself not needing to look at any of the code as long as the format and tidy passes come back clean.

3

u/puttak 2d ago

Does it able to replace unnecessary std::string with std::string_view? Best practices does not enforced on the logic since it is up to programmer.

1

u/mguerrette 1d ago

Yes it does, since it can read language quite well it can read the output of clang tidy which instructs for this among other things. Seems like you’ve just never used AI at all based on your comment

1

u/puttak 1d ago

Sorry I mean does it know how to avoid unnecessary allocation? I did not mean just changing parameter from std::string or return value to std::string_view.

Seems like you’ve just never used AI at all based on your comment

Unfortunately your assumption is wrong. I review a large amount of code generated by Fable 5 and Opus 5 everyday and I known its strength and its weakness. If you think AI generate an efficient code then it is likely you don't know how to make that code more efficient.

34

u/LevelKnee1556 2d ago

One PR.

30k lines of changes.

“LGTM”

6

u/Xiphoseer 2d ago

Having the Fuchsia Kernel be Rust too is a crazy turn of events

7

u/matthieum [he/him] 1d ago

Is it complete? The article is using the present continuous form, leading me to think it's ongoing work.

But yes, a Rust micro kernel deployed on every Android device will be a hell of a flex.

19

u/Zalenka 2d ago

It's crazy how well AI does porting from one thing to another. I've been writing native iOS and Android apps targeting a rust ble backend. I'll write, generate, and debug the iOS code and then have AI port it over to Android and although stylistically it still needs some hand-holding it's surprisingly comprehensive.

1

u/xAmorphous 2d ago

Have you found that it ends up writing an Android analogue of the iOS code? I've never developed for mobile, but I've found that even the best models try to "port" code over by keeping the general structure of the existing code, which makes a migration from something like python to rust really janky.

1

u/Zalenka 2d ago

It mostly follows the patterns already existing but they are similar since the original writer took it from my iOS from the beginning anyway (MVC).

The Android app does use KMP for the network layer which was mostly a bust since it was originally for code reuse but I wrote the iOS app so much faster than the Android app and didn't want to rework just to deal with those horrendous stack traces on iOS from KMP.

1

u/Fancyness 1d ago

We live in great times, computers never were as useful as today!

5

u/EmperorOfCanada 1d ago

Rust boils down to a very simple pair of wins:

  • It covers all the platforms I use. Embedded, mobile, server, desktop, wasm. C/C++ is pretty much the only other major language which does this properly.

  • It works. I don't mean it vaguely solves the problems like other languages, but the code I make with rust is rock solid. Granite rock solid. It runs, it runs and runs and runs. No segfaults, no weird threading crapouts, no WTF buffer crap.

If there are any mistakes, they are ones which I would make in every language out there, logic issues, spelling mistakes in GUIs, etc.

1

u/astar0n 1d ago

Is there way where I can apply at Google specifically to work on Rust ?

1

u/InternationalFee3911 1d ago

I thought you linked the wrong article! On a closer look, there are only two paragraphs about Rust. The first says what you cite.

The second is something about an apparently bad hand-rolled Rust with much slow SIMD. This LLM helped clean that up, after which it only came closer in performance to C++.

1

u/DavidXkL 1d ago

I only use LLMs for finding information 😂

I still code myself lol

0

u/Capable_Belt1854 22h ago

What about large "Donations" and "Contributions" to Rust? No? Ok, then.

3

u/andreicodes 17h ago

Google is a Foundation member, so they pay their fees. Plus from time to time they give out extra grants to Foundation to work on specific things like C++ interoperability.

-10

u/dashingThroughSnow12 2d ago

Now sure if still true but there was a common adage that Google would rewrite every line of code every two years (obviously not all at once). For related reasons, this is why the Go programming language was invented.

With that in mind, this sounds like status quo.

-16

u/zackel_flac 1d ago

In the age of LLMs, Rust seems superfluous really. The safety net was supposed to be for human beings, not machines. Rust comes at a cost, in terms of perf, verbosity and build infra.