r/devops • • 5d ago

Weekly Self Promotion Thread

21 Upvotes

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!


r/devops • • 17h ago

Discussion Am I cursed? Is this the end? Where did the clients go?

210 Upvotes

I run small DevOps/Could Management service company with 6 engineers. The team is quite experienced and I myself have more than 14 years between enterprise and startups. We built multi-cluster multi-regional deployments, automated and gated CI/CD pipelines, helped pass compliance audits, configured useful monitoring and response workflows, etc.

In the previous two years things were going well, I was able to find steady work for the team, even more than we could cover at times so we were growing. My channels were personal network, linkedin, upwork and even our very seo-unoptimized website. Our customers were mostly middle-sized companies and VC-backed (or otherwise financed) fast-scaling startups.

There's no work anymore. Instead of closing several deals per month in 2024 and 2025, this year I was able to get two new customers back in March. That's it, nothing ever since. We still had ongoing tasks but in Summer that dried up as well. My company bank account is empty and I have to let the team go.

I tried to contact our customer from the previous years and most of them have laid off a significant portion of their teams and scaled down infrastructure. Even people who were supposed to support and monitor the infrastructure we implemented were fired without replacement in many cases.

Is anyone else experiencing this? I understand AI-led changes in the industry, but it's like a switch turned off some time in early Spring. What am I doing wrong?


r/devops • • 18h ago

Ops / Incidents When does self service infrastructure become too much self service?

22 Upvotes

We pushed pretty hard to let dev teams handle more of their own infrastructure changes.

It worked, but now the platform team spends a lot more time reviewing Terraform, fixing edge cases, and explaining context that isn't obvious from the repo.

At some point it feels like the bottleneck just moved instead of disappearing.

Has anyone found a good balance between developer ownership and keeping infra changes sane?


r/devops • • 21h ago

Discussion Do self taught engineers still exist?

26 Upvotes

Or whatever you call yourselves? Since the LLM/agentic explosion to today - how many of you got to where you are purely through teaching yourselves, whether with/without certifications? How much of your experience was taught on the job post-2020?

Who in here got in the workforce in the last 6 years and had supervisors willing enough to train you on the job? How much extra hours did you have to put in to get up to speed?

I know it's a combination of projects, networking, certs (dependent on the company/your hiring manager), going to industry events, layering on fundamentals, breaking my projects on purpose, cold emails, cold Linkedin connects, Discord chats, surrounding yourself with people more knowledgeable. Feels like I've tried everything

EDIT: A lot of very insightful replies, thank you all. I do enjoy a chat and it's hard to get that organically through LinkedIn and the works


r/devops • • 22h ago

Tools What net cost-savings have you realized from cloud-agnostic infrastructure?

6 Upvotes

Many engineers feel wary of relying too much on one cloud computing vendor, and I understand why in principle. Using multiple vendors makes a system/application less likely to completely fail if one vendor gets compromised. Using cloud-agnostic tech also can also support switching platforms based on cost or standardized skillsets across teams. Furthermore even the most mature vendors have global single points of failure though they are not all well-known.

That said, the tradeoff is not free even if the software is. Using cloud vendors' proprietary tools allows smoother integrations with their other products and can reduce opportunity cost in cases that are time-sensitive. Mature cloud computing platforms also provide enough isolation and redundancy on their own that most DR/availability concerns can be addressed without separate cloud platforms. On the flip side, using non-proprietary tools insources significant security/compliance/maintenance responsibility and risk - all things that contribute to cost.

I'm not completely sold either way; I tend to be a pragmatist and feel that each fits its own niche. However the choice can sometimes be unclear.

If you use cloud-agnostic tooling, then have you observed evidence that it supported a real net cost-savings accounting for both cloud spend and labor vs the cloud vendor's alternative? If you observed cost-savings, then what about the circumstances pushed it into cost-savings territory?

I'm intentionally being open-ended here. The first tools which come to mind are Kubernetes and Terraform/OpenTofu; then apps like ELK, Grafana, Prometheus, Jenkins, the Apache family of software, etc. etc.


r/devops • • 1d ago

Discussion Do you guys feel lucky to be in devops or to have switched to devops ?

74 Upvotes

With uncertainty in tech , layoffs, advancement in llms . Do you feel lucky to be in devops which is somewhat AI resilient field compared to other tech jobs ?


r/devops • • 15h ago

Troubleshooting Forgejo runner needs to apply configuration on host as root

0 Upvotes

Hi all,

What is the best way to having a Forgejo runner applying configuration on a host, which requires root access?

I have a Forgejo runner running directly on the host. I have tried to use sudo, having forgejo-runner ALL=(ALL) NOPASSWD: ALL in my sudoers, but that user is not configured to run sudo:

sudo: Account expired or PAM config lacks an "account" section for sudo, contact your system administrator sudo: a password is required


r/devops • • 2d ago

Discussion The flood of "part-time remote US DevOps contracts" is an interview proxy & identity theft scam

52 Upvotes

Most of job post for DevOps on reddit are fake, If anyone asks you to interview under someone else's name, run.

If you ask for details, the real scheme comes out:

  • The ask: You jump on Zoom calls, fake an accent, and clear technical rounds pretending to be a US-based guy.
  • The deal: Split the paycheck 60/40 while they use their US citizenship to pass the background check, and you do all the actual work offshore.
  • The catch: It’s straight-up identity and wire fraud. There's zero contract, so when they ghost you on payday, you get nothing. Plus, hiring platforms are heavily flagging biometric and IP mismatches now.

Anyone else seeing an influx of these messages recently? How is your team filtering them out?


r/devops • • 20h ago

Ops / Incidents What do you think of OneUptime?

Thumbnail
github.com
0 Upvotes

I am curious to know if you tried it in a real production scenario and/or you found better cloud alternatives


r/devops • • 1d ago

Discussion Experience with XAML Builds in Azure DevOps 2020 on prem Upgrade to latest Azure DevOps Version

3 Upvotes

Hello everyone,

I am planning to upgrade our current DevOps environment to the latest Version (Upgrade path is given by Microsoft Docs).

Unfortunately, we don't have a second environment to test the upgrade so maybe someone of you has some experience.

What we do have is that the Machine where our Azure DevOps instance is running is fully virtualized. We can do snapshots and restore. (HyperV)

Our instance unfortunately still uses XAML Builds for an legacy application that is still used by a lot of customers and is Business critical. These are sre not easy to migrate and are highly intertwined with business logic and some other dependecies. In Short: A Nightmare. We also use Classic Pipelines and Releases.

I know that XAML Builds are deprecated, but i inherited the Environment and the old engineer that build the whole Thing is not in the company anymore. So it is still kind of a "blackbox".

We could just snapshot the Environment before the upgrade and rollback. Never done it, but i think it should work if shit would hit the fan.

Does anyone have experience with Upgrading to a newer Version and XAML Builds?

What would be the best approach here?


r/devops • • 22h ago

Discussion 7 years of running Linux servers (from 2019 shared hosting to bare-metal microVMs today) , 3 rules I never break anymore

0 Upvotes

Started in 2019 keeping shared hosting boxes alive at 3am. Today I run bare-metal infrastructure with Firecracker microVMs. Three lessons learned the hard way:

1-> Containers aren't VMs: A Docker container is just a Linux process sharing the host kernel. For untrusted code or AI agents, dedicated kernel per workload > shared kernel + prayers.

2-> Public IPv4 on boot is a trap: Bots hit port 22 in under 60 seconds. Keep outbound open and inbound dark by default until a domain or IP allowlist is explicitly mapped.

3-> 1:1 reserved RAM beats burst specs: Half of "random" 2am database OOMs on budget clouds are just noisy neighbors fighting over oversold host RAM.

What’s one infra rule you learned the hard way?


r/devops • • 1d ago

Discussion What is the best way to learn devops quickly? (Best free course)

0 Upvotes

Hello, I'm a junior sysadmin and I'm planning on upgrading to devops.
I'm trying to learn by myself but it's proven quite hard and time consuming, so I'm wondering is there like a really great free course on youtube or online that is up-to-date.

Thank you for your time!


r/devops • • 1d ago

Architecture AI Comisseration Follow up

3 Upvotes

I wanted to write a follow up on this post I made a few days ago: https://www.reddit.com/r/devops/comments/1wb3l8j/ai_comiseration_client_replacing_production/

So to summarize: I was more or less asked to review a portal implementation that was made using Claude and vibe coding by someone who doesn't know how to develop websites. I suspected it wouldn't end well and even on the surface found lots of issues.

So here's where I'm at now: I created a report doing quite a bit of analysis. Yes I used some AI tools to try and comb through, and I made sure that any claim I made based on its finding I dug in to looking at the actual code and outputs. It made things more manageable but there's still way too much garbage to look through.

My conclusion is this (and it's obvious I think to most here) LLMs need guardrails and lots of them or at least clear ones. I do find that it's tough to actually communicate these low level issues that have a clear (to me) underlying architectural issue and a human issue (the human doesn't know what their doing or knows how to validate output beyond the surface level), when it takes one minute to go "please fix the list of findings" and Claude goes "it's fixed."

Many mentioned that this thing is just going to fail, and it is. I'm mildly concerned about the fix being "just fix problem x" and then they move on until the next thing blows up. And there is a real lack of concern around impact. "This is a public website so it's fine if we have keys in our code and expose endpoints only secured with a SAS key (which is plain text in JavaScript)."

I'm curious how many of you have run in to this even at an operational level what pains are you feeling and have you been able to deal with them? I do get what the solution is, create a proper architecture and define guardrails for Claude to follow, iterate over that until the LLM generates outputs in a way that is in-line with said architecture. I think LLMs may work (though whether they are as cost effective I press X to doubt), and if forced to make it work: this is probably the way. So are you guys maybe managing some guidance markdown files and what are the things you learned from that?

Finally, I think LLMs suck for this purpose and it's clear. I knew this but now I'm living it. Which is nice and validating, but also it's tough to rely on my tried and true "look mayor companies have been doing x for years you are not special you should follow the standards" because LLMs in this space are so new and there isn't much precedence.


r/devops • • 1d ago

Discussion When someone leaves, how do you actually revoke their access to servers and databases?

0 Upvotes

Curious how teams handle this in practice. When a dev or contractor leaves, there are SSH keys in authorized_keyson a dozen boxes, a DB password everyone knew, staging .env files in Slack DMs, maybe AWS keys on their laptop.

  • Do you rotate everything they could have touched, or just disable SSO and hope?
  • Is there a checklist, or is it tribal knowledge?
  • What tool (if any) made this not painful?

Asking because we got burned by this at a small team and I'm trying to figure out what "good" looks like without a dedicated security person.


r/devops • • 2d ago

Observability Nobody Is Listening on Port 8125

Thumbnail
yeet.cx
3 Upvotes

Hey all! I wrote this write up talking about how I was able to implement eBPF technology in order to re-implement StatsD Exporter on my local observability stack.

I thought it was interesting exploring how I could find the same information available in the kernel and preserve the fire and forget behavior without needing the running port or userspace overhead. Hope you enjoy!


r/devops • • 3d ago

Discussion Is it just me or Github actions is overrated?

289 Upvotes

We're adopting github actions at work and... they look clunky, bloated and overcomplicated?

I've used both Jenkins in the past and Gitlab CI/CD (and bare shell scripts in a past life).

It seems to me that gitlab-ci was the pinnacle of code-driven ci/cd (despite having some sharp edges).

Also, running github runners in kubernets is quite a fight. Gitlab's runner was so easy and simple to run.

Am I missing something ?


r/devops • • 3d ago

Discussion If AI makes everyone a 10x developer… who gets promoted?

112 Upvotes

Random thought — if pretty much every developer is using AI for coding now, how does career progression work?

Like, what makes one dev stand out from another?

How does a manager decide, “Yeah, this person is ready for the next level,” when everyone has access to the same AI tools?

Genuinely curious how people see this playing out.


r/devops • • 3d ago

Security Critical RCE Alert: Full takeover of HashiCorp Vault and OpenBao. OpenBao is patched. Vault remains exposed

170 Upvotes

https://control-plane.io/posts/unauthed-to-rce-in-vault-and-openbao/

OpenBao engineers at ControlPlane have chained 4 vulnerabilities to show how under certain conditions, an OpenBao or Vault server can be completely compromised from an unauthenticated position. This is only the second RCE ever found in the Vault codebase.

The exploit is highly plausible in real-world environments, requiring only an unauthenticated entry path and a defined Raft snapshot policy to trigger a complete server compromise.

If you are impacted, upgrade as soon as possible to OpenBao 2.6.3 or 2.7.0

While OpenBao is fully patched, HashiCorp Vault remains exposed as of writing. Unfortunately, IBM's unwillingness to coordinate a mutual disclosure policy means Vault users currently lack an official mitigation


r/devops • • 2d ago

AI content Do your clients or employer require AI use disclosures?

3 Upvotes

I work in a position where what and how I use AI for coding/devops is heavily constrained by regulations and policies, and where AI use disclosures may soon be required. I'm curious how common it is in the industry now for AI use disclosures to be required vs voluntary vs not discussed.

Does your workplace discuss AI use disclosures, and if so are they voluntary or mandatory? Also, what do disclosure requirements look like for integrating AI into live devops tooling vs using it for development only?


r/devops • • 2d ago

Troubleshooting How to set a "env" var on github actions using shell script?

2 Upvotes

i need to get the app version and i have this awk command for that:

` awk -F'"' '/^version[[:space:]]*=/ {print $2; exit}' pyproject.toml`

but how do i set the variable to the output of this command?


r/devops • • 2d ago

Tools LambdaTest/TestMu - Is KaneAI Worth Your Time?

2 Upvotes

Has anyone out there used TestMu's KaneAI? I do not see a lot of conversations out there on it. I wanted to use it for non-technical users to be able to generate tests without guidance. However, even when being as descriptive as possible in the prompt, I see test generation often failing or failing during the hyper execute phase. Thoughts? Experiences? Success stories anyone?


r/devops • • 3d ago

Ops / Incidents Has AI made application development faster than infrastructure teams can safely support?

14 Upvotes

Something I've been wondering about lately:

Developers can now generate features, tests and entire PRs much faster with coding agents.

But infrastructure hasn't suddenly become easier to reason about.

Terraform, networking, IAM, Kubernetes, observability, cost and security still require quite a bit of context.

So instead of AI removing bottlenecks, are we just moving the bottleneck downstream to DevOps/platform teams?

Especially interested in teams where developers are now generating their own Terraform/Kubernetes changes with AI.

Has your infra PR volume increased?

And more importantly: has the quality increased with it?


r/devops • • 2d ago

Tools Need easy to follow education(al materials)

0 Upvotes

Devops people of Reddit, I could use some pointers towards some good quality learning materials that cover the basics of devops.

I'm not trying to become a devops, I'm one of those pesky PM types and am supposed to lead some devops/infra projects soon which - from a PM perspective is fine but I hate sitting in meetings where people are speaking Chinese at me so Ive decided to learn something. Not enough to contribute constructively but enough to have a general idea of whats being said without wasting everyone's time and need some help.

Any (preferably video) resources you can recommend that cover basics in a digestible format, preferably someone who knows how to transfer knowledge and not just confuse me further.


r/devops • • 3d ago

Tools Ran AskUI in our pipeline for three months, notes on getting desktop UI tests working in CI

4 Upvotes

Posting because getting GUI tests to run headless is the part nobody talks about. everything passes on a dev machine and then does nothing on the runner, because the VM boots to a logon screen and there's no interactive session for input injection to land in.

Their runtime installs as a Windows service, which is what got us past that. Runs as SYSTEM, survives RDP disconnects, works at the logon screen. AskUI run for the headless job. tests are plain markdown files so they version with the app, and secrets get referenced by name and stay out of the logs.

Two honest caveats. runs cost money, so "run everything on every commit" stopped being a free decision and we moved the full suite to nightly. And their WARN status is a judgement call by the agent, which is fine when a human is watching and less fine at 3am when it just lands in a report nobody reads until standup lol.

Hope this helped :)

Anyone else running agent based UI tests in CI? curious how you handle the nondeterminism in a gate.


r/devops • • 2d ago

Architecture The Lunch Incident: A Short Novel About Accidental Architecture

0 Upvotes

Everything here is based on real experiences of a single application, but composed into a story by AI. All company/individual names have been replaced by fictional names. And yes, it was exactly as bad as the story makes it seem. None of these are made up. Reposted now that I have enough subreddit karma to actually post here.

The Lunch Incident

A Short Novel About Accidental Architecture

Chapter One: The System Works

On Eli’s second week at Meridian Therapy Systems, someone told him not to worry about the billing application.

“It’s old,” his manager said, “but customers love it.”

This was true.

CareLedger—an acquired application used by thousands of small therapy practices—was fast enough, familiar, and remarkably well suited to the people who used it. Clinicians could schedule patients without thinking like programmers. Billers could enter payments quickly. Small practices could submit claims without buying an enterprise revenue-cycle system.

It was also, Eli soon discovered, held together by a series of engineering decisions that would have been rejected from a freshman software-design project.

The original developer had understood users exceptionally well.

Computers, somewhat less so.

Every practice had its own database.

Not its own schema.

Its own database.

Thousands of them.

Eli stared at the architecture diagram.

“Why?”

“Isolation,” said an engineer.

“That’s actually not terrible.”

“No,” the engineer said. “Keep looking.”

There was a pause.

“Oh.”

The web application had administrative privileges over all of them.

Chapter Two: Lunch

The incident arrived through Support.

A clinic reported that a large portion of its appointment schedule had been destroyed.

The appointments themselves had not disappeared.

That would almost have been easier.

Instead, real patient appointment data had been overwritten by a recurring calendar event.

Its name was:

Lunch

Names, details, and other appointment information that clinicians relied upon had been replaced. Where the clinic expected to find its patient schedule, CareLedger confidently informed them that they were having lunch.

Engineering reconstructed what had happened.

A clinician had created an ordinary schedule block called Lunch.

Somewhere beneath the friendly calendar interface, CareLedger had decided that the clinician’s request applied to rather more appointments than the clinician intended.

There was, fortunately, tenant isolation.

The bug did not turn every appointment in the company into Lunch.

Only a substantial portion of one clinic’s schedule.

The clinic also had something CareLedger did not expect:

paper.

Staff reconstructed enough of the schedule from printed records and other offline information to continue seeing patients while engineering investigated.

And Meridian had database snapshots.

The damaged appointment data could therefore be restored rather than reconstructed entirely by hand.

This fact was discussed as good news.

From that day forward, engineers referred to it simply as the Lunch Incident.

Nobody needed clarification.

Chapter Three: The Free Customer

Meridian charged practices monthly for using CareLedger.

At least, that was the idea.

One practice filed an unusually large number of bug reports. Support knew them well. Engineering knew them better.

One day someone noticed something odd.

They weren’t paying.

Not because their card had failed.

Not because their account had been comped.

CareLedger simply wasn’t charging them.

The billing system worked by scheduling a future event.

When that event ran, it charged the customer.

Then—and only then—it scheduled the next billing event.

Eli stared at the code.

“So if the event doesn’t run…”

“There’s no next event.”

“And something later notices they’re overdue?”

“No.”

“There’s a reconciliation job?”

“No.”

“Accounts receivable?”

“Not really.”

“So if the server is down when the event fires…”

“They stop getting billed.”

“For how long?”

“Forever.”

This had happened to more than one customer.

Deployments could cause it.

Outages could cause it.

Anything capable of preventing the event from firing could permanently transform a paying customer into a free customer.

The billing application had implemented recurring revenue as a chain letter.

Break the chain and the customer escaped.

Chapter Four: The Faster Server

Eventually Meridian migrated CareLedger to newer servers.

No application code changed.

No database code changed.

No UI code changed.

Soon afterward, clinicians began opening appointments and finding large sections of patient information mysteriously blank.

There was no deletion audit.

There was no obvious failed update.

The data seemed simply to disappear.

Engineering searched for days.

Then Eli asked a question that sounded ridiculous.

“What if the new server is too fast?”

Silence.

Two operations were occurring without explicit sequencing. On the old server, one was reliably slow enough that the other always completed first.

Nobody had designed it that way.

Nobody knew it depended on that.

The server upgrade changed the timing.

The race appeared.

A faster computer had made the healthcare application less correct.

They fixed the race.

Afterward, Eli added a question to his personal design checklist:

What is this system accidentally depending on being slow?

It would prove surprisingly useful.

Chapter Five: The Payment That Had No Name

The payment system was worse.

Different screens calculated balances differently.

A payment could be correct in one part of the application and wrong in another.

Support had historically fixed individual displays.

The customer would edit something.

The fix would disappear.

Eventually Eli was assigned one of the worst payment cases.

A biller had submitted an exceptionally large batch.

Nothing appeared to happen.

So she clicked Submit again.

Still nothing.

She clicked again.

Eventually CareLedger finished.

Several times.

There was no safe way to identify the duplicate payments.

Eli searched for the payment ID.

There wasn’t one.

A payment was identified by…

every field in the payment row.

All of them.

If two rows contained exactly the same values, then as far as the system was concerned, they were the same description of a payment with no independent identity.

Unfortunately, two legitimate payments could also have identical values.

“So which rows are the accidental duplicates?” someone asked.

Eli studied the production data.

There was no answer stored anywhere.

The application had discarded the fact at creation time.

“Best evidence says these,” he said.

“You’re sure?”

“No.”

They deleted them.

Eli described the procedure later as:

YOLO accounting.

To prevent recurrence, the team added a progress display.

Processing 317 of 2,846 payments…

It seemed like a modest UI improvement.

It was actually preserving financial integrity by convincing users not to submit the same operation repeatedly.

The backend mechanism made this even funnier.

CareLedger processed the payments quickly by invoking shell commands which launched HTTP requests back into CareLedger itself.

PHP called the shell.

The shell launched requests.

The requests called PHP.

The web application was its own distributed job queue.

The ampersand character was its message broker.

Chapter Six: One Table

CareLedger had been created when MySQL commonly defaulted to MyISAM.

The original developer rarely specified the storage engine.

There was no reason to.

Until years later, when MySQL defaulted to InnoDB.

A database administrator replaced one table.

Same columns.

Same indexes.

Same data.

Production nearly collapsed.

The DBA investigated queries.

CPU.

Disk.

Indexes.

Connections.

Nothing made sense.

Eli had recently been studying database internals.

“What storage engine is the new table?”

Silence.

They checked.

InnoDB.

The surrounding tables were MyISAM.

CareLedger contained a substantial number of queries joining tables and relying, unknowingly, on the locking behavior of the original engine.

Replacing a single table had changed the concurrency model of part of the application.

The schema looked the same.

The application behaved as though civilization had ended.

The team documented another invariant:

DO NOT MODERNIZE ONE TABLE.

Chapter Seven: We Made It Faster

Another team was asked to improve an integration process.

The process was slow.

Painfully slow.

They found an obvious missing database index.

They added it.

Then they added parallelism.

The improvement was monumental.

Data that had crawled through the system now flew.

Everyone celebrated.

Then the backend server began dying.

The database was fine.

The CPU was fine.

The filesystem was not.

The old process had been so slow that the database bottleneck had accidentally rate-limited filesystem writes.

Once the database became fast, CareLedger could generate writes faster than the storage subsystem could sustain them.

The optimization had removed the system’s flow-control mechanism.

That mechanism had been:

being terrible.

They implemented actual backpressure.

Eli added another question to his checklist:

If I remove this bottleneck, what was it protecting?

Chapter Eight: Custom Reporting

CareLedger offered customizable reports.

Customers loved them.

They could select friendly fields such as:

Patient Name

Date of Service

Insurance Payment

Eli eventually discovered how these friendly selections mapped to the database.

The value behind the dropdown option was SQL.

Not an abstract report field.

Not a server-side identifier.

SQL.

Sometimes a column.

Sometimes an expression.

The report builder was effectively a graphical interface for assembling pieces of a query.

The frontend was part of the database abstraction layer.

This discovery led to a philosophical discussion.

One engineer said, “Technically, it’s very flexible.”

Nobody disagreed.

Chapter Nine: The Lock

CareLedger had a lock-screen feature.

A clinician could lock the application before walking away.

The screen became inaccessible until the password was re-entered.

The password was checked on the backend.

This sounded reassuring.

The lock itself was a frontend overlay.

The authenticated session remained authenticated.

Removing or bypassing the overlay revealed the still-active application beneath it.

Worse, before Meridian hardened it, the password-checking endpoint was vulnerable to SQL injection.

Thus the security model had three layers:

A visual barrier.

A real password check.

Bobby Tables.

The team fixed it.

Nobody commemorated the occasion.

Chapter Ten: The Protocol

XHR debugging had its own protocol.

Endpoints returned JSON.

Then came a separator.

Then arbitrary debugging output.

The JavaScript split the response at the separator, parsed the JSON, and displayed or ignored the rest.

This was not entirely unreasonable for a home-grown tool built before modern browser development workflows matured.

The remarkable part was that different pages used different separators.

Therefore CareLedger did not have one strange application protocol.

It had many.

Every page spoke a slightly different dialect of:

JSON, followed by whatever the developer wanted to print today.

Chapter Eleven: Twelve Thousand Lines

Eventually Eli identified one script that, despite everything, served as a reasonable candidate for the authoritative payment-write path.

It was enormous in responsibility.

But at least it existed.

He designed a replacement consisting of roughly fifty classes.

His team divided them up.

Eli wrote the characterization test.

The test was somewhere between ten and fifteen thousand lines.

Management was nervous.

They should have been.

Before go-live, the company ran production data through both implementations and compared the results.

Old system.

New system.

Same inputs.

Nearly identical outputs.

The few differences were cases Eli had already identified and deliberately changed.

They deployed.

Nothing exploded.

In CareLedger, this counted as a transcendent engineering achievement.

Chapter Twelve: The Most Stable Application in the Company

Years passed.

Many of the worst failure modes were removed.

The codebase remained ugly.

There were still ancient pages.

Still strange abstractions.

Still enough questionable code to make a new engineer reconsider his career.

But something unexpected happened.

CareLedger became stable.

Very stable.

Eventually, more stable than Meridian’s primary application.

This seemed impossible until Eli understood why.

The team had been burned by almost every category of accidental architecture imaginable.

They had learned to ask questions other teams considered paranoid.

What happens if the job runs twice?

What happens if it never runs?

What happens if this becomes faster?

What happens if this becomes slower?

What if the server restarts here?

What if two values are identical?

What if the cache returns somebody else’s state?

What if the storage engine changes?

What if this screen is lying?

What if the thing we are fixing is currently protecting us from something worse?

The code had not become beautiful.

The engineers had become careful.

CareLedger had begun life as an application that worked because an extraordinary number of accidents remained true simultaneously.

It survived long enough for a team to discover those accidents one at a time and replace many of them with guarantees.

At the next architecture meeting, someone proposed replacing another old subsystem.

Eli looked at the diagram.

Everyone waited.

Finally he asked:

“What terrible thing is currently depending on this terrible thing?”

Nobody laughed.

It was a serious question.

And that, more than anything else, was how they knew CareLedger had taught them something.