Everything here is based on real experiences of a single application, but composed into a story by AI. All company/individual names have been replaced by fictional names. And yes, it was exactly as bad as the story makes it seem. None of these are made up. Reposted now that I have enough subreddit karma to actually post here.
The Lunch Incident
A Short Novel About Accidental Architecture
Chapter One: The System Works
On Eli’s second week at Meridian Therapy Systems, someone told him not to worry about the billing application.
“It’s old,” his manager said, “but customers love it.”
This was true.
CareLedger—an acquired application used by thousands of small therapy practices—was fast enough, familiar, and remarkably well suited to the people who used it. Clinicians could schedule patients without thinking like programmers. Billers could enter payments quickly. Small practices could submit claims without buying an enterprise revenue-cycle system.
It was also, Eli soon discovered, held together by a series of engineering decisions that would have been rejected from a freshman software-design project.
The original developer had understood users exceptionally well.
Computers, somewhat less so.
Every practice had its own database.
Not its own schema.
Its own database.
Thousands of them.
Eli stared at the architecture diagram.
“Why?”
“Isolation,” said an engineer.
“That’s actually not terrible.”
“No,” the engineer said. “Keep looking.”
There was a pause.
“Oh.”
The web application had administrative privileges over all of them.
Chapter Two: Lunch
The incident arrived through Support.
A clinic reported that a large portion of its appointment schedule had been destroyed.
The appointments themselves had not disappeared.
That would almost have been easier.
Instead, real patient appointment data had been overwritten by a recurring calendar event.
Its name was:
Lunch
Names, details, and other appointment information that clinicians relied upon had been replaced. Where the clinic expected to find its patient schedule, CareLedger confidently informed them that they were having lunch.
Engineering reconstructed what had happened.
A clinician had created an ordinary schedule block called Lunch.
Somewhere beneath the friendly calendar interface, CareLedger had decided that the clinician’s request applied to rather more appointments than the clinician intended.
There was, fortunately, tenant isolation.
The bug did not turn every appointment in the company into Lunch.
Only a substantial portion of one clinic’s schedule.
The clinic also had something CareLedger did not expect:
paper.
Staff reconstructed enough of the schedule from printed records and other offline information to continue seeing patients while engineering investigated.
And Meridian had database snapshots.
The damaged appointment data could therefore be restored rather than reconstructed entirely by hand.
This fact was discussed as good news.
From that day forward, engineers referred to it simply as the Lunch Incident.
Nobody needed clarification.
Chapter Three: The Free Customer
Meridian charged practices monthly for using CareLedger.
At least, that was the idea.
One practice filed an unusually large number of bug reports. Support knew them well. Engineering knew them better.
One day someone noticed something odd.
They weren’t paying.
Not because their card had failed.
Not because their account had been comped.
CareLedger simply wasn’t charging them.
The billing system worked by scheduling a future event.
When that event ran, it charged the customer.
Then—and only then—it scheduled the next billing event.
Eli stared at the code.
“So if the event doesn’t run…”
“There’s no next event.”
“And something later notices they’re overdue?”
“No.”
“There’s a reconciliation job?”
“No.”
“Accounts receivable?”
“Not really.”
“So if the server is down when the event fires…”
“They stop getting billed.”
“For how long?”
“Forever.”
This had happened to more than one customer.
Deployments could cause it.
Outages could cause it.
Anything capable of preventing the event from firing could permanently transform a paying customer into a free customer.
The billing application had implemented recurring revenue as a chain letter.
Break the chain and the customer escaped.
Chapter Four: The Faster Server
Eventually Meridian migrated CareLedger to newer servers.
No application code changed.
No database code changed.
No UI code changed.
Soon afterward, clinicians began opening appointments and finding large sections of patient information mysteriously blank.
There was no deletion audit.
There was no obvious failed update.
The data seemed simply to disappear.
Engineering searched for days.
Then Eli asked a question that sounded ridiculous.
“What if the new server is too fast?”
Silence.
Two operations were occurring without explicit sequencing. On the old server, one was reliably slow enough that the other always completed first.
Nobody had designed it that way.
Nobody knew it depended on that.
The server upgrade changed the timing.
The race appeared.
A faster computer had made the healthcare application less correct.
They fixed the race.
Afterward, Eli added a question to his personal design checklist:
What is this system accidentally depending on being slow?
It would prove surprisingly useful.
Chapter Five: The Payment That Had No Name
The payment system was worse.
Different screens calculated balances differently.
A payment could be correct in one part of the application and wrong in another.
Support had historically fixed individual displays.
The customer would edit something.
The fix would disappear.
Eventually Eli was assigned one of the worst payment cases.
A biller had submitted an exceptionally large batch.
Nothing appeared to happen.
So she clicked Submit again.
Still nothing.
She clicked again.
Eventually CareLedger finished.
Several times.
There was no safe way to identify the duplicate payments.
Eli searched for the payment ID.
There wasn’t one.
A payment was identified by…
every field in the payment row.
All of them.
If two rows contained exactly the same values, then as far as the system was concerned, they were the same description of a payment with no independent identity.
Unfortunately, two legitimate payments could also have identical values.
“So which rows are the accidental duplicates?” someone asked.
Eli studied the production data.
There was no answer stored anywhere.
The application had discarded the fact at creation time.
“Best evidence says these,” he said.
“You’re sure?”
“No.”
They deleted them.
Eli described the procedure later as:
YOLO accounting.
To prevent recurrence, the team added a progress display.
Processing 317 of 2,846 payments…
It seemed like a modest UI improvement.
It was actually preserving financial integrity by convincing users not to submit the same operation repeatedly.
The backend mechanism made this even funnier.
CareLedger processed the payments quickly by invoking shell commands which launched HTTP requests back into CareLedger itself.
PHP called the shell.
The shell launched requests.
The requests called PHP.
The web application was its own distributed job queue.
The ampersand character was its message broker.
Chapter Six: One Table
CareLedger had been created when MySQL commonly defaulted to MyISAM.
The original developer rarely specified the storage engine.
There was no reason to.
Until years later, when MySQL defaulted to InnoDB.
A database administrator replaced one table.
Same columns.
Same indexes.
Same data.
Production nearly collapsed.
The DBA investigated queries.
CPU.
Disk.
Indexes.
Connections.
Nothing made sense.
Eli had recently been studying database internals.
“What storage engine is the new table?”
Silence.
They checked.
InnoDB.
The surrounding tables were MyISAM.
CareLedger contained a substantial number of queries joining tables and relying, unknowingly, on the locking behavior of the original engine.
Replacing a single table had changed the concurrency model of part of the application.
The schema looked the same.
The application behaved as though civilization had ended.
The team documented another invariant:
DO NOT MODERNIZE ONE TABLE.
Chapter Seven: We Made It Faster
Another team was asked to improve an integration process.
The process was slow.
Painfully slow.
They found an obvious missing database index.
They added it.
Then they added parallelism.
The improvement was monumental.
Data that had crawled through the system now flew.
Everyone celebrated.
Then the backend server began dying.
The database was fine.
The CPU was fine.
The filesystem was not.
The old process had been so slow that the database bottleneck had accidentally rate-limited filesystem writes.
Once the database became fast, CareLedger could generate writes faster than the storage subsystem could sustain them.
The optimization had removed the system’s flow-control mechanism.
That mechanism had been:
being terrible.
They implemented actual backpressure.
Eli added another question to his checklist:
If I remove this bottleneck, what was it protecting?
Chapter Eight: Custom Reporting
CareLedger offered customizable reports.
Customers loved them.
They could select friendly fields such as:
Patient Name
Date of Service
Insurance Payment
Eli eventually discovered how these friendly selections mapped to the database.
The value behind the dropdown option was SQL.
Not an abstract report field.
Not a server-side identifier.
SQL.
Sometimes a column.
Sometimes an expression.
The report builder was effectively a graphical interface for assembling pieces of a query.
The frontend was part of the database abstraction layer.
This discovery led to a philosophical discussion.
One engineer said, “Technically, it’s very flexible.”
Nobody disagreed.
Chapter Nine: The Lock
CareLedger had a lock-screen feature.
A clinician could lock the application before walking away.
The screen became inaccessible until the password was re-entered.
The password was checked on the backend.
This sounded reassuring.
The lock itself was a frontend overlay.
The authenticated session remained authenticated.
Removing or bypassing the overlay revealed the still-active application beneath it.
Worse, before Meridian hardened it, the password-checking endpoint was vulnerable to SQL injection.
Thus the security model had three layers:
A visual barrier.
A real password check.
Bobby Tables.
The team fixed it.
Nobody commemorated the occasion.
Chapter Ten: The Protocol
XHR debugging had its own protocol.
Endpoints returned JSON.
Then came a separator.
Then arbitrary debugging output.
The JavaScript split the response at the separator, parsed the JSON, and displayed or ignored the rest.
This was not entirely unreasonable for a home-grown tool built before modern browser development workflows matured.
The remarkable part was that different pages used different separators.
Therefore CareLedger did not have one strange application protocol.
It had many.
Every page spoke a slightly different dialect of:
JSON, followed by whatever the developer wanted to print today.
Chapter Eleven: Twelve Thousand Lines
Eventually Eli identified one script that, despite everything, served as a reasonable candidate for the authoritative payment-write path.
It was enormous in responsibility.
But at least it existed.
He designed a replacement consisting of roughly fifty classes.
His team divided them up.
Eli wrote the characterization test.
The test was somewhere between ten and fifteen thousand lines.
Management was nervous.
They should have been.
Before go-live, the company ran production data through both implementations and compared the results.
Old system.
New system.
Same inputs.
Nearly identical outputs.
The few differences were cases Eli had already identified and deliberately changed.
They deployed.
Nothing exploded.
In CareLedger, this counted as a transcendent engineering achievement.
Chapter Twelve: The Most Stable Application in the Company
Years passed.
Many of the worst failure modes were removed.
The codebase remained ugly.
There were still ancient pages.
Still strange abstractions.
Still enough questionable code to make a new engineer reconsider his career.
But something unexpected happened.
CareLedger became stable.
Very stable.
Eventually, more stable than Meridian’s primary application.
This seemed impossible until Eli understood why.
The team had been burned by almost every category of accidental architecture imaginable.
They had learned to ask questions other teams considered paranoid.
What happens if the job runs twice?
What happens if it never runs?
What happens if this becomes faster?
What happens if this becomes slower?
What if the server restarts here?
What if two values are identical?
What if the cache returns somebody else’s state?
What if the storage engine changes?
What if this screen is lying?
What if the thing we are fixing is currently protecting us from something worse?
The code had not become beautiful.
The engineers had become careful.
CareLedger had begun life as an application that worked because an extraordinary number of accidents remained true simultaneously.
It survived long enough for a team to discover those accidents one at a time and replace many of them with guarantees.
At the next architecture meeting, someone proposed replacing another old subsystem.
Eli looked at the diagram.
Everyone waited.
Finally he asked:
“What terrible thing is currently depending on this terrible thing?”
Nobody laughed.
It was a serious question.
And that, more than anything else, was how they knew CareLedger had taught them something.