A story from the field, with names changed to protect the innocent.
A few years ago, we were brought in to help a client whose platform had a reliability problem. Almost every week, the site went down, and the cause was usually the same: the database buckled under load. Every new feature took longer to ship than the last, incidents were hard to diagnose, and each new batch of users made things worse.
The root cause wasn’t a weak team or a bad product idea. It was a technology choice made years earlier, and everything that had been built around it ever since.
The setup
The platform had been built from scratch by an independent consultant, whom we’ll call Mark X. He built it on XSpaces, an open-source component framework built on JSF and publicly available. What made the choice unusual is that Mark had written XSpaces himself. Nobody knew the technology better, and for a while the platform did what the client needed.
But a technology can be well crafted and still be the wrong one for the job. This platform needed to grow, and its foundation wasn’t built for that.
What we found
- The stack was heavy to write, read and maintain. Even simple changes meant navigating layers of component abstractions and configuration, which made the code slow to understand for anyone who hadn’t lived in it for years.
- The architecture didn’t scale with users. JSF’s server-side component model keeps a lot of state per session, and the way the application interacted with the database put growing pressure on it as concurrency increased.
- Workarounds were piled on workarounds. Whenever the platform hit the limits of the framework, the answer was a patch. Over the years these patches became load-bearing: some compensated for the framework’s shortcomings, others for the side effects of earlier patches, and some of them added to the very database load that was causing the outages.
The result was a system that fell over almost weekly, where changing one thing risked breaking three others, and where every incident was a firefight rather than a fixable problem.
The human side
This wasn’t purely a technical story. XSpaces was Mark’s own creation, and the platform was built on it, so he was understandably protective. Suggestions to simplify or move away from parts of the stack were often met with resistance, because when the technology is your own work, every critique of it feels personal.
We tried to approach it with empathy. The choices had made sense at the time, and he had kept the platform running for years. But the risk to the client was real: a critical system whose inner workings, and all its accumulated workarounds, lived mostly in one person’s head.
After a few years, Mark moved on, and that’s when the real work began.
Taking over
A handover like this is never clean. Our first priority wasn’t to rewrite anything, but to stop the bleeding and understand what we had.
- Reduce the outages. We started with the database: finding the heaviest queries and access patterns, fixing the worst offenders, and adding monitoring so we could see what the platform was really doing under load.
- Map and document. We traced the critical flows end to end, separated essential workarounds from historical accidents, and wrote down what had only existed as tribal knowledge.
- Add safety nets. We put safeguards in place around the most critical flows, so we could change things without fear.
The phase-out strategy
A full rewrite was tempting, and it would have been the wrong call. The business depended on the platform, and big-bang rewrites have a long record of running late and over budget.
Instead, we designed a gradual phase-out of the JSF-based stack, in the spirit of the strangler fig pattern:
- New functionality was built on a simpler, more mainstream architecture from day one, so the legacy part stopped growing.
- Existing modules were migrated one at a time, ordered by business value and pain rather than technical elegance.
- Workarounds whose reason for existing had disappeared were removed instead of ported over.
- The overall architecture was simplified, with fewer layers, clearer boundaries and less hidden state, and a much lighter footprint on the database.
Throughout, the platform stayed live and users kept working.
The results
Over time, the platform went from fragile to dependable:
- Incidents went from almost every week to once every few months.
- It can handle far more concurrent users, with much more headroom than before.
- Maintenance is predictable, and new developers get productive much faster on a simpler codebase built on mainstream technology.
- The client no longer depends on any single person to keep the system running.
What we took away from it
Choose technology for where the product is going, not just where it is today. A stack that works for hundreds of users can quietly become the bottleneck at thousands.
Be wary of niche technology, even when it’s open source. Publishing a framework doesn’t create a community around it. If the only real expert is the person who wrote it, the bus factor is still one.
Recurring incidents are a signal, not a fact of life. If the platform goes down every week, the fix is rarely another patch. It’s usually the architecture.
Workarounds are debt with interest. Each one is cheap to add and expensive to live with, and eventually someone pays the bill.
Migrate incrementally. Evolution beats revolution when the system has to keep running.
Treat the people with respect. Legacy code is never just code. Someone made those choices under real constraints, and the best rescues honor that while still moving forward.
At bitGloss, this is the kind of work we enjoy: stepping into a complicated situation, understanding it before judging it, and leaving the client with a platform that’s simpler, faster and easier to own. If your team is dealing with a system that has become harder to change than to live with, we’d be glad to talk.