Building systems for working with AI agents is software engineering, whether or not the builder knows it. And increasingly the builders aren’t engineers: a growing crowd of vibe architects with little technical background are crafting complex systems that rely on large context libraries to help them in their daily work. In a recent study of these users, we recognized a familiar pattern that has plagued software architects for decades.
Software engineers learned long ago that systems patched together to meet today’s needs turn fragile and unmaintainable. They learned to plan their architecture so that changing a system later wouldn’t cost more time, money, and expertise than they had. AI makes all these resources feel unlimited, and we saw the result: with nothing forcing them to prioritize, these vibe architects were building big balls of mud. In the AI era, a big ball of mud works surprisingly well — until accumulated complexity makes the system difficult to direct, diagnose, or reconstruct.
What Is a Big Ball of Mud?
A big ball of mud is a system scraped together for short-term functionality, not architected for long-term durability.
The term comes from Brian Foote and Joseph Yoder’s 1997 paper, which observed that most real software at the time looked this way. Such a ball-of-mud system may be useful, but its fragility makes it costly to maintain. Like a shantytown built hastily out of necessity, it needs constant patching and jury-rigging to keep running, rather than holding up like the infrastructure of a master-planned city.
Big balls of mud typically emerge when scarcity meets reality. There’s not enough time, money, or expertise, yet something must be created. Early software teams worked in a “code-and-fix” mode, which sufficed until problems got harder and their big balls of mud became unsustainable. The industry responded with the heavily planned “waterfall” era, modeled on mature engineering disciplines. But when it proved impossible to foresee every need up front, it moved toward an Agile framework.
Our participants were retreading this history. They coded and fixed, with little forethought about architecture or durability, planning to keep molding the clay to fit whatever problems arose.
The difference today is that AI has turned that scarcity into abundance. Building is cheap, so nothing forces a builder to prioritize, and people overbuild. One advanced participant said:
“When I talk to people who say they are advanced in AI, what I often see is what I consider overengineered. It has a lot of fancy words, a lot of flashy tools, but it doesn't get you that much more value for the time you spend in it … Because AI makes it so easy for you to build something, people actually start overbuilding because it feels productive."
Most of our participants built this way, often working alone, guided by the AI itself and by posts on X or Reddit, with subscriptions that put vast numbers of tokens at their disposal.
5 Ball-of-Mud Patterns
Foote and Yoder described the patterns by which a big ball of mud forms and persists. We saw 5 of them in our sessions. These 5 aren’t strict stages — a system keeps growing piecemeal while its owner patches and sweeps — but they tend to appear roughly in the order below.
1. Throwaway Code
It starts when a user has the AI build a proof of concept for something that seems useful. For example, they create a rules.md file to restrain their agents from annoying or problematic behaviors. The file isn’t planned or thoroughly tested; it is a first draft. But the user quickly moves on to other things while the AI diligently continues referencing the context that was created, as instructed. And before long, the user’s workflows have adapted to rely on it.
One participant described a daily automation that summarized his emails, tasks, and meeting notes into that day’s priorities:
“I didn't write the structure myself… I worked with Claude […] to just build the structure for itself […] now I'm kind of […] stuck in the way that it works. And I've never really taken the time to greenfield improve what I did."
These throwaway pieces are rarely useless, so they stay. Participants seemed to have two reasons for not having the AI rebuild them: a lack of time and fear of breaking the existing value. The system was like a pair of worn-out shoes: functional enough to keep wearing, but not necessarily comfortable. Participants kept wearing them instead of taking the time to get a new pair that might do a worse job or not fit as well. Eventually, the shoes wore out, and reconstruction became necessary, but users deferred that investment and risk as long as they could.
2. Piecemeal Growth
As new needs arise, users bolt new functionality onto the original version of their system that they’re hanging on to. One participant’s tool went through three stages:
Stage 1: The participant noticed that her team had limited visibility into work tracked in Asana, so she used AI to build an Asana widget that surfaced each project’s status without requiring stakeholders to sift through more than 50 tasks.
Stage 2: Then, a new need arose. Stakeholders using the widget assumed projects had stalled because many tasks in Asana hadn’t been checked off properly. So, she had the AI add a dashboard component alongside her widget to show more granularity.
Stage 3: She then saw a chance to improve the whole organization’s visibility into every project. She and the AI merged the original widget and newer dashboard into one tool, adding visualizations and weighted algorithms to predict completion dates and time tracking for finance.
This participant could not write code. Yet, with AI removing the brakes, a small throwaway widget grew into a crossfunctional dashboard. This is piecemeal growth.
Our participants were proud of what they had built, and rightly so. The problem is that the number of ways to extend the system never stops growing — particularly once others have input. Several participants regularly asked the AI itself to suggest improvements, but none worked within planned release cycles. Their creations had no “finished” state, even temporarily — and that is when their systems started turning into a big ball of mud.
One of the most advanced users in the study described his wakeup call:
“[Claude is] always hammering me that I'm stuck in this prerevenue infrastructure loop. And basically, what it tells me is, ‘you're constantly trying to refine the system, and you're not actually going out and using the system.’”
Other participants could not locate, in front of us, important files their systems relied on; they had to ask the AI to find them and explain what they contained. The system had slowly grown to the point where they no longer understood exactly where everything lived, how the AI was using it, or what it contained. If they had to recreate the system, they probably couldn’t.
3. Keep It Working
Eventually, a system breaks, drifts, or goes stale, and the user must patch it, abandon it, or overhaul it. Our participants mostly patched, because they relied on their systems. A startup founder described her feelings this way:
“I might even be dependent on it […] If I experience even a slight amount of challenge […] I ask it to help me instead of actively thinking about it myself […] I think I would definitely feel frustrated if the tool was taken away.”
When one participant's AI desktop application crashed and deleted his chat history, he retrieved the files from his trash, dropped them into his Downloads folder, and began running every session out of it. This is what a big ball of mud looks like.
But there are other ways the system can fail, too. For example, one participant connected Todoist (a to-do app) to the terminal where he primarily worked with the AI. The connection failed regularly. Rather than switching to-do apps or reconfiguring the system, he reconnected it by hand each time.
Luckily, AI lowers the barrier tremendously for changing the system. If the context files that an agent uses aren’t working, fixing them might just be a quick prompt away. This was how our participants maintained their systems: they described the change and let the AI fix files wherever it saw fit. Such a process is easy, but it can erode the user’s understanding of where information lives in the system. Users end up dependent on the AI to find, understand, and update it for them. They are more focused on short-term fixes than long-term clarity and durability.
4. Sweeping It Under the Rug
When users can’t or won’t clean up a mess, they contain it: everything goes into one place so the rest of the system stays tidy. In agentic systems, the rug is really a junk drawer — the user keeps needing what’s in it, but only the AI can find anything. One participant emailed himself articles and links throughout the day for the AI to collect. There were too many to process, so they went into a folder called to be synthesized, with no naming scheme. When he later asked the AI to pull relevant material from that folder for a project, it retrieved the wrong articles, and he spent 45 minutes fixing the instructions instead of doing real work.
Another participant let Claude save the files it created as Google Docs in the top-level folder of his Drive. Claude would have filed them anywhere he asked; he just never did.
In both cases, the folder was the junk drawer: everything the AI compiled or created went in, and nothing was sorted.
This pattern resembles Harry Potter visiting the Gringotts bank. Harry knows he has a vault, and he vaguely knows what’s in it, but he can’t find it or access it without a goblin guide. Our participants relied on the AI to navigate their mess, and it did so surprisingly well, which makes it tempting to keep sweeping more and more under the rug. Whether the AI can keep compensating as the pile grows is an open question; the 45 minutes lost to the wrong files suggest it has limits.
5. Reconstruction
Eventually, many systems, AI-centric or not, are rebuilt or abandoned. We saw three triggers for reconstructing an agentic AI system.
The System No Longer Fits the Need
In traditional software, this trigger came sooner because changes were costly and got deferred. A conscientious vibe architect can stave it off longer by having the AI make updates as needs change.
But the fit between user and system can also expire as the user outgrows the system. As users learn what the AI can do, their way of working changes, until rebuilding looks easier than more tweaking. One participant drew a flowchart of how he now wanted every session to run (clarify intent, gather context, orchestrate the model, check the outputs) — steps his current system didn’t support. Unlike his piecemeal first build, this one was planned. He uploaded the flowchart to the AI and asked it to implement it.
Other times, the system was built for someone else. One participant inherited an agentic workflow and a context library from his manager; over time, he found they didn’t perfectly fit how he worked, so he personalized the system. He wasn’t patching holes to preserve the original design; he was renovating it. He called this drift the hardest part of sharing an agentic workflow with another person.
Nobody Understands the System Anymore
As a system grows piecemeal, even its creator can lose track of how it fits together. One participant, a product lead, had built a context library that his whole team was required to use. By the time we met him, he had rebuilt it once and was rebuilding it again:
"At some point around December, this had grown to something that was beyond what I was still able to keep in my head. And so, I had to rework this all to make it manageable again."
This can also happen when a system is handed off to others who did not create it. Someone else’s big ball of mud is hard to understand because you have no mental model for how it's put together. The AI can help the new owner get up to speed, but they will likely want to renovate.
A Better Way of Working Emerges
Sometimes the system works fine, but the user discovers a better way to do the same work. One participant did this in two ways: he regularly picked up new skills and techniques on social media and implemented them in his system. He also regularly asked the AI to review how he worked with it and suggest improvements — and then, usually, to implement them wherever it saw fit.
How Do I Know If I Should Reconstruct?
The stages building up to a big ball of mud aren't a mistake; they're a legitimate strategy that AI has made more manageable than ever. So, when is reconstruction necessary, and how can vibe architects prepare for it?
Start with a very simple question: If you wanted to change how the system works, would you know what to change in the context library the model relies on? The context library is made up of 3 types of contexts with different roles:
- Global context: instructions or knowledge that apply across the system
- Local context: information and instructions for a particular task, project, or workflow
- Ambient context: information streams such as email, Slack, or meeting transcripts that are available to the AI but aren't deliberately maintained as persistent context
If the answer is yes, meaning that you still know which of these types of context needs to change, then total reconstruction is likely far off because you still understand the system well enough to make small renovations without overhauling it.
If the answer is no, the different types of context have blurred together. Nothing is necessarily broken, but the AI is the sole navigator, and you are probably just hoping the system doesn't break.
Keep Tabs on Each Context Type Separately
The key to maintaining the integrity of an agentic system is keeping tabs on each context type. Maintaining each one independently makes it possible to avoid forced reconstruction for a long time, much like how many buildings last for hundreds of years, with each piece slowly maintained or replaced as needed. It's when the context types start blurring together that users accelerate toward system abandonment and reconstruction. And at that point, users will struggle to know exactly what needs to change and will rely heavily on the AI's understanding of the system to salvage what’s working and rebuild from there. It will be difficult to precisely guide the reconstruction process.
Keeping tabs on contexts doesn't mean memorizing the system. It means a few small habits:
- Keep the types of context separate. When one markdown file contains both global and local context, the lines have blurred, making it difficult to parse which parts of the system play which roles.
- Name the context type when directing a change. Even if users don't know exactly where each file lives or every word inside it, knowing what's happening in each context type allows them to tell the AI, "go update XYZ in the global context" rather than "go fix this wherever you think it matters," retaining an understanding of the impact changes have.
- Audit the global context from start to finish once in a while. This context type affects everything the system does, so it should be short enough to review in one sitting. If it isn’t, instructions that don’t apply everywhere have probably crept in; delete them or move them to local context.
- Refresh local context when the outputs suffer. Local context is meant to serve a specific purpose and should be shaped accordingly when it fails to do so.
Efforts like these are renovations, not reconstruction. This is what makes reconstruction a choice rather than an emergency.
All that said, users might decide to reconstruct anyway because their ways of working have evolved. If they can articulate the roles each context type plays, they're in good shape. The goal was never to avoid reconstruction — it's to stay in a position to direct it.
Conclusion
We saw mud everywhere in our sessions — even from the most advanced participants. Foote and Yoder argued that mud is often a pragmatic response to constraints. With AI, many constraints are gone, but the mud remains, because nothing forces builders to ask whether a new piece of the system is worth its maintenance cost or to carefully plan the system’s architecture.
Unlike traditional software, an agentic system is fast and cheap to reconstruct with AI. The hard part is knowing what the new system should be. The participants who seemed best positioned were those who could articulate the limitations of their systems and what they wanted to change. The discipline that matters most is not tidier folders but periodic reflection: knowing what each piece of the system does, whether you still use it, and what you would rebuild if you started over. Mud is tolerable while the AI can still navigate it for you; the time to reconstruct is before it can’t.
References
Brian Foote and Joseph Yoder. 1999. Big Ball of Mud. In Pattern Languages of Program Design 4. Addison-Wesley, Boston, MA, 653–692.