
The usual response to why companies aren’t using AI: "Our data is a mess. We have to clean it up first."
It sounds responsible and disciplined, but really, it’s the corporate version of saying you’ll start exercising once you’re in shape. The data will never be clean enough. There’s always one more system to reconcile, one more master file to scrub, one more region that reports things its own way.
Leaders who wait around for pristine data are going to be waiting a long time. While they wait, competitors are leapfrogging.
Bad data isn’t a reason to avoid AI. It’s a good reason to use it.
Before writing off data as unusable, figure out what’s wrong with it. In our experience, "dirty" data tends to fall into two different buckets.
The first is sloppiness, or genuine human error. Someone fat-fingered a posting, miscoded a transaction, or entered the same vendor three different ways. This is what everyone pictures when they think of bad data.
The second kind is dirty by design. The data looks wrong to an outsider, but it’s built that way for a reason. Salespeople and sales administrators sit in the same category because that’s how the system was set up. Twelve ERP business units don’t line up neatly with how leadership actually thinks about the business. It’s not error, it’s architecture. It’s something AI can handle well.
Companies get stuck because they confuse the two. A company-wide data cleanup is not always the answer.
Companies don’t have to overhaul historical records to start getting value. The trick is to work with the mess, using a few smart rules to help AI make sense of data exactly as it sits today.
Here’s what this looks like in practice.
1. Build groupings on the fly instead of re-mapping the source.
The classic example is the ERP that carves the business into twelve units when leadership only thinks in terms of seven or eight regions. The instinct is to kick off a project to re-map everything at the source, which means months of work, meetings, and change requests. You don’t need any of that. Train the model to interpret those twelve units into the eight regions your executives report on. The ERP stays untouched, and the regional view shows up on its own.
The same idea applies to a legacy vendor master file you’ll never realistically clean. When the same vendor shows up under three slightly different names, just put simple rules in place, like a parent-child relationship that ties the variants together, or an asterisk flag on related entries, so the model groups them correctly on the fly. Now you get consolidated vendor spend without the large cleanup project.
2. Separate "sloppy" from "by design," then let AI police the difference.
Not everything that looks wrong is wrong. Salespeople and sales administrators might sit in the same category because that’s how the system was built, and that’s fine. Genuine sloppiness is a different animal, like the miscoded posting or the transaction that landed in the wrong bucket. The move here is to train the model on what "normal" looks like so it can automatically put real anomalies into an exception report. Now teams don’t have to hunt through everything. They can work from a short, prioritized list of things that are broken.
3. Mask sensitive fields so privacy stops being a blocker.
Sometimes the hesitation isn’t messy data at all. It’s sensitive data. Say you need to link records across systems using a birth date but can’t have that value exposed. Process the data inside a secure, internal environment and convert the sensitive value into a shared integer. The model can match records across systems using that shared key without ever seeing or exposing the raw personal info. You get the linkage and private data never leaves the safe zone.
4. Define one specific problem, not the whole data landscape.
Don’t try to fix everything at once. Pick a single, narrow, high-value question, like one regional report, one vendor consolidation, or one exception workflow, and solve it. This creates little concrete wins over time. After solving one problem, the next one gets easier, because you’ve already built the groupings, the exception logic, and trust.
Let’s be clear about what AI does and doesn’t do here. It doesn’t make data hygiene optional forever. Clean data still matters, and the exception reports AI generates should feed a real cleanup effort over time. What AI does is take away the excuse for sitting still. Think of it as a bridge. It lets companies pull value out of the business in the short-term while identifying to where future cleanup should focus.
The next time someone says the data is too dirty to start, don’t let that be an excuse. It’s not caution. It’s avoidance.
At Trenegy, we help organizations implement AI tools that bring about tangible efficiencies and immediate impact. To chat more about AI in your organization, emailinfo@trenegy.com.