Somewhere in the first serious production incident on each of these products, a client asked a version of the same question. They never asked “can you fix it.” They asked “how do we know this won't happen again, and whose fault was it when it did.” That question doesn't have a technical answer. It has a governance answer. If you haven't built one before the incident, you're improvising it in front of a client who's already lost some trust.
Across Sustainability Intelligence, Carbon Pre-Audit, Carbon Credit Management, and Sustainability Benchmarking, the products that earned real trust weren't the ones that were never wrong. None of them are never wrong. They were the ones where being wrong was something the system, and the relationship with the client, was actually built to handle.

Traditional software fails loudly. AI fails quietly and plausibly.
Clients coming from traditional enterprise software have a mental model of failure that doesn't transfer. A report generator either runs or throws an error. A calculation either uses the formula in the spec or it's a bug you can point to and fix. That model assumes determinism: same input, same output, every time, and when something's wrong, it's wrong in a way you can reproduce and root-cause.
AI output doesn't fail that way. Ask the same underlying question with slightly different phrasing and you can get a different answer. That doesn't mean anything broke. It's simply how the technology works. The first time a client noticed two answers from Sustainability Benchmarking that weren't word-for-word identical for what they considered “the same question,” the client didn't react with “interesting, tell me about non-determinism.” They reacted with “which one is the mistake, and how many other mistakes are in here.” Explaining that variability is expected and doesn't itself indicate an error, while also being honest that it makes verification harder, is a conversation you need to have proactively, in plain language, before launch. Having it for the first time during an incident makes you sound like you're rationalizing a bug.
Where a human needs to be in the loop, and where they don't
We put a judge model in front of every output. About 30% still needed a human anyway, which is the real automation ceiling on this kind of work.
The instinct after a mistake is often “put a human in front of everything.” That's not free, and if a client is paying for automation, unwinding it into full manual review defeats the point. The harder and more useful question is where specifically a human checkpoint is worth its cost.
With Sustainability Benchmarking, the checkpoint we chose sits at a very specific moment: the first time a new company is introduced into the peer set. That's exactly the moment where unit conversions, normalization assumptions, and the calculations that make one company comparable to another are being decided: the volume-to-weight conversion, the per-employee normalization, the judgment calls we've talked about in earlier posts. We made the call that it's too much risk to let the model make those decisions in real time, unreviewed, the first time a new peer shows up. So a human reviews and confirms exactly how that new company's data was normalized before it's allowed to feed into any benchmark a client sees. Once a peer has been through that review, the ongoing comparisons against it don't need the same scrutiny every single time. The risk was never in running the comparison. It sat in the one-time judgment calls that decided how that company's numbers would be made comparable in the first place. That single checkpoint became the honest answer to “who's accountable here”: the model proposes the normalization, a named person confirms it before it's trusted, and the system records which one did what.
Sustainability Intelligence needed a different balance. Requiring human sign-off on every regulatory applicability question would have made the product unusably slow for its actual use case, a sustainability team scanning a large and constantly changing regulatory landscape. There, the accountability structure changed from “a human checks every answer” to “the system is explicit and specific about its own confidence, and routes only the genuinely ambiguous calls, new regulations, edge-case corporate structures, conflicting geography rules, to a human for review.” Not every product needs the same checkpoint in the same place. It needs a checkpoint calibrated to where the actual risk of being wrong is highest.
Setting expectations before launch, not after the incident
The single highest-leverage thing we started doing was writing down, before a client ever used the product, what the system is and isn't responsible for, in language a non-technical stakeholder would actually read. Not buried in a terms-of-service document, but in an actual conversation, and ideally in the onboarding materials themselves: what the model does well, where it's known to need human judgment, what “confidence” means in this product's specific outputs, and what the client's own responsibility is for reviewing output before it's used downstream (in a filing, a board presentation, an audit response).
The clearest version of this is drawing a hard boundary around what the product is not, and saying it before a client can assume otherwise. Sustainability Intelligence is a monitoring and news feed tool, AI-enabled and genuinely useful at surfacing what's changed and what's relevant to a company's specific profile, but it is explicitly not a compliance determination. It doesn't tell a client “you are compliant” or “you are not compliant,” and that boundary is stated to every client going in, not discovered by them later when they assume a green checkmark means more than it does. Carbon Pre-Audit draws the same kind of line in a different place: it is not an auditor, and it does not issue anything resembling a limited assurance letter. What it does is warn the customer, ahead of time, about the specific issues a real auditor is likely to flag, so they can fix them before the actual audit, not so they can treat the tool's output as the audit itself. In both cases, the guardrail is a sentence rather than a technical feature, said clearly and early, about what the product will never claim to be.
Who owns the mistake internally
The other half of accountability is internal, and it's just as easy to leave undefined until you're forced to figure it out live. When something goes wrong, is it an engineering bug, a domain-modeling gap, a data quality issue the client introduced, or a case where the product worked exactly as designed and the design itself was wrong for that situation? Those are different owners, different fixes, and different things to tell the client.
We got this wrong early, informally treating almost every issue as “an engineering bug to patch,” which meant real domain-modeling gaps, where the model didn't understand a nuance of GHG accounting that a domain expert would have caught immediately, got treated as one-off code fixes instead of systematic gaps in how the product understood the domain. Separating “this is a software defect” from “this is a domain knowledge gap” from “this is a data problem the client needs to fix on their end” changed both how fast we actually solved the underlying issue and what we were able to tell the client about what happened and why it won't recur in exactly that form.
What actually builds trust
Explain non-determinism before a client discovers it themselves. If two runs of the same question can produce different-looking answers, say so upfront, in plain language, and explain what that does and doesn't mean about correctness.
Put human checkpoints where the risk is highest, not everywhere. A blanket “human reviews everything” policy is both unaffordable and a tell that you haven't actually figured out where the real risk sits.
Make sign-off visible and attributable, not implicit. “Someone could have caught this” is not the same as “a named person confirmed this, and the system recorded it.”
Set expectations about responsibility at kickoff, not after an incident. What the system does, what it doesn't, and what the client is responsible for reviewing before relying on the output, all said out loud, early, to the actual stakeholders.
Separate the internal root-cause categories before you need them under pressure. Engineering bug, domain gap, data issue, and “working as designed but wrong for this case” are different problems with different owners and different client conversations.

The takeaway
Trust in an AI product isn't built by being right all the time. Nobody actually believes that's possible once they've used it for a while. It's built by having a clear, honest, pre-agreed answer to “what happens when it's wrong”: who catches it, who's accountable, and what changes afterward. Get that structure in place before the first real incident, and being wrong sometimes becomes a manageable part of the relationship instead of a crisis that erodes it.
That's the third lesson from the front line. Next up: the cost curve nobody warns you about, and what token economics actually do to your margins as an AI product scales.


