Talk to Us +91 98847 45599

AI & Automation

The Hard Part of Enterprise AI Is Not the Model

Choosing a model is the most discussed decision in an enterprise AI project and, increasingly, the easiest to reverse. The decisions that determine success are organisational.

Vianmax Editorial7 min read

Technical illustration of a small dark square at the centre of a large surrounding structure of data channels, gates and workflow lanes, drawn in fine lines, with one blue path running through it.

Most enterprise AI projects begin with a question about models. Which one is best? Should we use an open-weight model or a hosted one? Do we need to fine-tune? These are reasonable questions, and they get a great deal of attention, partly because they are the most visible choice and partly because they are the easiest to discuss in the abstract.

They are also, increasingly, the least decisive. Capable models are available from several providers, and systems designed with a clean boundary around the model can switch between them with modest effort. What cannot be switched easily is everything else: the data the system depends on, the permissions it respects, the workflow it fits into, the way its quality is judged, and the people who have to trust it.

When enterprise AI projects stall, they rarely stall on the model. They stall somewhere in that list. What follows is an account of where, drawn from the patterns that recur across organisations rather than from any single project.

The demo ran on clean data

Almost every AI initiative has a moment of early excitement. Someone assembles a handful of documents, asks a model questions about them, and the answers are impressive. Leadership sees the demo. A pilot is approved.

The documents in that demo were chosen by someone who knew what they contained. They were current, they did not contradict each other, and there were few enough of them that retrieval was trivial. Enterprise data is not like that. There are four versions of the same policy, and only one is in force. A product was renamed two years ago and half the documentation still uses the old name. The authoritative figure lives in a spreadsheet on someone's desktop, and the version in the shared drive is last quarter's.

A model cannot resolve these contradictions, because the organisation has not resolved them either. The first real work in most enterprise AI projects is therefore not AI work at all. It is deciding what the sources of truth are, retiring or labelling what is obsolete, and establishing who keeps each source current. That work is slow, it crosses departmental lines, and it rarely appears in the project plan. It is also the work that determines whether answers can be trusted.

An assistant can see too much

Consider a general example. An organisation builds an internal assistant over its document stores, and in testing it works well. Then someone asks it about salary bands, or a pending restructuring, or a customer dispute, and it answers helpfully from documents the person asking was never meant to see.

The underlying problem is that most document stores have permissions that were set loosely, on the assumption that obscurity was enough. Nobody would think to look in that folder. An assistant that indexes everything removes the obscurity. Suddenly every overly broad permission becomes an active disclosure.

The technical answer is well understood: retrieval must respect the access rights of the person asking, document by document, at query time. The organisational answer is harder, because implementing it properly often reveals how many permissions were wrong to begin with. Projects that treat this as a late security review tend to be delayed by it. Projects that treat it as a design constraint from the first week tend not to be.

Every overly broad permission in a document store becomes an active disclosure the moment an assistant can search it.

A separate window is a detour

A common pattern in early deployments is the standalone assistant: a new tool, in a new browser tab, where people can ask questions. Usage starts high and then falls away. The tool works, but using it means leaving the system where the actual work happens, copying context across, and copying the answer back.

The value of AI in enterprise settings tends to appear when it arrives inside the existing workflow. A support agent sees a drafted reply alongside the ticket they are already looking at. An accounts clerk sees extracted invoice fields already populated in the form they would otherwise type into. A reviewer sees a contract with unusual clauses highlighted, in the document itself.

This is integration work: connecting to the systems of record, understanding their data models, fitting into their interfaces and their permission schemes. It is often more effort than the AI component itself. It is also the difference between a tool people try and a tool people use.

Nobody agreed what good looks like

To know whether an AI system is working, you need to know what a correct output is. That sounds obvious until you try to write it down.

What is a good summary of a customer complaint? Should it preserve the customer's tone, or neutralise it? Should it recommend an action, or only describe the problem? Ask five experienced people and you may get five answers, each defensible. The AI system is being asked to meet a standard the organisation has never articulated.

Building an evaluation set, a collection of real cases with agreed judgements about the right output, forces those questions into the open. It is uncomfortable, because it exposes disagreements about how work should be done. It is also essential. Without it, quality is judged by anecdote: someone saw one bad answer and lost confidence, or someone saw one good answer and assumed the rest are fine. With it, the team can measure whether a change made things better, and have an honest conversation about whether the system is good enough for a given use.

People have to change how they work

Introducing an AI system changes jobs, even when it changes them only slightly. The support agent who used to write replies now edits drafts. The analyst who used to read every document now reviews the ones the system flagged. These are real changes, and people respond to them in real ways: with relief, with suspicion, with worry about what comes next.

Trust is earned slowly and lost quickly. A system that is right on routine cases, visibly uncertain on hard ones, and honest about its limitations builds trust over weeks. A system that is confidently wrong in front of an important customer can lose it in an afternoon. This argues for deployments that start narrow, with the cases where the system is most reliable, and widen as confidence grows. It also argues for involving the people who will use the system in defining its evaluation, because they know which errors matter.

Who owns it on Monday?

Once an AI system is in production, it needs the same things any production system needs, plus a few that are new. Someone must own the prompts and configuration and be accountable for changing them. Someone must own the evaluation set and keep it representative as the work evolves. Someone must decide when a new model version is adopted, after checking that it performs at least as well on the cases that matter. Someone must respond when the system produces a harmful or embarrassing output, and be able to reconstruct what happened.

In many organisations, none of these responsibilities has an obvious home. The data team built the pipeline, the IT team runs the infrastructure, the business unit uses the tool, and a vendor supplied the model. When something goes wrong, each can reasonably say it is someone else's problem. Settling ownership before launch is less exciting than any technical decision in the project, and more important than most of them.

The pilot that succeeded and then didn't

There is one more failure pattern worth naming, because it is common and demoralising. A pilot succeeds. Metrics look good, users are positive, and the system is rolled out more widely. Then quality drops.

Often the reason is that the pilot was quietly supported by manual effort. Someone cleaned the input data by hand. Someone reviewed edge cases before users saw them. Someone adjusted the prompt every few days in response to feedback. None of that scales, and none of it was visible in the metrics. When the system meets the full variety of real work without that hidden support, it performs as it would have all along.

The defence is to ask, before declaring a pilot successful, what human effort went into it and whether that effort will exist at scale. If it will not, the pilot has not yet shown what it appears to show.

Where model choice does matter

None of this means the model is irrelevant. For some tasks the differences between models are material: reasoning over long documents, following complex instructions, working in languages other than English, or running within tight latency and cost limits. Cost in particular deserves attention, because the most capable model is rarely the right default for high-volume, well-defined tasks where a smaller one performs acceptably.

But these are questions best answered with an evaluation set in hand, by testing candidates against real cases. Asked in the abstract, at the start of a project, they tend to absorb attention that the harder questions need.

Questions worth settling first

Before choosing a model, it is worth being able to answer a short list of questions. What are the authoritative sources, and who keeps them current? Whose access rights does the system inherit, and how are they enforced? Where in the existing workflow will the output appear? What does a good output look like, and who has agreed that definition? Who owns the system after launch? What happens when it is wrong?

Organisations that can answer those questions usually find the model decision straightforward. Organisations that cannot will find that no model, however capable, makes up for the gap. That is the approach behind our AI & Automation work: start with the system the model has to live in, and choose the model once the system is understood. For the engineering view of the same problem, see AI as a layer in software.

Filed under AI & Automation · Vianmax Editorial ·

Examples in this article are general and conceptual unless stated otherwise. Where Vianmax products or work are mentioned, the description matches what is published elsewhere on this site.

Keep reading

Product & Digital Systems

When a Product Needs Fewer Features

Every feature is paid for twice: once when it is built, and then indefinitely, mostly by people who never asked for it. Mature products get better by subtraction more often than their roadmaps admit.

7 min read

From pilot to production

We design AI systems around the data, permissions and workflows they have to live in, with evaluation and human review built in from the start.