Talk to Us +91 98847 45599

AI & Automation

AI Is Becoming a Layer in Software, Not the Software Itself

The model is the one part of the system whose output you cannot specify in advance. Good AI engineering is mostly about everything built around it.

Vianmax Editorial7 min read

Technical illustration of a cross-section of stacked deterministic layers, drawn as precise rectangles, with one band of scattered points in the middle. A blue path passes through every layer.

For a while, the dominant picture of an AI product was a chat window. You typed, a model answered, and the model was the product. That picture is changing, and not because chat is going away. It is changing because the teams putting language models into systems people rely on keep arriving at the same conclusion: the model is one component among many, and usually not the one that decides whether the system works.

The more useful mental model is a layer. A model sits inside ordinary software, which prepares its inputs, constrains what it can do, checks what it produces and decides what happens next. The model contributes something conventional code cannot: reading unstructured text, interpreting intent, drafting language, classifying things that resist rules. The surrounding code contributes everything else, including most of the reliability.

A probabilistic part in a deterministic machine

Traditional software is, for practical purposes, deterministic. The same input produces the same output, and when it doesn't, that is a bug you can reproduce. Language models are different. Their output varies with phrasing, with context, and sometimes between identical calls. They are extremely capable within that variability, but they are not specifiable in the way a function is.

This has a direct engineering consequence. You cannot test a model into correctness the way you test a calculation. What you can do is design the contract around it: define what it receives, define the shape of what it must return, and check every response against that shape before anything downstream relies on it.

In practice, that often means asking the model for structured output rather than free text, validating the structure strictly, and treating anything that fails validation as a failure to be handled, not a response to be passed along. It means deciding in advance what happens when the model is uncertain or wrong, rather than discovering it in production. The model becomes a component with a known interface and known failure modes, which is something engineers know how to work with.

RequestPermissions& policycodeRetrievalcodeModelprobabilisticmodelValidationcodeActlow riskPersonreviewLog inputs, context, versions and outcomes · evaluate against a curated set
The model is one stage in a deterministic pipeline. Permissions, retrieval and validation are ordinary code; routing by risk decides where a person reviews.

Retrieval changes the question

Much of the anxiety about AI reliability comes down to one worry: the model might say something that isn't true. It will state it fluently and confidently, and nothing in the output will tell you it is invented.

Retrieval does not eliminate that risk, but it changes its character. When a system retrieves relevant documents and asks the model to answer from them, the question shifts from "what does the model know?", which is opaque, to "what did we give it, and did it use it faithfully?", which can be inspected. You can log the retrieved passages. You can check whether the answer is supported by them. You can ask the model to cite which passage supports each claim, and verify that the citation exists.

This is also where much of the real work sits. Retrieval quality depends on how documents are split, indexed and ranked, on whether stale versions are excluded, and on whether the system can tell the difference between a policy document and an email that quotes it. A mediocre model with good retrieval often outperforms a strong model with poor retrieval. The unglamorous parts of the pipeline carry more weight than the model choice.

Tools, and who holds the keys

The next step beyond answering questions is acting: looking up an order, creating a ticket, drafting and sending a reply, updating a record. Models can now call tools, and that is where the layered view matters most.

A principle we hold to firmly: permissions belong to the tool, not the model. The model can request an action. Whether that action is allowed must be decided by conventional code that knows who the user is, what they are entitled to do, and what this particular tool may touch. If a user cannot see another customer's records, an assistant acting on their behalf must not be able to either, regardless of how the request is phrased. Relying on instructions in a prompt to enforce access control is relying on the one component you have already accepted is unpredictable.

The same logic applies to consequences. Tools that only read are relatively safe. Tools that write, send or spend deserve additional gates: confirmation from a person, limits on volume or value, or a queue where actions wait for review. The model proposes; the system disposes.

Relying on a prompt to enforce access control is relying on the one component you have already accepted is unpredictable.

Hallucination is managed, not solved

It is tempting to treat hallucination as a defect that the next model release will fix. Models do keep improving, but the property that makes them useful, generating plausible language from patterns, is closely related to the property that makes them occasionally wrong. Engineering around it is more productive than waiting for it to disappear.

The techniques are well understood, if not always applied. Constrain answers to supplied sources where accuracy matters. Make "I don't know" or "this isn't in the documents" an acceptable, even rewarded, outcome rather than a failure. Verify claims that can be verified: if the model extracts an invoice total, check it against the line items. Show sources to the user so they can judge for themselves. And be honest in the interface about what the system is doing, so that people calibrate their trust to its actual reliability.

Evaluation is a test suite that speaks in percentages

Conventional tests pass or fail. Evaluating a model-based component is statistical: across a representative set of inputs, how often is the output correct, how often is it acceptable, how often is it harmful? That requires something many teams skip, a curated evaluation set drawn from real cases, with agreed judgements about what good looks like.

Without one, every change is a guess. A prompt revision that fixes one reported problem may quietly break three others. A new model version may be better on average and worse on exactly the cases your users care about. With an evaluation set, those changes become measurable, and model and prompt updates can be treated like any other release: run the suite, compare, decide.

Building the set is slow and slightly tedious, and it forces conversations about quality that people would rather avoid. It is also the single most valuable asset in most AI systems, more durable than any particular prompt or model.

Where the person sits

"Human in the loop" is often stated as a safeguard and left there. It deserves more design than that. Where exactly does the person sit? Reviewing every output, or only the ones the system flags as uncertain? Approving actions before they happen, or auditing them afterwards? Seeing the model's sources, or only its conclusion?

These choices have costs in both directions. Reviewing everything is safe but can make the system slower than doing the work by hand, and reviewers who approve hundreds of correct outputs in a row stop reading carefully. Reviewing nothing is fast and occasionally disastrous. The useful middle ground is to route by risk and confidence: let routine, reversible, well-evidenced outputs through; send ambiguous or consequential ones to a person, with the context they need to judge quickly.

That is the pattern behind the kind of systems we build in AI & Automation: intelligence applied where it improves a decision, with a person in the loop where judgement matters. The emphasis on where is deliberate.

Where conventional software is simply better

A layered view also makes it easier to say something that enthusiasm tends to obscure: many problems should not involve a model at all.

Arithmetic should be done by code. So should anything governed by explicit rules with legal or financial weight: tax calculations, margin requirements, eligibility criteria, access control. So should anything that needs to be exactly reproducible for an audit. A model can help a person understand a rule, draft the code that implements it, or extract the inputs from a messy document. It should not be the thing that applies the rule, because "usually right" is not an acceptable standard for those tasks.

The best systems we have seen use models narrowly and precisely: to turn unstructured input into structured data, to interpret what someone is asking for, to draft language for a person to approve. Then they hand off to deterministic code for everything that must be exact.

Seeing what happened

Finally, a layered system needs to be observable in ways traditional systems do not. When a model-based feature produces a bad result, "the AI got it wrong" is not a diagnosis. You need to know what input it received, what context was retrieved, which prompt and model version were used, what it returned, and what the validation layer did with it. Without that record, problems cannot be reproduced, and what cannot be reproduced cannot be fixed with any confidence.

Logging all of this raises its own questions about privacy and retention, which need deliberate answers. It is also the kind of decision worth recording explicitly, in the way we describe in a piece on architecture as decision-making. But the principle holds: if you cannot reconstruct why the system did something, you do not really control it.

A component, not a character

Treating AI as a layer is not a way of diminishing it. The capabilities are real, and they open up problems that were impractical to automate a few years ago. It is a way of putting those capabilities where they can be relied upon: inside systems with clear contracts, enforced permissions, measured quality and a defined role for human judgement.

The teams that get lasting value from AI tend to talk less about the model and more about the system. That is usually a good sign. The organisational side of the same argument, covering data, ownership and trust, is the subject of a companion piece on enterprise AI.

Filed under AI & Automation · Vianmax Editorial ·

Examples in this article are general and conceptual unless stated otherwise. Where Vianmax products or work are mentioned, the description matches what is published elsewhere on this site.

Keep reading

AI & Automation

The Hard Part of Enterprise AI Is Not the Model

Choosing a model is the most discussed decision in an enterprise AI project and, increasingly, the easiest to reverse. The decisions that determine success are organisational.

7 min read

Learning & Future of Work

The New Learning Stack: People, Platforms and Intelligent Tools

Institutions tend to buy a platform to solve a teaching problem, or an AI tool to solve a content problem. The learning stack works when each layer does its own job and the layers share what they know.

7 min read

AI with a person in the loop

We build grounded assistants, document intelligence and workflow automation that are designed to be reliable in operation, not just impressive in a demo.