Why successful RAG systems are built on sound architecture rather than framework choices, starting with the fundamentals that determine long-term scalability and reliability.
When should enterprises route requests between frontier and smaller models? This article examines routing strategies, latency, cost, and quality trade-offs.
Starting at the Base Metal: Designing Model Routing for the Unhappy Path
What happens when a request looks simple, gets routed to the small model, and turns out it needed more? A two-layer escalation pattern — checking at dispatch and again before release — catches what a one-pass system misses.
Why benchmark scores aren't enough and how continuous evaluation improves enterprise AI reliability.
When does model monitoring become continuous evaluation?
Why continuous evaluation isn't just a technical capability — it's an organizational one, and how better collaboration between tech and business teams is the next evolution of AI operations.
How a dedicated intake layer — lightweight classifiers, guided UIs, or external agents — translates human intent into exact API calls, letting code resolve the data and the model reason over the result.
What happens when a routed answer turns out to be wrong? You need enough logging and tracing to reconstruct the decision — and the ability to fix the logic that produced it.
Good AI solutioning starts with understanding a client's pain points, appetite for change, and real-world constraints such as budget, regulatory compliance, data readiness, and competitive positioning.
I've been helping a medical practice think through how AI can reduce day-to-day operational friction. While prototyping an internal chatbot, I went to the base metal: mapping what information each prompt actually needs and which data source the system should use to answer it.
Beyond prompt engineering, the system has to handle actual operational constraints around data access, audit, provenance, and guardrails.
The biggest mistake in enterprise AI is treating all data the same.
In reality, the data dictates the path:
🔹 Structured data requires exact retrieval from a source of truth.
🔹 Unstructured knowledge requires semantic retrieval based on meaning and context.
An LLM should never be allowed to infer concrete, fixed information such as an appointment date or a medication list. Semantic RAG is not the solution when the answer lives as structured data.
The real architecture isn't the chatbot interface. It is the routing layer underneath: deciding when to use a structured query, when to use semantic retrieval, when to synthesize, and when to escalate instead of faking confidence.
#AIArchitecture #EnterpriseAI #HealthTech #SolutionsArchitect #SystemDesign

The most capable model isn't always the best model. Healthcare just happens to make that painfully obvious.
In a clinical setting, one request might be: "Pull the payer's prior authorization criteria for this medication." The next might be: "Read this 15-page faxed discharge summary, extract the failed step therapies, align them to payer criteria, and draft a prior authorization."
Those requests have fundamentally different computational needs. Treat them the same, and the tradeoffs show up immediately. Route everything to a frontier model and you incur unnecessary latency, token consumption, inference cost, and lower throughput. Route everything to a smaller model and you risk degraded accuracy, missed clinical context, and failures on the requests that actually require deeper reasoning.
The answer isn't a better model. It's better architecture.
A deterministic routing layer evaluates complexity, intent, and modality before a prompt is processed, then sends the request to the appropriate engine. Straightforward retrieval goes to a smaller model paired with deterministic systems like APIs or vector search, where you don't need an agent improvising against your schema. Complex multimodal reasoning goes to a frontier model, but only when the task actually warrants it.
The hard part isn't the routing logic. It's understanding the workload well enough to classify requests so you can optimize for quality, latency, throughput, and cost simultaneously. Going to the base metal means starting with the workload to identify its complexity, modalities, failure modes, and performance requirements and letting those characteristics drive the architecture.
These are the kinds of tradeoffs I enjoy thinking about. The interesting engineering challenge isn't finding the model that tops the latest SWE-bench leaderboard. It's building an architecture that consistently delivers the best outcome within real-world constraints.

Starting at the Base Metal: Designing Model Routing for the Unhappy Path
One of my older post argued for routing clinical AI tasks by complexity, sending straightforward retrieval to smaller models and complex reasoning to frontier ones. An excellent follow-up question came up: what happens when a request looks simple, gets routed to the small model, and turns out it needed the more complex reasoning after all? In a one-pass system, the answer has gone out by the time you would catch it.
My routing layer is deterministic by design, which makes the dispatch decision reproducible and explainable, but, by its very nature, it only recognizes the patterns you have encoded.
At dispatch, anything that doesn't clearly match a known pattern gets escalated rather than answered. That still misses the harder case, a request that matches a simple pattern cleanly and needed more, so a second check runs after the fast path and before anything returns: empty or off-topic retrieval, a failed validation, a tripped guardrail, any of which pulls the task back to escalate instead of release.
Escalation isn't always a larger model. Sometimes it's a deterministic lookup, a human, or declining to answer. When the system isn't confident it has the right path, it defers rather than forcing a decision it can't stand behind. That lowers the failure rate without driving it to zero, so what slips through gets logged and reviewed.
Some requests won't match a rule, and some will match the wrong one. That's a fact of the workload, so handling it belongs in the design from the start, not bolted on after the happy path works. That's what starting from the base metal means to me. Updated routing graphic below..
#AIArchitecture #HealthcareIT #EnterpriseAI

Over the next three posts, I want to go to the base metal of a question every organization deploying AI will eventually face: how do you know the models supporting your core workflows are still doing what you need them to do?
Every AI deployment follows the same arc: define evaluation criteria, test edge cases, validate performance, and run a pilot. The pilot succeeds, stakeholders sign off, and the model goes into production. And then, over time, production starts asking questions the pilot was never designed to answer. That's because many organizations make a critical category error.
Pre-deployment and production evaluation answer entirely different questions. A pilot proves that a model performs well on anticipated scenarios, much like testing a car on a closed track. It tells you very little about how that same model behaves in the "downtown traffic" of real users, evolving workflows, changing data, and unforeseen edge cases.
I saw this firsthand on a recent healthcare AI pilot. In healthcare, physician trust depends on having confidence that AI models continue to perform reliably after deployment. We built a rigorous pre-deployment evaluation framework using real patient data, and the model performed exceptionally well. But once a model enters production, the challenge shifts. Workflows and requirements evolve, user behavior changes, upstream systems change, and real-time ground truth isn't always available. Those are conditions a pilot simply cannot simulate.
Passing the pilot isn't the finish line for evaluating AI models. It's where continuous evaluation begins.
In Part 2, I'll explore what continuous evaluation should look like once an AI system is in production.
#ArtificialIntelligence #EnterpriseAI #AIArchitecture #AIModels #ModelEvaluation #ContinuousEvaluation #AIGovernance #HealthcareAI #AgenticAI #GenAI

In Part 1, I discussed one of the biggest mistakes organizations make: assuming that because an AI model performed well in pilot, it will continue performing well once it's deployed. If you missed it, you can read it here: https://lnkd.in/eFRfWedT
Production is where the real test begins.
Continuous evaluation is not just about monitoring a model. It is about detecting when something has changed and deciding what to do about it. A dashboard is reporting. Evaluation begins when the information leads to action.
In practice, that happens in three ways:
Each approach answers a different question. Drift tells you something may be changing. Review tells you what is actually going wrong. Validation gives you confidence that a new model is ready before it reaches users.
The goal isn't to monitor everything. It's to make sure what you detect leads to action. That's what separates reporting from continuous evaluation.
Part 3: Who should own continuous evaluation? Why it often falls between organizational silos, and why solving that ownership problem is becoming a competitive advantage.

In Part 1, I wrote about why passing a pilot doesn't predict production performance. In Part 2, I explored what to evaluate once a model is deployed, and asked who should own that work.
The more I've thought about it, the more I've wondered whether we've framed continuous evaluation too narrowly. Traditional software engineering assumes a functioning system keeps meeting its intended purpose until something changes. A bug is introduced, requirements evolve, an integration breaks, a release causes a regression.
AI models introduce a different reality. Their continued usefulness isn't something we can assume. It has to be continuously evaluated. That shift has implications most approaches miss. Most focus on monitoring: observability, drift detection, telemetry, dashboards. Those capabilities are essential. They're just not the whole story.
For all the criticism it gets, Agile changed one thing. It replaced long handoffs between business and engineering with continuous collaboration. Business wasn't involved only at the start and the end. It became part of the process.
I wonder if AI needs a similar evolution after deployment. Not another governance framework, not another review board, definitely not more meetings. What may be missing is an operating model that creates a continuous feedback loop between the teams monitoring model performance and the people who own the workflows those models are meant to improve.
Technical teams can tell us what changed. Business and clinical teams can tell us whether it matters. Neither perspective is complete on its own, which is why ownership cannot sit entirely with either at any given time. Engineering owns the signals. Business owns the judgment about whether the model still fits the workflow. AI systems work when both are accountable for the same outcome.
A model may drift. User behavior may change. A provider may update an underlying model. The workflow may evolve. Models can also hallucinate with remarkable confidence, which makes degradation harder to spot than a software defect. In high-impact domains like healthcare, those changes aren't only operational. They can affect patient care and whether clinicians continue to trust the system.
I'm starting to think continuous evaluation isn't just a technical capability. It's an organizational one. Monitoring technology is advancing rapidly. The harder and more valuable problem is designing how organizations learn from those signals and act on them.
Maybe that's the next evolution of AI operations. Not just better monitoring, but better collaboration.
How do your tech and business teams collaborate on continuous evaluation?
#ArtificialIntelligence #AIGovernance #EnterpriseAI #ResponsibleAI #MLOps #AIOperations #ContinuousEvaluation #HealthcareAI #AIStrategy

For one of my clients, I am designing an AI platform over decades of proprietary engineering data. We are dealing with a massive relational database, thousands of interdependent tables, air-gapped security, and output that must satisfy mandated standards exactly.
Work like this typically gets scoped as an end-to-end AI problem, and the natural instinct is to build it that way. I went to the base metal instead: Every question goes to the proprietary relational database first.
That is where the complexity lives. Answering one question can mean traversing a long chain of related records across tables designed decades apart. These relationships are already mapped in the database and are enforced.
That traversal is exact, repeatable work, and it belongs in code. A model asked to do it would produce a probable answer, and when there is exactly one correct answer, probable is not good enough.
But how does the system translate a vague human, NL prompt into the exact API calls needed to retrieve those records? Not by training users to be prompt engineers and definitely not letting the primary reasoning LLM guess the endpoint but rather by building a dedicated intake layer. Whether through a fast, lightweight classification model that extracts exact variables from unstructured text, guided UIs that parameterize the input, or external customer agents passing structured requests directly into the platform, this layer acts as a translator. It converts human intent into a strict, machine-readable query before any data is fetched.
What the database retrieves based on that API call is a perfectly resolved set of records. The primary LLM operates at this architectural layer.
The data itself is strictly structured, but its implications are not. The model is reserved for genuine reasoning: explaining what a set of related records implies, why a specific configuration exists, or what a chain of dependencies means for a higher-order decision.
The division of labor is settled in the architecture. Code resolves the data. The model reasons over the result. In highly constrained environments, abstracting the complexity before the model ever sees it is the only way AI adds significant value.

In a clinical AI system, routing checks reduce risk but they won't catch every case. So what happens when an answer goes out and is later found to be wrong?
You need enough logging and tracing to reconstruct the decision: what the request was, which routing rule matched, why it matched, what was retrieved, which checks passed, and what the system returned. Without that record, you know there was a failure, but you can't reliably identify where it happened.
Once you can trace the decision, you can fix the logic that produced it. The routing rule may have been too broad. A particular phrase or context may need to trigger escalation. Or the post-routing check may need to look for something it currently misses. You can then run the case through the revised workflow to confirm that it takes the right path.
The rules won't be perfect, especially in a clinical setting. Building in the ability to trace a wrong decision back to its source, correct the logic, and test the change is part of going down to the base metal.

#AIArchitecture #HealthcareIT #EnterpriseAI
AI Solutioning Perspectives