A shopper asks a retail product assistant whether a sale item can be returned. The assistant quotes last season’s return window; the order goes through, and the error surfaces only when the return is refused. No dashboard flagged the answer. No alert was fired. An LLM evaluation framework is the set of test questions, scoring rules, and release thresholds that catch a failure like this before a shopper does. Without one, the first warning is a user’s complaint.
The gap between demo confidence and production reality is not a one-team problem. In LangChain’s State of Agent Engineering survey of 1,340 practitioners, fielded in November and December 2025, quality was the top barrier to putting agents into production, named by 32% of respondents. The same survey found that 89% had observability in place, but only 37% ran evaluations on live production traffic. Teams can tell when the system is down. Only about a third score whether its answers are right.
Key takeaways
- Build the eval set from real questions users asked in production, including the ones the system answered wrongly.
- Track five metric groups: accuracy, groundedness, safety, latency, and cost.
- Record a baseline score for each version of the model, prompt, and retrieval index, and block releases that fall below it.
- Run a regression check on every change and score a sample of live traffic every week.
- Name an owner for the eval set and for each threshold, or the scores are reported and ignored.
What counts as an LLM evaluation framework?
An LLM evaluation framework is a repeatable way to score a language model feature on the tasks it performs for your users. It connects four parts into a single loop: an eval set of real inputs paired with correct answers or judging rules, the metrics each answer is scored on, a scoring method that can be a code check, an LLM-as-a-judge, or a human reviewer, and the thresholds that decide whether a version can ship. Remove any one part and the loop breaks. A set without thresholds produces scores nobody acts on. Thresholds without a set test: nothing real. Each part needs to be written down, so the same test runs the same way every time.
Teams also call this an LLM testing framework. Open-source tools such as DeepEval describe themselves as an LLM evaluation framework, but a tool only runs the tests. The team still must decide what to test, how to score it, and which result blocks a release. Evals for LLM applications differ from GenAI model evaluation on public benchmarks for the same reason: a benchmark score shows how a model performs general tasks, not whether your assistant quotes your current return policy correctly.
An LLM Evaluation Framework That Tests What Demos Hide
Gartner expects companies to cancel over 40% of agentic AI projects by the end of 2027. Weak risk controls are one reason. ViitorCloud’s GenAI model evaluation tests your agent against messy, real-world user behavior. You find the breaking points while they’re still cheap to fix.
Why does an LLM that passed every demo fail with production users?
A demo runs on clean questions the builders chose, while production brings conditions nobody rehearsed. Each of the following can lower answer quality without raising an error, so monitoring dashboards show no problem while users receive worse answers.
The first shift is the inputs themselves. Real users write short, misspelled, or mixed-language questions and ask about edge cases the demo never covered. Those unfamiliar inputs hit a second problem: changing data. In retail, new products, seasonal promotions, and price changes alter the documents the assistant retrieves, and when retrieval returns stale or wrong passages, the answer is wrong even if the model is working as expected. Meanwhile, model providers update and retire model versions on their own schedule, and a small prompt edit can shift behavior in cases nobody re-tested. Layer on multi-turn chats, where earlier context can pull later answers off track, and answers that were fast and cheap in a demo can slow down or cost more at peak traffic such as a holiday sale. Each layer compounds the last, and none of them throws an exception.
NIST’s AI Risk Management Framework (AI RMF 1.0, January 2023) speaks to this gap directly: its MEASURE 2.3 subcategory says AI system performance should be measured and demonstrated for conditions like deployment settings. The framework is voluntary, and NIST states it is being revised under the White House AI Action Plan. The principle still applies to any team, and an LLM evaluation framework puts it into practice by testing real production inputs.
The five metric groups an LLM evaluation framework can’t skip
An LLM evaluation framework should track five groups of metrics: accuracy, groundedness, safety, latency, and cost. Accuracy and groundedness show whether answers are right and supported by your data. Safety shows whether the system avoids harmful content and data leaks. Latency and cost show whether a correct answer arrives fast enough and cheaply enough to keep in production.
| Metric group | What it checks | How to score it | Related reference |
| Accuracy | The answer matches the expected answer or meets the rubric | Code checks for single right values; LLM-as-a-judge for open answers | NIST AI RMF: valid and reliable |
| Groundedness | Each claim in the answer is supported by the retrieved source | Judge model compares the answer with the retrieved passages; human spot checks | NIST AI 600-1 risk: confabulation |
| Safety | No harmful content, no private data, a refusal when the request is out of scope | Rule filters plus red-team cases kept in the eval set | NIST AI RMF: safe; AI 600-1 risks: data privacy, information security |
| Latency | Time to a complete answer, including the slowest 5% of requests (p95) | Measured on every request | No NIST mapping; operational metric |
| Cost | Cost per answer or per conversation | Token usage multiplied by price, tracked per route | No NIST mapping; operational metric |
Of these five, groundedness deserves the closest attention in any feature that answers your own data. NIST’s Generative AI Profile (NIST AI 600-1, July 2024) names this risk confabulation and describes it as “confidently stated but erroneous or false content.” Shoppers usually cannot tell when a confident answer is wrong, which is exactly why LLM-as-a-judge has become a standard way to score these LLM evaluation metrics. In the LangChain survey, 53% of respondents used a judge model and about 60% used human reviews. A judge model is only useful after you check it against human scores on a sample; if the judge and your reviewers often disagree, fix the rubric before you trust the judge’s numbers.
For agents, the final answer is not enough to score. Whether the agent called the right tool with the right input, and whether it stopped when it should have, matters as much as the output. Step scores depend on tracing, and AI agent observability supplies the step-level records they need.
Five steps to a baseline eval suite that holds up in production
A first suite for one feature can be small. What matters is that it reflects real use and runs the same way every time. The baseline eval suite is the core of an LLM evaluation framework, and building one follows five steps.
Collect real questions
Evals for LLM applications should start with production logs, support tickets, and chat transcripts. Cover each task the feature handles and add every question it has answered wrongly. For a retail assistant, that means order status, returns, sizing, stock, and promotions.
Write the expected result
For each case, record the correct answer or a rubric, which is the list of criteria a good answer must meet. A domain owner signs these off; in retail, that is often a merchandising or customer service lead.
Choose a scoring method for each metric
Use code checks where an answer has one right value, such as a price or a date. Use a calibrated judge model for open answers, plus a human review of a sample in each cycle.
Score the current version to set a baseline
Record the scores against the exact model version, prompt version, and retrieval index version. NIST’s Generative AI Profile suggests a similar step: considering baseline performance on benchmarks before fine-tuning a model or adding retrieval-augmented generation (action MS-2.3-001).
Agree on release thresholds
Decide in writing what blocks a release. For example: block a deploy if accuracy falls more than two points below baseline, if any safety case fails, or if p95 latency or cost per conversation rises above the agreed budget. The numbers are yours to set. Write them down before the next release, so the decision is already made when the scores arrive.
LLM Evaluation Metrics That Catch Problems Before Tickets Do
We score every AI agent on groundedness, task success, tool-call accuracy, and latency. Each release runs against the same scorecard. When quality slips, your team sees it first.
The cadence that keeps an LLM evaluation framework from going stale
Run the eval suite on every change to the model, prompt, or retrieval setup, and score a sample of live traffic every week. Refresh the eval set each month with new failures, and review thresholds each quarter. This cadence keeps the LLM evaluation framework running as LLM quality monitoring long after launch day.
- On every change (offline regression runs). A model version change, prompt edit, retrieval setting, or index rebuild triggers the full suite in your CI pipeline, the automated process that builds and tests each change. A failed threshold blocks the deployment. The engineering lead who owns this pipeline is the person accountable for keeping it green.
- Weekly (online evaluation). Score a random sample of real conversations, plus every conversation a user flagged, with the same metrics. This is the check teams skip most often: in the LangChain survey, 52% ran offline evaluations but only 37% evaluated production traffic.
- When a provider updates a model. Re-run the suite before switching, even when the provider describes the update as minor. The provider’s own GenAI model evaluation on public benchmarks does not cover your cases.
- Monthly. Move new failure cases from production into the eval set, so the suite grows with real use. A domain owner, such as merchandising or support to lead in retail, owns the expected answers in these new cases.
- Quarterly. A product owner checks that thresholds still match the business risk, for example, before a peak sales season, and decides on any exception.
An LLM evaluation framework without named owners at each of these levels produces scores that sit in a dashboard and never block a release.
What six weeks of live traffic taught one retail team
This example is illustrative and does not describe a specific client:
Two problems hid inside a retail product assistant for weeks, and neither one threw an error. The team had demoted the assistant on 30 hand-picked questions about products in stock. Every answer came back correctly. They also built an eval suite of 150 real questions drawn from six months of support tickets, scored against thresholds for accuracy, groundedness, latency, and cost. That second decision is what saved them.
The first problem surfaced in week three. A holiday promotion changed the return window, but the retrieval index still served last season’s policy page. The monitoring dashboard stayed green. A stale-but-confident answer looks exactly like a correct one to a logging tool. The eval suite, however, scored the answer against the current return window a human had written into the expected result. Groundedness dropped, the alert fired, and the team refreshed the index before more shoppers received the wrong answer. Those failing questions now sit permanently in the eval set, so the same gap cannot reopen quietly.
The second problem appeared in week six. An engineer shortened the system prompt to cut costs per conversation. Cost fell as planned, but sizing accuracy fell with it. The regression run caught the trade-off and blocked the deployment. The engineer revised the prompt, re-ran the suite, and shipped only after sizing accuracy cleared the threshold.
Neither catch came from the demo set, because the demo set tested only the questions the builders already expected. Every real failure came from questions real shoppers asked.
Where regressions slip through when evaluation stops short
Count how many of these habits apply to your team. Each one is a gap that lets a regression reach users without any alarm firing.
- Treating observability as evaluation. Traces show what happened and how long it took. They do not know whether the answer was correct.
- Testing only on demo questions. An LLM testing framework built on the builders’ own questions checks only the cases they already expected, which is why AI demos that looked strong can stall after launch.
- A golden set that never changes. When production failures never enter the eval set, the suite drifts away from real use.
- An unchecked judge model. LLM-as-a-judge scores are only as good as the rubric and the calibration against human reviewers.
- Ignoring latency and cost. A correct answer that arrives too slowly, or costs too much at peak load, still fails in the business case.
- Evaluating once at launch. Model, data, and prompt changes continue after launch, so the evaluation must continue too.
If three or more apply, the gap between what the team measures and what users experience is wider than any dashboard will show.
Why ViitorCloud locks the eval set before choosing the architecture
Most teams treat evaluation as a final gate. In the LLM application development approach ViitorCloud follows, the order is reversed: the eval set is the first deliverable, not the last. Real production questions define what “correct” means before a single design decision is made, which means the first prototype is already scored against real user questions, not builder-written demo questions.
That early start changes what happens during the build. Thresholds are agreed with the domain owner while the feature is still taking shape, and the suite runs in CI on every release. A deploy where accuracy, groundedness, or cost moves in the wrong direction is blocked automatically, not reviewed in a meeting after the fact. After launching, the cadence continues: weekly scoring of live traffic, monthly additions from new failures, and a fresh baseline before each peak season.
For a retail assistant delivered through custom AI solutions work, that translates to eval cases drawn from actual shopper questions about stock, returns, and promotions, a calibrated judge model checked against human reviewers, and a team that knows where quality stands before traffic spikes.
Your Users Run the Real Test. Ship Agents That Pass It.
ViitorCloud’s AI Agents Development team builds an LLM testing framework into every agent from the first sprint. Edge cases, regression checks, and human review run before any update ships. You go live with proof it holds up.
Start small: one baseline, one score to watch
Start with 50 real questions and a baseline score. That is enough to know, next week, whether answers improved or slipped after a model provider update. To build a full LLM evaluation framework around a feature already in production, talk to the ViitorCloud team.
Vishal Shukla
Vishal Shukla is Vice President of Technology at ViitorCloud Technologies.