← All notes

RAGA Framework, Part 7: Beyond RAG: Evaluating AI Agents with Ragas

RAGA Framework, Part 7: Beyond RAG: Evaluating AI Agents with Ragas

Part 7 of a series on evaluating LLM apps with Ragas.

Article content

infographic - AI Agent Evals using RAGA

Everything so far has been single-shot: a question comes in, we retrieve, we answer, we score the answer. But our policy assistant is outgrowing that. Employees started asking things it couldn't answer from documents alone, like "I've been here 3 years and have 12 unused vacation days, how many can I carry over and what's my payout if I don't?" That needs a calculation, not a paragraph. So we gave it a tool: a leave_calculator, plus the existing policy_search. Now it's an agent, and agents fail in ways a single answer never could.

Why agent evaluation is a different problem

When a system takes multiple steps, "was the final answer good?" is no longer enough. The agent can reach a right answer through a broken process, or fail halfway with the right idea. Between input and output there's now a trajectory, and any point on it can go wrong:

  • It picks the wrong tool, or no tool when it needed one.
  • It picks the right tool but passes bad arguments.
  • It calls tools in an order that doesn't make sense.
  • A tool errors and the agent doesn't recover.
  • It wanders off-topic or answers something it shouldn't.
  • The final task simply isn't accomplished.

You have to evaluate two different things: the journey (did it do the right steps?) and the destination (did the user get what they wanted?). Ragas has metrics for both.

The building blocks: messages

Agent metrics in Ragas work over a conversation represented as typed messages, not a flat string. That structure is what lets the metrics reason about tool calls.

from ragas.messages import HumanMessage, AIMessage, ToolMessage, ToolCall 

A HumanMessage is user input, an AIMessage is the agent talking (and may carry tool_calls), a ToolMessage is what a tool returned, and a ToolCall names a tool and its arguments. A logged agent run becomes a list of these.

Tool call accuracy: did it use tools correctly?

ToolCallAccuracy checks whether the agent called the tools you expected, with the right arguments, in the right order. It doesn't need an LLM judge, it compares against a reference list of tool calls, which makes it cheap and deterministic, ideal for regression tests.

import asyncio
from ragas.metrics.collections import ToolCallAccuracy
from ragas.messages import HumanMessage, AIMessage, ToolCall

async def main():
    user_input = [
        HumanMessage(content="I have 12 unused vacation days after 3 years. What can I carry over?"),
        AIMessage(
            content="Let me check the carryover policy and run the numbers.",
            tool_calls=[ToolCall(name="policy_search", args={"query": "vacation carryover limit"})],
        ),
        AIMessage(
            content="Now calculating your carryover.",
            tool_calls=[ToolCall(name="leave_calculator", args={"unused_days": 12, "tenure_years": 3})],
        ),
    ]
    reference_tool_calls = [
        ToolCall(name="policy_search", args={"query": "vacation carryover limit"}),
        ToolCall(name="leave_calculator", args={"unused_days": 12, "tenure_years": 3}),
    ]
    metric = ToolCallAccuracy()
    result = await metric.ascore(user_input=user_input, reference_tool_calls=reference_tool_calls)
    print(result.value)

asyncio.run(main()) 

The scoring is stricter than people expect, and that's worth understanding. It combines argument accuracy with sequence alignment, and the sequence part is effectively a gate: if the order is wrong in strict mode, the score is 0 even if every tool and argument is right. If order genuinely doesn't matter (fetching weather for three cities in parallel), set strict_order=False. If a call has three arguments and one is wrong, you get partial credit on the arguments, roughly two-thirds.

Tool call F1: a softer view

ToolCallAccuracy is binary about sequence, which is harsh while you're still iterating. ToolCallF1 gives a gentler read.

First, what F1 is, in case it's new: it's a single score (0 to 1) that balances two things.

  • Recall = did the agent make all the calls it needed? Miss a required tool call and recall drops.
  • Precision = did the agent avoid unwanted extra calls? Every unnecessary call it makes drops precision.

F1 combines both into one number, so you can't game it by only optimizing one, an agent has to be both complete (recall) and clean (precision) to score well. It's a standard metric borrowed from classification, applied here to tool calls.

ToolCallF1 computes exactly that over the agent's tool calls, matching each call on both name and arguments (wrong arguments count as a miss) and ignoring the order. So if the agent made every expected call plus one unnecessary extra, ToolCallAccuracy might read harshly, but F1 shows perfect recall with a precision dip, telling you the agent is close, not broken. Use F1 during development to track "are we getting warmer," and ToolCallAccuracy for a strict pass/fail gate before shipping.

Agent goal accuracy: did the user actually get what they wanted?

Sometimes you don't care which tools were used, only whether the goal was met. "Book me a table" is a success if a table is booked, regardless of the path. AgentGoalAccuracy captures that. It's a binary metric, 1 if the goal was achieved and 0 if not, and it comes in two flavors:

  • AgentGoalAccuracyWithReference compares the end state to an outcome you specify.
  • AgentGoalAccuracyWithoutReference infers the user's goal from the conversation and checks whether it was met, no reference needed.
from ragas.metrics.collections import AgentGoalAccuracyWithReference
# ... build user_input as a message list of the full conversation ...
metric = AgentGoalAccuracyWithReference(llm=llm)
result = await metric.ascore(
    user_input=user_input,
    reference="Correct carryover (5 days) and payout for the remaining 7 days reported to the employee",
) 

Tool-call metrics and goal accuracy answer different questions, and choosing between them is itself a design decision. Tool-call accuracy tells you the process was right. Goal accuracy tells you the outcome was right. You usually want both, because a system can hit the goal by luck through a messy process (fragile, will break later) or follow a clean process to a wrong goal (a logic bug downstream of the tools).

Topic adherence: did it stay in its lane?

Our policy assistant should talk about HR policy, not restaurant recommendations. TopicAdherence checks whether the agent stuck to allowed topics across a conversation. You pass reference_topics and choose a mode: precision (of the things it engaged with, how many were on-topic) or recall (of the on-topic things it should have handled, how many it did). For a scoped internal assistant, this is your guardrail metric against the model cheerfully answering things it has no business answering.

Choosing a simple agent to practice on

If you want to build intuition, pick something small with an obvious success condition:

  • Expense categorization
  • A travel-policy assistant
  • A calculator agent
  • A customer-support workflow
  • A document-search agent

Our leave-calculator policy assistant is deliberately in this bucket: two tools, a clear right answer, easy to construct reference tool calls and reference outcomes for.

Interview angle

Agent evaluation is where forward-deployed and AI engineering interviews are heading, because agents are what clients now want built.

The question: "How is evaluating an agent different from evaluating a RAG pipeline?"

The answer: a RAG pipeline is single-shot — one question, one retrieval, one answer — so you score a single output. An agent runs a multi-step trajectory: it chooses tools, fills in their arguments, calls them in some order, reacts to the results, and recovers (or fails to) before producing a final outcome. That changes evaluation in two ways:

  • You score the journey, not just the destination. Tool selection, arguments, order, and error recovery each need checking, because any one can go wrong even when the final answer looks fine.
  • Process-correct and outcome-correct are different things. An agent can follow a clean tool sequence to a wrong result, or fumble the steps and still land the right answer. So you measure both: tool-call metrics for the trajectory, goal accuracy for the outcome.

The line that lands it: "I use tool-call accuracy for the process and goal accuracy for the result, because a system can get one right and the other wrong."

The takeaway

An agent has a trajectory, not just an output, so you evaluate both. Ragas gives you tool-call accuracy and F1 for the process, goal accuracy for the outcome, and topic adherence for staying in scope. Deterministic metrics make excellent regression tests; LLM-judged ones cover the fuzzier goal.

We now have all the pieces: metrics, test data, experiments, judge validation, agents. The last article ties them into something that runs continuously, so evaluation isn't a thing you do once in a notebook but a gate that protects every change you ship.

Originally published on LinkedIn.