Testing AI Agents with Python: LLM Evaluation with DeepEval

Testing AI Agents with Python: LLM Evaluation with DeepEval

AI agent testing for QA and Python devs: RAG, tool calls, MCP, prompt injection, tracing, CI/CD gates. No API key.

Preview this Course

Testing AI Agents with Python: LLM Evaluation with DeepEval

Building an AI agent is only the beginning.

Once an agent can reason, call tools, retrieve information, and generate responses, a more difficult question appears:

How do you know whether the agent is actually working well?

Traditional software testing can verify whether a function returns the expected value. AI systems are different. Their outputs can vary, and an answer may be technically valid while still being incomplete, irrelevant, unsafe, or poorly grounded.

This is why LLM evaluation has become an important part of AI-agent development.

In this guide, we'll explore Testing AI Agents with Python: LLM Evaluation with DeepEval and look at how developers can build repeatable evaluation workflows for AI applications.

You'll learn what to test, how evaluation metrics work, how to create test cases, and how to use DeepEval as part of a Python-based evaluation workflow.

Why Testing AI Agents Is Different
Traditional applications usually have deterministic behavior.

For example:

def add(a, b):
    return a + b

You can test:

add(2, 3) == 5

AI agents are more complicated.

An agent may receive:

"Summarize the latest customer feedback and identify the three most important problems."

The response could vary between runs.

The output may also contain several dimensions that need evaluation:

Is it factually correct?

Did it answer the question?

Did it use the right information?

Did it follow instructions?

Did it hallucinate?

Did it use tools correctly?

Was the response sufficiently complete?

Did it provide useful reasoning or evidence?

This means AI-agent testing often requires evaluation criteria, not just exact output matching.

What Is DeepEval?
DeepEval is a Python-based evaluation framework designed for testing and evaluating LLM-powered applications.

It provides concepts such as:

Test cases

Evaluation metrics

Model-based evaluation

Dataset-based testing

Custom evaluation criteria

Evaluation pipelines

This makes it useful when you want to move from manually asking an AI application questions toward a repeatable testing process.

Instead of evaluating an agent once, you can create a collection of test cases and run evaluations repeatedly as your system changes.

Why Evaluate an AI Agent?
Imagine you improve your agent's system prompt.

The new version performs better on one test but worse on five others.

Without systematic evaluation, you might not notice the regression.

A repeatable evaluation suite can help answer questions such as:

Did the new prompt improve answer quality?

Did changing the model increase hallucinations?

Does the agent follow instructions consistently?

Is retrieval quality improving?

Are tool calls becoming more reliable?

Did a code change introduce regressions?

This is especially important when AI systems are continuously updated.

The Basic AI Evaluation Workflow
A simple evaluation pipeline looks like this:

Input → AI Agent → Output → Evaluation Metric → Score → Analysis

You provide an input to the agent.

The agent generates a response.

An evaluation metric analyzes the response.

The evaluation produces a score or judgment.

You then use those results to improve the system.

With multiple test cases, the workflow becomes:

Test Dataset → Agent → Outputs → Evaluation → Results → Regression Analysis

This is the foundation of an AI evaluation pipeline.

Step 1: Define What "Good" Means
Before writing evaluation code, define success.

This is one of the most important steps.

Suppose you're building a customer-support agent.

A good response might need to be:

Relevant

Helpful

Accurate

Grounded in company information

Concise

Consistent with support policies

These criteria can become evaluation dimensions.

Without clearly defined criteria, evaluation scores may not tell you much.

Step 2: Create Representative Test Cases
A good evaluation dataset should represent real usage.

For example:

test_cases = [
    {
        "input": "How do I reset my password?",
        "expected": "The response should explain the password reset process."
    },
    {
        "input": "Can I cancel my subscription?",
        "expected": "The response should accurately explain the cancellation policy."
    }
]

In practice, evaluation test cases can contain richer information, including the actual input, expected output, context, and metadata.

The important principle is to test realistic scenarios, not just easy examples.

Step 3: Install Your Evaluation Environment
A Python project can be configured with the relevant evaluation dependencies.

For example:

pip install deepeval

Then import the components needed for your tests.

A typical workflow may involve test cases and evaluation metrics.

The exact APIs and available metrics can evolve, so check the current DeepEval documentation when implementing a production evaluation suite.

Step 4: Create an Evaluation Test Case
A basic conceptual structure looks like this:

from deepeval.test_case import LLMTestCase

test_case = LLMTestCase(
    input="How do I reset my password?",
    actual_output="You can reset your password from the account settings page."
)

The test case represents one interaction that you want to evaluate.

Depending on your evaluation scenario, you may also provide additional information such as expected output or contextual data.

Step 5: Add an Evaluation Metric
The next step is deciding how the output should be judged.

For example, you may want to evaluate whether the answer is relevant.

Conceptually:

from deepeval.metrics import AnswerRelevancyMetric

metric = AnswerRelevancyMetric(
    threshold=0.7
)

You can then use the metric to evaluate the test case.

The exact metric configuration should match your application's requirements.

The important concept is:

Test Case + Metric = Evaluation

Important Metrics for AI-Agent Testing
Different applications require different evaluation criteria.

There is no single metric that can determine whether an AI agent is "good."

Answer Relevancy
Does the response actually answer the user's question?

An agent can generate grammatically perfect text that doesn't address the request.

Relevancy testing helps detect this problem.

Faithfulness
Is the generated answer supported by the information available to the system?

This is particularly important for retrieval-augmented generation systems.

Contextual Relevance
If your agent retrieves external context, does that context actually help answer the question?

Poor retrieval can lead to poor answers even when the language model itself is capable.

Contextual Precision
Does the system prioritize useful information rather than returning large amounts of irrelevant context?

This can be important for RAG-based agents.

Contextual Recall
Did the retrieval process obtain the information needed to answer the question?

Low recall can cause the agent to miss important evidence.

Hallucination
Does the agent generate unsupported or fabricated information?

Hallucination testing is especially important for systems that provide factual answers.

Task Completion
Did the agent actually accomplish the requested task?

For an agent that must call tools or complete a workflow, this can be more useful than evaluating the final text alone.

Testing an AI Agent Instead of Just an LLM
This distinction is important.

An LLM evaluation may focus primarily on the quality of generated text.

An AI agent introduces additional behavior.

For example:

User Request
     ↓
Agent
     ↓
Tool Selection
     ↓
Tool Call
     ↓
Tool Result
     ↓
Agent Reasoning
     ↓
Final Answer

There are therefore multiple layers to evaluate.

Layer 1: Final Answer
Is the response correct and useful?

Layer 2: Tool Selection
Did the agent choose the right tool?

Layer 3: Tool Arguments
Did it provide the correct parameters?

Layer 4: Tool Result Handling
Did it correctly interpret the tool's output?

Layer 5: Task Completion
Did the overall workflow accomplish the user's goal?

A mature AI-agent evaluation strategy should consider more than the final response.

Example: Evaluating a Research Agent
Imagine an AI research agent with three tools:

search_web()
retrieve_document()
summarize_document()

A user asks:

"Find the main causes of customer churn and summarize the evidence."

The agent might:

Search for relevant information.

Retrieve documents.

Extract relevant passages.

Summarize findings.

You can evaluate each stage.

Retrieval Evaluation
Did the agent find relevant sources?

Grounding Evaluation
Are the conclusions supported by those sources?

Answer Evaluation
Does the final response answer the original question?

Task Evaluation
Did the agent successfully complete the research workflow?

This layered approach makes debugging much easier.

Step 6: Create a Dataset of Evaluation Cases
Testing one example isn't enough.

Build a dataset.

For example:

Test 001 — Simple customer question
Test 002 — Ambiguous request
Test 003 — Missing information
Test 004 — Out-of-scope question
Test 005 — Multi-step request
Test 006 — Tool failure
Test 007 — Adversarial input
Test 008 — Long context

This gives you a more realistic picture of agent performance.

The goal is not simply to maximize one score.

The goal is to understand where the system succeeds and where it fails.

Step 7: Test Edge Cases
Strong evaluation suites include difficult cases.

Examples include:

Ambiguous Requests
"Can you update it?"

The agent may need clarification.

Missing Data
"Compare this month's revenue with last month."

What happens if last month's data isn't available?

Incorrect Assumptions
What happens if the user provides a false premise?

Tool Failure
What happens if an API returns an error?

Unexpected Input
What happens if the user asks the agent to perform an unsupported action?

Prompt Injection
What happens when untrusted content attempts to override the agent's instructions?

These cases often reveal weaknesses that simple happy-path testing misses.

Step 8: Use Thresholds Carefully
Evaluation frameworks often allow you to define thresholds.

For example:

threshold=0.7

This can help classify a result as passing or failing.

But don't assume that a threshold of 0.7 is universally meaningful.

The right threshold depends on:

Your application

Your users

Your risk tolerance

The metric

The evaluation model

The cost of failure

A customer-support system and an experimental creative-writing assistant may require very different standards.

Step 9: Build Regression Tests
One of the most valuable uses of LLM evaluation is regression testing.

Suppose version 1 of your agent scores well.

You change:

The system prompt

The model

The retrieval system

A tool

The agent framework

Run the same evaluation suite again.

Then compare:

Version 1
Answer Relevancy: 0.86
Faithfulness:     0.91

Version 2
Answer Relevancy: 0.90
Faithfulness:     0.73

The new version may appear better in one dimension while becoming substantially worse in another.

This is exactly the kind of regression that systematic evaluation can expose.

LLM-as-a-Judge
Many AI evaluations use another language model to judge an AI-generated response.

This is often called LLM-as-a-judge.

For example:

Agent A generates the answer.

Evaluator model assesses the answer according to defined criteria.

The evaluator might consider:

Relevance

Correctness

Completeness

Style

Grounding

This approach is useful because many AI outputs don't have one objectively correct string.

However, LLM-based evaluation has limitations.

The evaluator can also make mistakes.

It may:

Prefer certain writing styles

Misinterpret the expected answer

Produce inconsistent judgments

Overestimate fluent but incorrect responses

For high-stakes systems, combine LLM-based evaluation with deterministic tests and human review.

Deterministic Tests vs. LLM Evaluation
These two approaches complement each other.

Deterministic Tests
Useful for:

Exact values

API responses

Tool arguments

Database operations

JSON structure

Required fields

Status codes

LLM Evaluation
Useful for:

Relevance

Helpfulness

Quality

Tone

Grounding

Completeness

The strongest testing strategy often combines both.

A Practical Evaluation Architecture
A mature evaluation system can look like this:

                    Test Dataset
                         |
                         ↓
                    AI Agent
                         |
              +----------+----------+
              |          |          |
           Output      Tools      Traces
              |          |          |
              +----------+----------+
                         |
                         ↓
                  Evaluation Layer
                         |
          +--------------+--------------+
          |              |              |
       Relevancy      Faithfulness   Task Success
          |              |              |
          +--------------+--------------+
                         |
                         ↓
                  Evaluation Report

This gives developers visibility into both output quality and agent behavior.

Testing Tool Use
For AI agents, tool usage deserves special attention.

Suppose your agent has:

search_database()
send_email()
create_ticket()

A user asks:

"Find my order status."

The agent should probably query the database.

It should not send an email or create a support ticket unless necessary.

You can therefore test:

Did the agent select the correct tool?

Were the arguments correct?

Was the tool called the correct number of times?

Did the agent handle the result correctly?

Did it avoid unnecessary actions?

These are agent-specific evaluation questions.

Testing Task Completion
A final answer can look excellent while the task itself remains incomplete.

Imagine an agent responsible for creating a support ticket.

The agent responds:

"I've prepared the ticket details."

But the ticket was never actually created.

A text-quality metric might score the response highly.

A task-completion evaluation would correctly identify the failure.

This is why agent evaluation should measure actions and outcomes, not just language quality.

Human Evaluation Still Matters
Automated evaluation is powerful, but it shouldn't always be the only source of truth.

Human reviewers can evaluate areas that automated metrics may miss.

For example:

Strategic usefulness

Business relevance

Brand alignment

Subtle factual errors

User satisfaction

A practical system can combine:

Automated Tests + LLM Evaluation + Human Review

Each layer provides a different perspective.

Common Mistakes in LLM Evaluation
Mistake 1: Using Only One Metric
A single score rarely tells the complete story.

Use multiple metrics that reflect your application's goals.

Mistake 2: Testing Only Easy Examples
Real users create messy inputs.

Include edge cases and failures.

Mistake 3: Evaluating Only the Final Answer
Agent behavior includes tool use, retrieval, and task completion.

Test those layers too.

Mistake 4: Treating LLM Scores as Absolute Truth
LLM-based evaluators are themselves imperfect.

Use them as signals rather than unquestionable ground truth.

Mistake 5: Not Maintaining an Evaluation Dataset
Your test cases should evolve as you discover new failure modes.

Every important production failure can potentially become a new regression test.

Best Practices for Testing AI Agents
A reliable evaluation strategy should:

Define success criteria before testing.

Build a representative test dataset.

Include both normal and edge cases.

Test tool usage separately.

Use deterministic checks where possible.

Use LLM-based metrics for subjective criteria.

Track results across agent versions.

Investigate failures rather than focusing only on average scores.

Add production failures to your evaluation dataset.

Combine automated evaluation with human review where appropriate.

Most importantly, treat evaluation as an ongoing engineering process.

AI systems change.

Models change.

Prompts change.

Data changes.

Tools change.

Your evaluation suite should evolve with them.

A Practical DeepEval Workflow
A simplified workflow for a Python project might look like:

1. Define evaluation criteria
          ↓
2. Create test cases
          ↓
3. Run the AI agent
          ↓
4. Capture outputs and traces
          ↓
5. Run DeepEval metrics
          ↓
6. Review failed cases
          ↓
7. Improve the agent
          ↓
8. Run regression tests again

This creates a feedback loop:

Build → Test → Measure → Improve → Retest

That feedback loop is essential for developing reliable AI applications.

Final Thoughts
Testing AI Agents with Python: LLM Evaluation with DeepEval is ultimately about moving AI development from subjective experimentation toward measurable engineering.

A working AI agent is not necessarily a reliable AI agent.

To build systems that users can trust, you need to test:

What the agent says

What information it uses

Which tools it calls

How it handles uncertainty

Whether it completes the task

How its performance changes over time

DeepEval can provide a useful foundation for creating repeatable LLM evaluation workflows in Python.

Start small.

Create a handful of realistic test cases.

Choose metrics that reflect your application's goals.

Add difficult cases.

Track regressions.

Then gradually build a more comprehensive evaluation suite.

The most important shift is simple:

Don't just ask whether your AI agent works. Build a system that can measure how well it works.

FAQ for Google Search
What is DeepEval?
DeepEval is a Python-based framework for evaluating LLM-powered applications. It provides tools for creating test cases and applying evaluation metrics to AI-generated outputs.

How do you test an AI agent with Python?
You can test an AI agent by creating representative test cases, running the agent against them, evaluating the outputs with appropriate metrics, and tracking results across different versions of the system.

What should you test in an AI agent?
You should evaluate final-answer quality as well as tool selection, tool arguments, retrieval quality, task completion, instruction following, and failure handling.

What is LLM evaluation?
LLM evaluation is the process of measuring the quality, reliability, and behavior of an application powered by a large language model. Evaluation can use deterministic tests, model-based metrics, human review, or a combination of these methods.

What is LLM-as-a-judge?
LLM-as-a-judge is an evaluation approach in which one language model assesses the output of another AI system according to defined criteria such as relevance, correctness, or helpfulness.

Can DeepEval test AI agents?
DeepEval can be used as part of an evaluation workflow for LLM applications and agentic systems. Agent evaluation should consider both the final output and agent-specific behavior such as tool use and task completion.

Why is regression testing important for AI agents?
AI-agent behavior can change when you modify prompts, models, retrieval systems, tools, or other components. Regression testing helps identify when a new version improves one capability while unintentionally degrading another.

Post a Comment

Previous Post Next Post