AI agent testing for QA and Python devs: RAG, tool calls, MCP, prompt injection, tracing, CI/CD gates. No API key.
Preview this Course
Testing AI Agents with Python: LLM Evaluation with DeepEval
Building an AI agent is only the beginning.
Once an agent can reason, call tools, retrieve information, and generate responses, a more difficult question appears:
How do you know whether the agent is actually working well?
Traditional software testing can verify whether a function returns the expected value. AI systems are different. Their outputs can vary, and an answer may be technically valid while still being incomplete, irrelevant, unsafe, or poorly grounded.
This is why LLM evaluation has become an important part of AI-agent development.
In this guide, we'll explore Testing AI Agents with Python: LLM Evaluation with DeepEval and look at how developers can build repeatable evaluation workflows for AI applications.
You'll learn what to test, how evaluation metrics work, how to create test cases, and how to use DeepEval as part of a Python-based evaluation workflow.
Why Testing AI Agents Is Different
Traditional applications usually have deterministic behavior.
For example:
def add(a, b):
return a + b
You can test:
add(2, 3) == 5
AI agents are more complicated.
An agent may receive:
"Summarize the latest customer feedback and identify the three most important problems."
The response could vary between runs.
The output may also contain several dimensions that need evaluation:
Is it factually correct?
Did it answer the question?
Did it use the right information?
Did it follow instructions?
Did it hallucinate?
Did it use tools correctly?
Was the response sufficiently complete?
Did it provide useful reasoning or evidence?
This means AI-agent testing often requires evaluation criteria, not just exact output matching.
What Is DeepEval?
DeepEval is a Python-based evaluation framework designed for testing and evaluating LLM-powered applications.
It provides concepts such as:
Test cases
Evaluation metrics
Model-based evaluation
Dataset-based testing
Custom evaluation criteria
Evaluation pipelines
This makes it useful when you want to move from manually asking an AI application questions toward a repeatable testing process.
Instead of evaluating an agent once, you can create a collection of test cases and run evaluations repeatedly as your system changes.
Why Evaluate an AI Agent?
Imagine you improve your agent's system prompt.
The new version performs better on one test but worse on five others.
Without systematic evaluation, you might not notice the regression.
A repeatable evaluation suite can help answer questions such as:
Did the new prompt improve answer quality?
Did changing the model increase hallucinations?
Does the agent follow instructions consistently?
Is retrieval quality improving?
Are tool calls becoming more reliable?
Did a code change introduce regressions?
This is especially important when AI systems are continuously updated.
The Basic AI Evaluation Workflow
A simple evaluation pipeline looks like this:
Input → AI Agent → Output → Evaluation Metric → Score → Analysis
You provide an input to the agent.
The agent generates a response.
An evaluation metric analyzes the response.
The evaluation produces a score or judgment.
You then use those results to improve the system.
With multiple test cases, the workflow becomes:
Test Dataset → Agent → Outputs → Evaluation → Results → Regression Analysis
This is the foundation of an AI evaluation pipeline.
Step 1: Define What "Good" Means
Before writing evaluation code, define success.
This is one of the most important steps.
Suppose you're building a customer-support agent.
A good response might need to be:
Relevant
Helpful
Accurate
Grounded in company information
Concise
Consistent with support policies
These criteria can become evaluation dimensions.
Without clearly defined criteria, evaluation scores may not tell you much.
Step 2: Create Representative Test Cases
A good evaluation dataset should represent real usage.
For example:
test_cases = [
{
"input": "How do I reset my password?",
"expected": "The response should explain the password reset process."
},
{
"input": "Can I cancel my subscription?",
"expected": "The response should accurately explain the cancellation policy."
}
]
In practice, evaluation test cases can contain richer information, including the actual input, expected output, context, and metadata.
The important principle is to test realistic scenarios, not just easy examples.
Step 3: Install Your Evaluation Environment
A Python project can be configured with the relevant evaluation dependencies.
For example:
pip install deepeval
Then import the components needed for your tests.
A typical workflow may involve test cases and evaluation metrics.
The exact APIs and available metrics can evolve, so check the current DeepEval documentation when implementing a production evaluation suite.
Step 4: Create an Evaluation Test Case
A basic conceptual structure looks like this:
from deepeval.test_case import LLMTestCase
test_case = LLMTestCase(
input="How do I reset my password?",
actual_output="You can reset your password from the account settings page."
)
The test case represents one interaction that you want to evaluate.
Depending on your evaluation scenario, you may also provide additional information such as expected output or contextual data.
Step 5: Add an Evaluation Metric
The next step is deciding how the output should be judged.
For example, you may want to evaluate whether the answer is relevant.
Conceptually:
from deepeval.metrics import AnswerRelevancyMetric
metric = AnswerRelevancyMetric(
threshold=0.7
)
You can then use the metric to evaluate the test case.
The exact metric configuration should match your application's requirements.
The important concept is:
Test Case + Metric = Evaluation
Important Metrics for AI-Agent Testing
Different applications require different evaluation criteria.
There is no single metric that can determine whether an AI agent is "good."
Answer Relevancy
Does the response actually answer the user's question?
An agent can generate grammatically perfect text that doesn't address the request.
Relevancy testing helps detect this problem.
Faithfulness
Is the generated answer supported by the information available to the system?
This is particularly important for retrieval-augmented generation systems.
Contextual Relevance
If your agent retrieves external context, does that context actually help answer the question?
Poor retrieval can lead to poor answers even when the language model itself is capable.
Contextual Precision
Does the system prioritize useful information rather than returning large amounts of irrelevant context?
This can be important for RAG-based agents.
Contextual Recall
Did the retrieval process obtain the information needed to answer the question?
Low recall can cause the agent to miss important evidence.
Hallucination
Does the agent generate unsupported or fabricated information?
Hallucination testing is especially important for systems that provide factual answers.
Task Completion
Did the agent actually accomplish the requested task?
For an agent that must call tools or complete a workflow, this can be more useful than evaluating the final text alone.
Testing an AI Agent Instead of Just an LLM
This distinction is important.
An LLM evaluation may focus primarily on the quality of generated text.
An AI agent introduces additional behavior.
For example:
User Request
↓
Agent
↓
Tool Selection
↓
Tool Call
↓
Tool Result
↓
Agent Reasoning
↓
Final Answer
There are therefore multiple layers to evaluate.
Layer 1: Final Answer
Is the response correct and useful?
Layer 2: Tool Selection
Did the agent choose the right tool?
Layer 3: Tool Arguments
Did it provide the correct parameters?
Layer 4: Tool Result Handling
Did it correctly interpret the tool's output?
Layer 5: Task Completion
Did the overall workflow accomplish the user's goal?
A mature AI-agent evaluation strategy should consider more than the final response.
Example: Evaluating a Research Agent
Imagine an AI research agent with three tools:
search_web()
retrieve_document()
summarize_document()
A user asks:
"Find the main causes of customer churn and summarize the evidence."
The agent might:
Search for relevant information.
Retrieve documents.
Extract relevant passages.
Summarize findings.
You can evaluate each stage.
Retrieval Evaluation
Did the agent find relevant sources?
Grounding Evaluation
Are the conclusions supported by those sources?
Answer Evaluation
Does the final response answer the original question?
Task Evaluation
Did the agent successfully complete the research workflow?
This layered approach makes debugging much easier.
Step 6: Create a Dataset of Evaluation Cases
Testing one example isn't enough.
Build a dataset.
For example:
Test 001 — Simple customer question
Test 002 — Ambiguous request
Test 003 — Missing information
Test 004 — Out-of-scope question
Test 005 — Multi-step request
Test 006 — Tool failure
Test 007 — Adversarial input
Test 008 — Long context
This gives you a more realistic picture of agent performance.
The goal is not simply to maximize one score.
The goal is to understand where the system succeeds and where it fails.
Step 7: Test Edge Cases
Strong evaluation suites include difficult cases.
Examples include:
Ambiguous Requests
"Can you update it?"
The agent may need clarification.
Missing Data
"Compare this month's revenue with last month."
What happens if last month's data isn't available?
Incorrect Assumptions
What happens if the user provides a false premise?
Tool Failure
What happens if an API returns an error?
Unexpected Input
What happens if the user asks the agent to perform an unsupported action?
Prompt Injection
What happens when untrusted content attempts to override the agent's instructions?
These cases often reveal weaknesses that simple happy-path testing misses.
Step 8: Use Thresholds Carefully
Evaluation frameworks often allow you to define thresholds.
For example:
threshold=0.7
This can help classify a result as passing or failing.
But don't assume that a threshold of 0.7 is universally meaningful.
The right threshold depends on:
Your application
Your users
Your risk tolerance
The metric
The evaluation model
The cost of failure
A customer-support system and an experimental creative-writing assistant may require very different standards.
Step 9: Build Regression Tests
One of the most valuable uses of LLM evaluation is regression testing.
Suppose version 1 of your agent scores well.
You change:
The system prompt
The model
The retrieval system
A tool
The agent framework
Run the same evaluation suite again.
Then compare:
Version 1
Answer Relevancy: 0.86
Faithfulness: 0.91
Version 2
Answer Relevancy: 0.90
Faithfulness: 0.73
The new version may appear better in one dimension while becoming substantially worse in another.
This is exactly the kind of regression that systematic evaluation can expose.
LLM-as-a-Judge
Many AI evaluations use another language model to judge an AI-generated response.
This is often called LLM-as-a-judge.
For example:
Agent A generates the answer.
Evaluator model assesses the answer according to defined criteria.
The evaluator might consider:
Relevance
Correctness
Completeness
Style
Grounding
This approach is useful because many AI outputs don't have one objectively correct string.
However, LLM-based evaluation has limitations.
The evaluator can also make mistakes.
It may:
Prefer certain writing styles
Misinterpret the expected answer
Produce inconsistent judgments
Overestimate fluent but incorrect responses
For high-stakes systems, combine LLM-based evaluation with deterministic tests and human review.
Deterministic Tests vs. LLM Evaluation
These two approaches complement each other.
Deterministic Tests
Useful for:
Exact values
API responses
Tool arguments
Database operations
JSON structure
Required fields
Status codes
LLM Evaluation
Useful for:
Relevance
Helpfulness
Quality
Tone
Grounding
Completeness
The strongest testing strategy often combines both.
A Practical Evaluation Architecture
A mature evaluation system can look like this:
Test Dataset
|
↓
AI Agent
|
+----------+----------+
| | |
Output Tools Traces
| | |
+----------+----------+
|
↓
Evaluation Layer
|
+--------------+--------------+
| | |
Relevancy Faithfulness Task Success
| | |
+--------------+--------------+
|
↓
Evaluation Report
This gives developers visibility into both output quality and agent behavior.
Testing Tool Use
For AI agents, tool usage deserves special attention.
Suppose your agent has:
search_database()
send_email()
create_ticket()
A user asks:
"Find my order status."
The agent should probably query the database.
It should not send an email or create a support ticket unless necessary.
You can therefore test:
Did the agent select the correct tool?
Were the arguments correct?
Was the tool called the correct number of times?
Did the agent handle the result correctly?
Did it avoid unnecessary actions?
These are agent-specific evaluation questions.
Testing Task Completion
A final answer can look excellent while the task itself remains incomplete.
Imagine an agent responsible for creating a support ticket.
The agent responds:
"I've prepared the ticket details."
But the ticket was never actually created.
A text-quality metric might score the response highly.
A task-completion evaluation would correctly identify the failure.
This is why agent evaluation should measure actions and outcomes, not just language quality.
Human Evaluation Still Matters
Automated evaluation is powerful, but it shouldn't always be the only source of truth.
Human reviewers can evaluate areas that automated metrics may miss.
For example:
Strategic usefulness
Business relevance
Brand alignment
Subtle factual errors
User satisfaction
A practical system can combine:
Automated Tests + LLM Evaluation + Human Review
Each layer provides a different perspective.
Common Mistakes in LLM Evaluation
Mistake 1: Using Only One Metric
A single score rarely tells the complete story.
Use multiple metrics that reflect your application's goals.
Mistake 2: Testing Only Easy Examples
Real users create messy inputs.
Include edge cases and failures.
Mistake 3: Evaluating Only the Final Answer
Agent behavior includes tool use, retrieval, and task completion.
Test those layers too.
Mistake 4: Treating LLM Scores as Absolute Truth
LLM-based evaluators are themselves imperfect.
Use them as signals rather than unquestionable ground truth.
Mistake 5: Not Maintaining an Evaluation Dataset
Your test cases should evolve as you discover new failure modes.
Every important production failure can potentially become a new regression test.
Best Practices for Testing AI Agents
A reliable evaluation strategy should:
Define success criteria before testing.
Build a representative test dataset.
Include both normal and edge cases.
Test tool usage separately.
Use deterministic checks where possible.
Use LLM-based metrics for subjective criteria.
Track results across agent versions.
Investigate failures rather than focusing only on average scores.
Add production failures to your evaluation dataset.
Combine automated evaluation with human review where appropriate.
Most importantly, treat evaluation as an ongoing engineering process.
AI systems change.
Models change.
Prompts change.
Data changes.
Tools change.
Your evaluation suite should evolve with them.
A Practical DeepEval Workflow
A simplified workflow for a Python project might look like:
1. Define evaluation criteria
↓
2. Create test cases
↓
3. Run the AI agent
↓
4. Capture outputs and traces
↓
5. Run DeepEval metrics
↓
6. Review failed cases
↓
7. Improve the agent
↓
8. Run regression tests again
This creates a feedback loop:
Build → Test → Measure → Improve → Retest
That feedback loop is essential for developing reliable AI applications.
Final Thoughts
Testing AI Agents with Python: LLM Evaluation with DeepEval is ultimately about moving AI development from subjective experimentation toward measurable engineering.
A working AI agent is not necessarily a reliable AI agent.
To build systems that users can trust, you need to test:
What the agent says
What information it uses
Which tools it calls
How it handles uncertainty
Whether it completes the task
How its performance changes over time
DeepEval can provide a useful foundation for creating repeatable LLM evaluation workflows in Python.
Start small.
Create a handful of realistic test cases.
Choose metrics that reflect your application's goals.
Add difficult cases.
Track regressions.
Then gradually build a more comprehensive evaluation suite.
The most important shift is simple:
Don't just ask whether your AI agent works. Build a system that can measure how well it works.
FAQ for Google Search
What is DeepEval?
DeepEval is a Python-based framework for evaluating LLM-powered applications. It provides tools for creating test cases and applying evaluation metrics to AI-generated outputs.
How do you test an AI agent with Python?
You can test an AI agent by creating representative test cases, running the agent against them, evaluating the outputs with appropriate metrics, and tracking results across different versions of the system.
What should you test in an AI agent?
You should evaluate final-answer quality as well as tool selection, tool arguments, retrieval quality, task completion, instruction following, and failure handling.
What is LLM evaluation?
LLM evaluation is the process of measuring the quality, reliability, and behavior of an application powered by a large language model. Evaluation can use deterministic tests, model-based metrics, human review, or a combination of these methods.
What is LLM-as-a-judge?
LLM-as-a-judge is an evaluation approach in which one language model assesses the output of another AI system according to defined criteria such as relevance, correctness, or helpfulness.
Can DeepEval test AI agents?
DeepEval can be used as part of an evaluation workflow for LLM applications and agentic systems. Agent evaluation should consider both the final output and agent-specific behavior such as tool use and task completion.
Why is regression testing important for AI agents?
AI-agent behavior can change when you modify prompts, models, retrieval systems, tools, or other components. Regression testing helps identify when a new version improves one capability while unintentionally degrading another.
