Testing¶
ScriptedProvider replays prepared model turns so real Tools and the real
RunLoop can be tested offline. Use Evals for answer quality, Tool
choice, or behavioral consistency.
from lovia import Agent, tool
from lovia.testing import ScriptedProvider, call, text
@tool
def add(a: int, b: int) -> int:
"""Add two numbers."""
return a + b
def make_agent() -> Agent:
return Agent(
name="calc",
model=ScriptedProvider([
call("add", {"a": 2, "b": 3}, call_id="c1"), # turn 1: request the tool
text("The answer is 5."), # turn 2: final answer
]),
tools=[add],
)
async def test_calc_uses_the_tool():
result = await make_agent().run("What is 2 + 3?")
assert result.output == "The answer is 5."
assert result.turns == 2
The script is the model's side of the conversation: each entry answers one model call, in order. Real tools run for real — only the LLM is scripted — so this exercises the actual loop: schema validation, parallel execution, approval gates, structured-output parsing, session persistence.
Building scripts¶
| Helper | Produces |
|---|---|
text("Done.") |
a plain-text turn (streams character by character) |
text("Done.", reasoning="hmm...") |
text preceded by reasoning deltas — for testing ReasoningDelta consumers |
call("search", {"q": "tides"}) |
a turn requesting one tool call (call_id defaults to call_<name>) |
batch(("a", {...}), ("b", {...})) |
a turn requesting several calls at once — for testing parallel execution |
A script that runs dry raises
AssertionError("ScriptedProvider ran out of canned responses") — a wrong
turn count fails loudly instead of hanging.
Asserting on what the agent saw¶
The provider records every prompt it received:
provider = ScriptedProvider([text("ok")])
agent = Agent(name="bot", model=provider, instructions="Be terse.")
await agent.run("hello")
first_prompt = provider.calls[0] # list[Message] for turn 1
assert first_prompt[0].role == "system"
assert "Be terse." in first_prompt[0].content
provider.calls[i] is the chat-format view of turn i's input — the
sharpest tool for testing dynamic instructions,
view injectors, and
compaction behavior ("did the cleared result really leave
the view?").
What to test, and with what¶
- Tool business logic — keep the implementation as a separate callable
and test it with pytest; use
ScriptedProviderfor Tool schema and wiring. - Loop behavior (routing, tool choice, repair, guardrails, handoff) —
ScriptedProvider, as above. Handoffs and agent-as-tool sub-runs each consume from their own agent's provider, so give each agent its own script. - Event consumers / UIs — script a run and iterate
Runner.stream(...); deltas stream character-by-character so consumers see realistic fragmentation. - Behavioral quality ("does it answer well?") — that's
evals, which use the same
ScriptedProviderfor their offline mode and real models for the live one. - Live smoke tests — mark them, skip by default, run on demand (this
repo uses
pytest -m live_providergated behindLOVIA_LIVE_TESTS=1).
Sharp edges¶
- A
ScriptedProvideris single-use. It pops a shared queue — neither repeat- nor concurrency-safe. Build a fresh provider (and agent) per run; withevaluate()pass an agent factory for exactly this reason. supports_json_schemaisFalse, so structured output takes the prompt path: your scripted final turn must be the JSON document itself (text('{"title": "..."}')), and the schema instructions land inprovider.calls[0][0]where you can assert on them.- Async tests need an async runner — the repo uses
pytest-asyncio;Runner.run_syncworks in plain tests when no loop is running.
See also¶
- Evals — the same double, measuring quality instead of wiring
- Providers —
ScriptedProviderdoubles as the referenceProviderimplementation - Example:
10_custom_provider.py(offline), plus this repo'stests/directory