AI Eval helps measure and improve AI integrations in Drupal by allowing you to define test datasets, run them against AI agents or providers, and score the results. It offers a dashboard to monitor quality, an optimizer to suggest prompt improvements, and pluggable graders for various evaluation criteria.
AI Eval helps you find out whether AI features on a Drupal site are getting better or worse. Describe a good answer, run a fixed set of questions, score each answer with judges and checks, then compare the result with a quality gate. Agent mode tests an AI Agents plugin from end to end. Chat mode tests a provider, model, and prompt without an agent framework.
Built by a human using an AI assistant: ๐ค โ ๐งMost of this module's code was generated with an AI assistant under human direction, review, and testing, using Strikethroo & Kenkeep for improved Drupal code generation.
Try the example
The Riverside Library example is a small library help desk. It comes with
a test, a six-question exam, and 36 recorded replies to review. Two demo
videos use it, and you can follow them on your own machine.
You need DDEV and an OpenRouter API key. Run this command from an empty folder:
bash <(curl -fsSL https://git.drupalcode.org/project/ai_eval/-/raw/main/installer/riverside.sh)The installer asks for a project name and your key. The key is saved in
the project's .ddev/.env file and read from the environment; it
is not stored in config or the database. When the install ends, look for
"Riverside example: READY" just above the installer's summary. A run of the
example exam costs a few cents.
On an existing site with any AI provider, apply the recipe instead:
drush recipe modules/contrib/ai_eval/recipes/riverside_libraryThe recipe asks for the provider and models to use. The example's README walks through both videos step by step.
Start with a guided test
Create a test in three steps. First choose the AI, graders, and quality gate. Next add examples or choose an existing dataset. Finally review the test, then save it, or save it and go straight to the run button. Drafts remain available for seven days. When several tests share a dataset, you can update the shared dataset or create a copy for one test.
Read a result
Each result starts with โN of M examples met your expectations.โ Questions that need attention appear first. Run history records completed, cancelled, errored, and interrupted runs. You can resume an interrupted run or start a new one.
Compare two runs of the same test. When both runs used the same judge, prompt, and settings, the page reports a direct comparison. When something differed, it names what changed, so read the difference with care. When a run did not record its settings, the page shows the two results side by side without a verdict.
Know whether to trust the judge
Validate each LLM judge against answers labeled by people. AI Eval measures how often the judge agrees with people on good answers and on bad answers. A trust badge on the Dashboard shows the current state. Results warn you when a quality gate used a judge that has not passed validation.
Find out why it fails
Import traces, the recorded conversations between users and your AI feature, review them, and label what happened. Add failure modes to group recurring problems. Promote a conversation into a dataset so the same failure becomes a question in later runs.
Graders and rubrics
Nine graders ship with the module. Six LLM judges measure relevance, completeness, accuracy, whether the answer gives the reader useful next steps, fact match against verified facts, and whether claims are supported by source material. Three deterministic graders check the answer format, the rubric checks described below, and which tools the agent called.
Rubrics let many questions share the same expectations. The editor provides form fields for the two text-list check kinds and JSON editing for every other kind. The rubric page shows the whole rubric and provides an Edit button.
Improve the prompt
The optimizer reads failed answers and asks a model to propose a better system prompt. Review a proposal before you apply it. The optimizer repeats the current prompt and each candidate because judge scores can vary between runs. It accepts a candidate only when the measured improvement clears the configured margin.
Import and export
Export completed runs as Every Eval Ever (EEE) 0.2.2 JSON files. Import EEE files into the results table to inspect results produced elsewhere and compare results from different test tools. EEE files contain aggregate metrics and run details, not question or answer text. You can also import OpenTelemetry AI traces from a JSON file or directory.
Drush commands
ai-eval:run(aer): run tests or preview what would run.ai-eval:optimize(aeo): generate and measure candidate prompts.ai-eval:distill(aed): turn stored results into summary files.ai-eval:sample-traces(aest): export answers for human labeling.ai-eval:validate-judge(aevj): compare a judge with human labels.ai-eval:judges(aej): list judge trust and validation results.ai-eval:export-envelope(aee): export stored results as EEE files.ai-eval:import-envelope(aeie): import EEE result files.ai-eval:import-traces(aeit): import OpenTelemetry AI traces.ai-eval:target-fixtures(aetf): list the proving-ground fixtures a test requires.
Requirements
- Drupal 11.3 or later
- PHP 8.3 or later
- AI 1.0 or later
Agent mode also requires AI Agents. Trace review can use AI Observability and OpenTelemetry. These features are optional.
Documentation
Read the published documentation for installation, configuration, usage, and extension guides. The repository README provides an overview and quick start.
Depends on
Dependencies of the latest stable release
No dependencies recorded for this project.
Required by
Tracked projects that depend on this one
No tracked projects depend on this one yet.
More in the AI ecosystem
Most installed first
Activity
Release Timeline
Releases
| Version | Type | Core | Release date | |
|---|---|---|---|---|
| 1.0.0-beta6 | Pre-release | 11 | Sep 25, 2026 | |
| 1.0.0-beta5 | Pre-release | 11 | Sep 25, 2026 | |
| 1.0.0-beta4 | Pre-release | 11 | Sep 24, 2026 | |
| 1.1.x-dev | Dev | 11 | Sep 10, 2026 | |
| 1.0.0-beta3 | Pre-release | 11 | Aug 15, 2026 | |
| 1.0.0-beta2 | Pre-release | 11 | Aug 7, 2026 | |
| 1.0.0-beta1 | Pre-release | 11 | Jul 16, 2026 | |
| 1.0.0-alpha10 | Pre-release | 11 | Apr 25, 2026 | |
| 1.0.0-alpha8 | Pre-release | 11 | Apr 23, 2026 | |
| 1.0.0-alpha7 | Pre-release | 11 | Apr 15, 2026 | |
| 1.0.0-alpha6 | Pre-release | 11 | Apr 9, 2026 | |
| 1.0.0-alpha5 | Pre-release | 11 | Apr 9, 2026 | |
| 1.0.0-alpha4 | Pre-release | 11 | Apr 8, 2026 | |
| 1.0.0-alpha3 | Pre-release | 11 | Apr 6, 2026 | |
| 1.0.0-alpha2 | Pre-release | 11 | Apr 5, 2026 | |
| 1.0.0-alpha1 | Pre-release | 11 | Apr 5, 2026 |