Skip to main content
drupalreleases
Release: Cms 2.2.3 — Update released for Drupal core (2.2.3)! Release: Easy Breadcrumb 2.0.11 — Minor update available for module easy_breadcrumb (2.0.11). Release: Bootstrap 8.x-3.42 — Minor update available for theme bootstrap (8.x-3.42). Release: Editoria11y Accessibility Checker 3.0.10 — Minor update available for module editoria11y (3.0.10). Release: Editoria11y Accessibility Checker 2.2.24 — Minor update available for module editoria11y (2.2.24). Release: Leaflet 10.4.13 — Minor update available for module leaflet (10.4.13). Release: Flag 5.1.1 — Minor update available for module flag (5.1.1). Release: MCP Sentinel 2.33.0 — Minor update available for module mcp_sentinel (2.33.0). Module Revived: Bootstrap 8.x-3.41 — Theme bootstrap updated after 6 months of inactivity (8.x-3.41). Security Coverage: Component Library — Module component_library now has official Drupal security advisory coverage.

AI Eval

3 sites No security coverage

Part of the AI ecosystem · 340 projects

View on drupal.org

AI Eval helps measure and improve AI integrations in Drupal by allowing you to define test datasets, run them against AI agents or providers, and score the results. It offers a dashboard to monitor quality, an optimizer to suggest prompt improvements, and pluggable graders for various evaluation criteria.

AI Eval helps you find out whether AI features on a Drupal site are getting better or worse. Describe a good answer, run a fixed set of questions, score each answer with judges and checks, then compare the result with a quality gate. Agent mode tests an AI Agents plugin from end to end. Chat mode tests a provider, model, and prompt without an agent framework.

Built by a human using an AI assistant: ๐Ÿค– โž• ๐Ÿง 

Most of this module's code was generated with an AI assistant under human direction, review, and testing, using Strikethroo & Kenkeep for improved Drupal code generation.

Try the example

The Riverside Library example is a small library help desk. It comes with
a test, a six-question exam, and 36 recorded replies to review. Two demo
videos use it, and you can follow them on your own machine.

You need DDEV and an OpenRouter API key. Run this command from an empty folder:

bash <(curl -fsSL https://git.drupalcode.org/project/ai_eval/-/raw/main/installer/riverside.sh)

The installer asks for a project name and your key. The key is saved in
the project's .ddev/.env file and read from the environment; it
is not stored in config or the database. When the install ends, look for
"Riverside example: READY" just above the installer's summary. A run of the
example exam costs a few cents.

On an existing site with any AI provider, apply the recipe instead:

drush recipe modules/contrib/ai_eval/recipes/riverside_library

The recipe asks for the provider and models to use. The example's README walks through both videos step by step.

Start with a guided test

Create a test in three steps. First choose the AI, graders, and quality gate. Next add examples or choose an existing dataset. Finally review the test, then save it, or save it and go straight to the run button. Drafts remain available for seven days. When several tests share a dataset, you can update the shared dataset or create a copy for one test.

Read a result

Each result starts with โ€œN of M examples met your expectations.โ€ Questions that need attention appear first. Run history records completed, cancelled, errored, and interrupted runs. You can resume an interrupted run or start a new one.

Compare two runs of the same test. When both runs used the same judge, prompt, and settings, the page reports a direct comparison. When something differed, it names what changed, so read the difference with care. When a run did not record its settings, the page shows the two results side by side without a verdict.

Know whether to trust the judge

Validate each LLM judge against answers labeled by people. AI Eval measures how often the judge agrees with people on good answers and on bad answers. A trust badge on the Dashboard shows the current state. Results warn you when a quality gate used a judge that has not passed validation.

Find out why it fails

Import traces, the recorded conversations between users and your AI feature, review them, and label what happened. Add failure modes to group recurring problems. Promote a conversation into a dataset so the same failure becomes a question in later runs.

Graders and rubrics

Nine graders ship with the module. Six LLM judges measure relevance, completeness, accuracy, whether the answer gives the reader useful next steps, fact match against verified facts, and whether claims are supported by source material. Three deterministic graders check the answer format, the rubric checks described below, and which tools the agent called.

Rubrics let many questions share the same expectations. The editor provides form fields for the two text-list check kinds and JSON editing for every other kind. The rubric page shows the whole rubric and provides an Edit button.

Improve the prompt

The optimizer reads failed answers and asks a model to propose a better system prompt. Review a proposal before you apply it. The optimizer repeats the current prompt and each candidate because judge scores can vary between runs. It accepts a candidate only when the measured improvement clears the configured margin.

Import and export

Export completed runs as Every Eval Ever (EEE) 0.2.2 JSON files. Import EEE files into the results table to inspect results produced elsewhere and compare results from different test tools. EEE files contain aggregate metrics and run details, not question or answer text. You can also import OpenTelemetry AI traces from a JSON file or directory.

Drush commands

  • ai-eval:run (aer): run tests or preview what would run.
  • ai-eval:optimize (aeo): generate and measure candidate prompts.
  • ai-eval:distill (aed): turn stored results into summary files.
  • ai-eval:sample-traces (aest): export answers for human labeling.
  • ai-eval:validate-judge (aevj): compare a judge with human labels.
  • ai-eval:judges (aej): list judge trust and validation results.
  • ai-eval:export-envelope (aee): export stored results as EEE files.
  • ai-eval:import-envelope (aeie): import EEE result files.
  • ai-eval:import-traces (aeit): import OpenTelemetry AI traces.
  • ai-eval:target-fixtures (aetf): list the proving-ground fixtures a test requires.

Requirements

  • Drupal 11.3 or later
  • PHP 8.3 or later
  • AI 1.0 or later

Agent mode also requires AI Agents. Trace review can use AI Observability and OpenTelemetry. These features are optional.

Documentation

Read the published documentation for installation, configuration, usage, and extension guides. The repository README provides an overview and quick start.

Depends on

Dependencies of the latest stable release

No dependencies recorded for this project.

Required by

Tracked projects that depend on this one

No tracked projects depend on this one yet.

Activity

Tracked releases
16
Tracked since
Apr 2026
Latest release
1 week ago
Releases (12 mo)
16 ▲ from 0
Maintenance
Active

Release Timeline

Releases

Version Type Core Release date
1.0.0-beta6 Pre-release 11 Sep 25, 2026
1.0.0-beta5 Pre-release 11 Sep 25, 2026
1.0.0-beta4 Pre-release 11 Sep 24, 2026
1.1.x-dev Dev 11 Sep 10, 2026
1.0.0-beta3 Pre-release 11 Aug 15, 2026
1.0.0-beta2 Pre-release 11 Aug 7, 2026
1.0.0-beta1 Pre-release 11 Jul 16, 2026
1.0.0-alpha10 Pre-release 11 Apr 25, 2026
1.0.0-alpha8 Pre-release 11 Apr 23, 2026
1.0.0-alpha7 Pre-release 11 Apr 15, 2026
1.0.0-alpha6 Pre-release 11 Apr 9, 2026
1.0.0-alpha5 Pre-release 11 Apr 9, 2026
1.0.0-alpha4 Pre-release 11 Apr 8, 2026
1.0.0-alpha3 Pre-release 11 Apr 6, 2026
1.0.0-alpha2 Pre-release 11 Apr 5, 2026
1.0.0-alpha1 Pre-release 11 Apr 5, 2026