EvaliQA
Sign inStart free
AI evaluation & observability

Ship AI changes with evidence. Not intuition.

EvaliQA is the workspace for evaluating AI systems end-to-end — text, multi-turn, redteam and voice. Connect your endpoint, generate real datasets, score with the metrics you actually need, and review the evidence in one place.

Start free No card · BYOK · Export anytime
Assess. Assure. Accept. — 100+ model providers, seven CI/CD tools, five voice engines. One workspace.
OpenAI · Anthropic · GeminiGitHub Actions · GitLab · JenkinsTwilio · Deepgram · ElevenLabsSlack · Webhooks · CSV export
One reviewable loop

From requirement to release verdict.

Every step of the eval is an artifact — the plan, the dataset, the run, the verdict, the reasoning. All in the same workspace, all reviewable by anyone on the team.

Nine-step wizard

A test plan is a spec. The eval-run is its acceptance test.

Start with a plan — behaviors, edge cases, judges and thresholds. EvaliQA turns it into datasets, connectors and reproducible runs. The result is a record you can compare across releases.

  • Basics · Mode · Judge · Metrics · Params · Docs · Dataset · Generate · Review
  • AI suggestions on goals, mode and metrics — you keep control
  • Documents attach to the plan; datasets ground on them
evaliqa.app / test-plans / new
01 · Basics02 · Mode03 · Judge04 · Metrics05 · Params06 · Docs07 · Dataset08 · Generate09 · Review
01
Write the test plan
Define behavior, edge cases and the metrics that make a release acceptable.
02
Build the dataset
Generate rows across five failure classes or edit them by hand — every row is reviewable.
03
Run against your system
Point EvaliQA at any HTTP connector. Text, multi-turn, redteam or voice.
04
Decide with evidence
Compare runs, open failed rows, read the judge's reasoning — then ship.
One workspace, four eval modes

Text, chat, redteam, voice. Same scoring pipeline.

The five modes share one metric verdict shape — whether the input was a JSON prompt, a whole conversation, an adversarial attacker or a live phone call.

Single-turn

Text eval

Score prompts, summaries and classifications against a metric set. The baseline for RAG answers and structured outputs.

Multi-turn

Chatbot eval

Simulation, Scripted or Adaptive. Fifteen personas drive the whole conversation, not just turn one.

Redteam

Safety & attacks

Ten vulnerability classes and nine attack techniques, single-turn or escalating multi-turn. Severity filter per plan.

Voice

Live phone call

The platform agent dials your voice bot over Twilio, drives the scenario, and hands the transcript to the same metric pipeline.

Read the run like a senior engineer

Every failing row carries its own reasoning.

Each result row shows the input, the actual output, the metric verdict and the judge's reasoning — plus latency, cost and token counts. Open a failure, see exactly why the metric refused, and jump to the fix.

  • Filter by pass, fail or error — with counts on the chip
  • Compare against any previous run to spot regressions
  • Export the full run as CSV or generate an AI report
evaliqa.app / eval-runs / 4f92c1a0
Eval run · 4f92c1a0
Mode: eval_multi_turn · Model: claude-sonnet-4 · Cost: $0.128
ReportCSV
Rows
50
Pass rate
74%
Avg latency
1.24 s
Total cost
$0.128
All 50Failed 13
#12Multi-turn — support/refund escalationfail
Groundedness (built-in)0.34 / ≥ 0.70
Cited SLA of 4h; retrieved passages state 24h for the customer's tier.
Tone (custom · GEval)0.81
Assistant stayed on-tone across four turns; empathy intact through escalation.
#18failMulti-turn — pricing follow-up · Groundedness 0.512.1 s
Multi-turn

Chatbots fail on turn four, not turn one.

Pick a strategy per test case, then let EvaliQA drive the whole exchange — deterministic when you need it, adaptive when the reality of your system matters more.

SimulationScriptedAdaptive · recommended
  • Simulation. Fifteen personas, deterministic seeds. Predictable coverage.
  • Scripted. You write the user turns. Byte-exact replays for regression.
  • Adaptive. The platform agent improvises based on your system's live answers.
Voice eval

Score a live phone call the same way you score a prompt.

The platform agent dials your voice bot over Twilio, drives the scenario with STT and TTS in the loop, records the audio, and hands the transcript to the same metric pipeline — the audio player and scores live next to the run.

  • Twilio dial → Media Streams WebSocket
  • STT — Deepgram or Whisper
  • TTS — ElevenLabs or OpenAI
  • Audio + transcript + metrics in one run detail
Metrics you write

Judges tuned to your product, not a library default.

Compose custom judges from natural-language criteria, plug in your own Python scoring code, or reuse the built-ins from eval-ai-library. Every metric produces the same reviewable evidence.

Judge · natural languagev3
Refund policy compliance

Answer must not offer a refund outside the 30-day window unless the customer cites a billing error.

Pass rate
87%
Code · your PythonCustom
Cited-doc precision

Compares cited passages against the retrieved chunks. Fails when a citation is fabricated.

Pass rate
64%
Built-in · eval-ai-libraryGroundedness
Answer faithfulness

Standard groundedness score against retrieved context. Threshold configurable per plan.

Pass rate
92%
New custom metric
Save to your workspace catalog
Tone of voice
Fails when the assistant sounds cold, dismissive or overly formal for a support context.
GEval
0.75
GEval parameters
The assistant should sound warm, patient and empathetic. Reject responses that read as dismissive, condescending or scripted.
3
CancelSave metric
After deploy

Watch production the same way you tested it.

A failing metric links back to the exact input, output, model and timing. Thresholds turn into alerts. Triggers turn into CI/CD checks. The evaluation loop keeps running past release day.

Alerts

Notify Slack, email or a webhook when quality drifts.

Bind a threshold to any metric, wire it to a channel, and route by project. Firing history is stored — you can prove the alert triggered, and when.

evaliqa.app / alerts
Rules
2 rules on Acme production
+ New rule
Groundedness droppass_rate < 70% over 15m#eng-qualityon
Voice refund policyrefund_compliance < 0.6 · any runwebhookemailon
Legacy CSAT dropcsat < 4.0 · daily#supportoff
CI/CD triggers

Kick off an eval-run from your pipeline.

Issue a token from Integrations, drop the snippet in your workflow, and every merge runs the acceptance suite before it ships. Seven pipeline tools ready out of the box.

evaliqa.app / integrations
Integrations
API tokens · CI/CD tools
ci-github-actionsevx_live_a1b2•••••active
GitHub ActionsGitHub Actions
GitLab CI/CDGitLab CI/CD
JenkinsJenkins
CircleCICircleCI
+3 more: CircleCI · Bitbucket · ShellSet up →
Architecture, not promises

Keep model access and evidence under your control.

Evaluation targets and judges use your provider credentials. EvaliQA charges for the workspace around them — not resold tokens or an opaque score you can't inspect.

Your keys

Provider credentials stay in your workspace, encrypted at rest with AES-256-GCM.

Your boundary

Use the hosted product or deploy the same service set inside your perimeter.

Your system

Connect any HTTP endpoint. No proprietary agent runtime, no SDK to install.

Your evidence

Export datasets, run results and traces in open, inspectable formats.

Pricing

Simple enough to approve.

Bring your own model keys. Platform pricing covers the evaluation workspace, collaboration and evidence — not the tokens.

Free
$0

Forever. No card required.

  • Seats: 2
  • Projects: 2
  • Eval cases: 1 000 / mo
  • Runtime traces: 5 000 / mo
  • Single-turn and multi-turn text eval
  • Test plan wizard with AI suggestions
  • Dataset generation and editor
Start free
Team
$149/month

$1,490 billed annually.

  • Seats: 5
  • Projects: 10
  • Eval cases: 10 000 / mo
  • Runtime traces: 25 000 / mo
  • Everything in Free
  • Voice evaluation — outbound calls, transcript, scoring
  • Red teaming — 10 vulnerability classes, 9 attack techniques
Subscribe
Custom
Custom

For your perimeter and process.

  • Seats: Unlimited
  • Projects: Unlimited
  • Eval cases: Unmetered
  • Runtime traces: Unmetered
  • Everything in Team
  • Workspace roles and per-project access control
  • Audit log of every action, exportable
Ready when you are

Replace "looks good" with evidence.

Start free with your own provider keys. Export anytime.

Start free