Text eval
Score prompts, summaries and classifications against a metric set. The baseline for RAG answers and structured outputs.
EvaliQA is the workspace for evaluating AI systems end-to-end — text, multi-turn, redteam and voice. Connect your endpoint, generate real datasets, score with the metrics you actually need, and review the evidence in one place.
Every step of the eval is an artifact — the plan, the dataset, the run, the verdict, the reasoning. All in the same workspace, all reviewable by anyone on the team.
Start with a plan — behaviors, edge cases, judges and thresholds. EvaliQA turns it into datasets, connectors and reproducible runs. The result is a record you can compare across releases.
The five modes share one metric verdict shape — whether the input was a JSON prompt, a whole conversation, an adversarial attacker or a live phone call.
Score prompts, summaries and classifications against a metric set. The baseline for RAG answers and structured outputs.
Simulation, Scripted or Adaptive. Fifteen personas drive the whole conversation, not just turn one.
Ten vulnerability classes and nine attack techniques, single-turn or escalating multi-turn. Severity filter per plan.
The platform agent dials your voice bot over Twilio, drives the scenario, and hands the transcript to the same metric pipeline.
Each result row shows the input, the actual output, the metric verdict and the judge's reasoning — plus latency, cost and token counts. Open a failure, see exactly why the metric refused, and jump to the fix.
Pick a strategy per test case, then let EvaliQA drive the whole exchange — deterministic when you need it, adaptive when the reality of your system matters more.
The platform agent dials your voice bot over Twilio, drives the scenario with STT and TTS in the loop, records the audio, and hands the transcript to the same metric pipeline — the audio player and scores live next to the run.
Compose custom judges from natural-language criteria, plug in your own Python scoring code, or reuse the built-ins from eval-ai-library. Every metric produces the same reviewable evidence.
Answer must not offer a refund outside the 30-day window unless the customer cites a billing error.
Compares cited passages against the retrieved chunks. Fails when a citation is fabricated.
Standard groundedness score against retrieved context. Threshold configurable per plan.
A failing metric links back to the exact input, output, model and timing. Thresholds turn into alerts. Triggers turn into CI/CD checks. The evaluation loop keeps running past release day.
Bind a threshold to any metric, wire it to a channel, and route by project. Firing history is stored — you can prove the alert triggered, and when.
Issue a token from Integrations, drop the snippet in your workflow, and every merge runs the acceptance suite before it ships. Seven pipeline tools ready out of the box.
Evaluation targets and judges use your provider credentials. EvaliQA charges for the workspace around them — not resold tokens or an opaque score you can't inspect.
Provider credentials stay in your workspace, encrypted at rest with AES-256-GCM.
Use the hosted product or deploy the same service set inside your perimeter.
Connect any HTTP endpoint. No proprietary agent runtime, no SDK to install.
Export datasets, run results and traces in open, inspectable formats.
Bring your own model keys. Platform pricing covers the evaluation workspace, collaboration and evidence — not the tokens.
Forever. No card required.
$1,490 billed annually.
For your perimeter and process.
Start free with your own provider keys. Export anytime.