Continuous AI-powered QA for your web platform.

Hand us a URL and a test account. AI Auto Test crawls behind your login, generates a test plan you approve, then runs it every day — or every hour — and tells you exactly what broke.

Runs on a schedule

Hourly to weekly. A regression introduced at 14:32 is reported by 14:45.

Evidence on every failure

Screenshots, video, full Playwright trace, console, and network HAR per case.

Claude triages each bug

Real bug vs flaky vs environment vs platform change — with a suggested fix.

How it works

Three stages. You stay in the loop where it matters — and out of it where it doesn't.

01 — Scan

Crawl your platform behind login. Build a structured site map.

Point us at your URL and a test account. A headless Playwright crawler authenticates, walks your authenticated surface, and captures every page, form, button, and API call it sees — desktop and mobile screenshots included.

  • Three scan modes: Quick (5 min, depth-2), Deep (full site), Guided (Claude decomposes flows you describe in plain English)
  • Authenticates with your stored test credentials — never your real users
  • Egress restricted to allowed domains; never wanders to third parties
  • Versioned site maps with a diff view between releases

02 — Plan

Claude turns the site map into a test plan you actually review.

Cases grouped by feature, each with preconditions, steps, expected result, severity, and the credentials role to use. Every step is grounded in a real selector or URL — edit a step, the platform re-verifies before saving.

  • Edit any case in a three-column editor with grounded selectors
  • Section-level updates from a URL, a text template, or a button
  • Grounding probes catch broken edits before they hit production runs
  • Approve once; sub-cases re-verify only when the underlying site drifts

03 — Run + triage

Inngest schedules. Playwright executes. Claude triages every failure.

Approved cases compile into Playwright scripts. Inngest orchestrates parallel workers, retries flaky cases up to three times, then hands genuine failures to Claude for classification — real bug vs flaky vs environment vs platform drift — with a suggested fix.

  • Parallel worker fan-out per case, with per-workspace concurrency caps
  • Pre-flight credit checks pause runs cleanly between cases — never mid-case
  • Bug reports auto-dedupe by failure signature across services and time
  • Hybrid + Pure Human tiers add Ubalitics staff verification on top

Fits into the tools you already use

Trigger runs from your pipeline, get notified where your team lives, and integrate programmatically.

Trigger from your CI/CD

Fire a regression test on every deploy without leaving your pipeline.

  • GitHub Actions
  • GitLab CI
  • Vercel deploy hooks

Get notified where you work

Bug alerts and run summaries land in the same channels your team already watches.

  • Slack
  • Telegram
  • Email digests
  • Signed webhooks

Programmatic access

Call /api/v1 from CI, an AI agent, or any HTTP-capable system. Sandbox keys return mock results so integrations test safely.

  • REST API
  • AI agent tools
  • Sandbox keys (pk_test_*)

Three tiers, one platform

From always-on AI runs to humans executing the whole plan by hand.

$99/month

Pure AI

Continuous Playwright regression coverage.

$499/month

Hybrid AI + Human

AI runs plus human verification on every bug.

$999–$14,999/month

Pure Human

Testers execute the whole plan, signed reports.

Frequently asked questions

Ready to start testing?

Sign in, register your first platform, and run a Quick scan in under 5 minutes.