Continuous AI-powered QA for your web platform.
Hand us a URL and a test account. AI Auto Test crawls behind your login, generates a test plan you approve, then runs it every day — or every hour — and tells you exactly what broke.
Runs on a schedule
Hourly to weekly. A regression introduced at 14:32 is reported by 14:45.
Evidence on every failure
Screenshots, video, full Playwright trace, console, and network HAR per case.
Claude triages each bug
Real bug vs flaky vs environment vs platform change — with a suggested fix.
How it works
Three stages. You stay in the loop where it matters — and out of it where it doesn't.
01 — Scan
Crawl your platform behind login. Build a structured site map.
Point us at your URL and a test account. A headless Playwright crawler authenticates, walks your authenticated surface, and captures every page, form, button, and API call it sees — desktop and mobile screenshots included.
- Three scan modes: Quick (5 min, depth-2), Deep (full site), Guided (Claude decomposes flows you describe in plain English)
- Authenticates with your stored test credentials — never your real users
- Egress restricted to allowed domains; never wanders to third parties
- Versioned site maps with a diff view between releases
02 — Plan
Claude turns the site map into a test plan you actually review.
Cases grouped by feature, each with preconditions, steps, expected result, severity, and the credentials role to use. Every step is grounded in a real selector or URL — edit a step, the platform re-verifies before saving.
- Edit any case in a three-column editor with grounded selectors
- Section-level updates from a URL, a text template, or a button
- Grounding probes catch broken edits before they hit production runs
- Approve once; sub-cases re-verify only when the underlying site drifts
03 — Run + triage
Inngest schedules. Playwright executes. Claude triages every failure.
Approved cases compile into Playwright scripts. Inngest orchestrates parallel workers, retries flaky cases up to three times, then hands genuine failures to Claude for classification — real bug vs flaky vs environment vs platform drift — with a suggested fix.
- Parallel worker fan-out per case, with per-workspace concurrency caps
- Pre-flight credit checks pause runs cleanly between cases — never mid-case
- Bug reports auto-dedupe by failure signature across services and time
- Hybrid + Pure Human tiers add Ubalitics staff verification on top
Fits into the tools you already use
Trigger runs from your pipeline, get notified where your team lives, and integrate programmatically.
Trigger from your CI/CD
Fire a regression test on every deploy without leaving your pipeline.
- GitHub Actions
- GitLab CI
- Vercel deploy hooks
Get notified where you work
Bug alerts and run summaries land in the same channels your team already watches.
- Slack
- Telegram
- Email digests
- Signed webhooks
Programmatic access
Call /api/v1 from CI, an AI agent, or any HTTP-capable system. Sandbox keys return mock results so integrations test safely.
- REST API
- AI agent tools
- Sandbox keys (pk_test_*)
Three tiers, one platform
From always-on AI runs to humans executing the whole plan by hand.
Pure AI
Continuous Playwright regression coverage.
Hybrid AI + Human
AI runs plus human verification on every bug.
Pure Human
Testers execute the whole plan, signed reports.
Frequently asked questions
Ready to start testing?
Sign in, register your first platform, and run a Quick scan in under 5 minutes.