AI applied to testing

AI writes tests. It still doesn't decide which ones to write.

We are running it in delivery for two years, not just as a demonstration. What has changed is not the nature of the job: it is the volume a consultant covers in a day.

Make an appointment

Where she reports

Four uses, listed in order of observed profitability.

1 — The first draft

Generate test cases from a specification or a story. The benefit is not the quality of the text, it's no longer starting from a blank page.

  • Proposed load cases and boundary conditions
  • Mandatory human proofreading, always

2 — Anchor maintenance

Fixing a selector that shifted after a screen redesign. It's the most thankless task in automation, and the most mechanical.

  • Fewer pointless red tests
  • The true cost item of an automated base

3 — Sorting the Failures

Differentiate application faults from unstable environments, and group fifty failures that share the same cause.

  • A red campaign becomes readable again
  • The analysis goes from two hours to ten minutes

4 — Bulk comparison

Comparing thousands of values against a reference document—a brief, a catalog, an interface contract.

  • 25,704 comparisons in one minute on a real fleet
  • Where the human samples, the machine covers everything

What she does not do

She decides nothing. And it is structural, not temporary.

Three judgments remain beyond his reach, and they are the ones that commit.

What deserves to be tested. Weighing whether to cover a payment flow or an administration screen requires knowing what the company loses in each case. No model knows this figure.

What an anomaly costs. A display issue on an internal screen and the same issue on a customer kiosk do not carry the same weight. Prioritization is a business decision.

If a version goes to production. It is a commitment, with an identified person in charge. It cannot be delegated to a probability.

Let's add a more prosaic limit: a model produces the plausible. On a test, the plausible that is false is worse than nothing — it creates a trust that no human is going to reopen. Hence our rule: everything produced by the AI is reviewed, and everything that fails is checked by hand before being declared a defect.

In practice

Two years in production, not an experiment.

2 years
for delivery use, not in a laboratory
5 s
to replay a route, count on 30 to 50 minutes
756
Items of a menu proofread and compared in one minute

What this actually produces: coverage that goes from a sample to the totality. What we were verifying on a few restaurants, we verify across an entire network — not because the tool is smarter, but because it doesn't get tired and doesn't skip line 431.

Frank questions

What we are asked for

Will AI replace testers?

No — and we are well-placed to say so, since we have been running it in delivery for two years. It writes a first draft, maintains selectors, sorts failures, and compares in bulk. What deserves to be tested, what an anomaly costs, and whether a version goes into production remain matters of judgment. What has changed is the volume covered.

Do our data go into a public model?

Not without your explicit consent, and never by default. Depending on your policy, we work with anonymized games, on a model hosted on your servers, or without any AI at all for sensitive areas. This is a decision that must be made at the outset, not along the way.

How to know if a generated test is good?

By the same proofreading as for a handwritten test: is it replayable, is its answer binary, does it test something valuable? A generated case that fails these three questions is deleted, not kept because it's free.

Where to begin without turning everything upside down?

Through anchor maintenance and failure triage: two invisible tasks for the business, with no risk to quality, that free up time starting the first week. Test generation comes afterwards, once review is in place.

Let's talk about what you can't test yet.

A no-obligation audit of your testing process to find out where you really stand.

Make an appointment
en_USEnglish