AskYourQA
← Back to Blog
test automation · October 10, 2026

What to Automate in Testing

Learn what to automate in testing with a clear framework for choosing tests, ranking ROI, and avoiding anti-patterns. Practical guide for QA leaders.

test automationQA strategyautomation ROIwhat to automatetest prioritization

Most teams answer what to automate in testing with “everything repetitive.” That advice creates bloated suites, slow pipelines, and red builds nobody trusts. A test that fails randomly, breaks after harmless UI changes, or provides no clue about the defect isn't protecting the release. It's consuming attention.

The better target is trusted signal per release. Automate the checks that protect revenue, customer access, compliance, security, and core operations. Favor behavior that's stable, repeatable, and easy to diagnose. Leave ambiguous judgment, one-off investigation, and rapidly changing interfaces to people until the product settles.

The cost of late discovery makes this a business decision, not a tooling preference. A NIST analysis of software failures estimated that inadequate testing and defect removal cost the U.S. economy approximately $59.5 billion annually, with users absorbing about 64% of that cost and developers absorbing 36%. The same analysis estimated that improved testing could reduce the impact by roughly one-third, or about $22.5 billion. The practical lesson is simple: move detection toward the point where defects enter the system, but don't confuse more scripts with better prevention.

Table of Contents

Why Most Teams Automate the Wrong Tests

The easiest test to automate is rarely the most valuable one. A stable button, a familiar login path, or a visible form field gives a QA engineer quick progress. Meanwhile, the difficult risks stay exposed: a tenant-isolation failure, an invalid API contract, a retry that charges twice, or a settlement job that drops a transaction without notice.

That trade-off looks good in a dashboard because script counts rise. It looks bad during a release because the suite takes too long, produces ambiguous failures, and still misses the defects executives care about. Every unreliable result teaches developers to rerun first and investigate later. Eventually, a red pipeline becomes background noise.

Senior QA rule: A test earns automation when its failure should change a release decision.

Automation should detect defects early and repeatedly, not merely replace manual clicking with code. Requirements-linked checks, unit tests, integration tests, API contracts, security controls, and critical end-to-end journeys usually create stronger feedback than broad visual coverage. NIST's software testing guidance recommends selecting cases from anticipated risks and executing higher-risk cases first.

Start with the business journey, then choose the cheapest reliable layer that can verify it. A payment authorization belongs in an API or integration check before it belongs in a browser flow. A permission boundary deserves an executable security assertion. A new checkout layout may need human exploration before selectors become stable enough for automation.

The standard I use with a CTO is strict:

  • Protect critical outcomes: Automate behavior tied to revenue, access, compliance, customer trust, or core operations.
  • Prefer deterministic checks: Automate inputs and expected outputs that can be reproduced precisely.
  • Demand useful failures: A failure should identify the broken contract, boundary, or workflow.
  • Retire noise: A test that repeatedly fails for environmental reasons must be fixed, quarantined, or removed.

The result is often a smaller suite. That is a feature, not a failure.

The Four Criteria That Decide What Is Worth Automating

Use four filters before writing a script. They turn a vague backlog discussion into a fast decision about value, effort, and trust.

An infographic titled The Four Criteria That Decide What Is Worth Automating, listing business value, frequency, stability, and cost.

Business criticality

Ask what happens if this behavior fails in production. Does the customer lose access, money, data, or confidence? Does the company breach a control or interrupt an essential operation? A tenant authorization check scores high. A minor spacing rule in an internal settings screen does not.

A useful test is to name the consequence without mentioning QA. “The customer can't submit an order” is a strong automation candidate. “The test team wants more coverage” isn't a business case.

Stability

Automation captures an agreement between the product and the test. If the underlying behavior changes every sprint, that agreement hasn't matured. Automating a stable API contract is sensible. Automating a prototype screen whose labels and navigation are still being debated creates maintenance work before the design is finished.

Stability doesn't mean the product never changes. It means changes are intentional, reviewable, and represented by a clear requirement.

Repeatability

A check that runs on pull requests, merges, releases, or recurring data sets benefits from automation. A scenario performed once to investigate an unexpected behavior usually doesn't. Ask how often the team will need the answer and whether a person can execute it consistently under delivery pressure.

A boundary test for malformed input may run repeatedly across builds. A one-time exploratory investigation into confusing onboarding behavior should remain human-led until the behavior becomes defined.

Diagnostic value

When the test fails, can the developer act without reconstructing the entire system? A contract assertion that reports an unexpected status, schema field, or authorization result has high diagnostic value. A browser test that reports only “element not found” has low value, even if the business scenario matters.

The four criteria should be applied together. High business value cannot rescue a test that is unstable and impossible to diagnose. If the risk is high but the interface is volatile, move the check down to a more stable layer or keep the scenario exploratory until the product settles.

Where Each Test Type Belongs in the Pipeline

The right test in the wrong pipeline stage becomes a delivery bottleneck. Put fast, precise checks close to code changes. Reserve expensive environment-wide verification for places where lower layers cannot provide enough confidence.

A diagram explaining the software testing pyramid and where each test type fits into the CI/CD pipeline.

Pull requests should prove local correctness

Unit tests belong on the pull-request path when they can isolate a function, rule, or transformation. They should cover calculations, validation, state transitions, and known regression cases. Integration tests can join this gate when they use controlled dependencies and provide a clear failure location.

Don't push a browser journey into every pull request if an API or service-level check can verify the same rule. The browser should prove that the critical path is wired together, not repeat every lower-level assertion.

Merges and scheduled runs should exercise contracts and interactions

API and contract checks are strong candidates for merge, staging, and scheduled pipelines. For REST, gRPC, GraphQL, and WebSocket interfaces, automate schemas, status or error behavior, authentication, authorization, boundary inputs, and compatibility expectations. These checks expose failures between services without paying the full cost of a UI session.

NIST's combinatorial testing research supports systematic coverage of meaningful interactions among configuration parameters, platforms, protocols, and operating conditions. Use a small smoke set for every change, then run broader combinations on scheduled or release pipelines.

Releases need a narrow end-to-end gate

A release gate should contain only the journeys that must work for customers to use the product. Login, account creation, checkout, payment confirmation, critical data access, and a primary workflow may qualify. Keep the list short enough that failures receive immediate attention.

Security checks should target authorization boundaries and threat-model risks throughout the pipeline, with deeper scans and environment checks scheduled appropriately. Performance tests belong in environments that represent the relevant workload and infrastructure. They shouldn't delay every small code change unless the change directly affects a performance-critical path.

AI features need their own evaluation checks around routing, output structure, permissions, safety, retrieval, and tool use. A green UI smoke test doesn't prove that an agent selected the right tool or refused an unsafe request.

A practical reference for separating fast functional checks from broader validation is functional testing versus unit testing.

A Prioritization Matrix for Your Test Backlog

Don't begin with a framework. Begin with a spreadsheet containing every candidate test and six scores from 1 to 5. The score isn't a scientific truth. It forces the team to make trade-offs visible before engineering time gets committed.

DimensionScore 1 (Low)Score 3 (Medium)Score 5 (High)
Business criticalityCosmetic or low-impact behaviorImportant workflow with a workaroundRevenue, access, compliance, or core operation
Failure probabilityRarely changes or failsModerate change or defect exposureFrequent change or known failure risk
Execution frequencyOne-off or occasionalRepeated during releasesNeeded on most builds or changes
Diagnostic valueAmbiguous failure locationRequires investigationPinpoints a contract, rule, or boundary
Flakiness riskDeterministic and isolatedSome environmental sensitivityTiming, shared state, network, or unstable selectors
Data or environment costSimple local fixturesManaged service or staging dataExpensive, scarce, or difficult to reproduce

For the first four dimensions, higher is better. For flakiness risk and environment cost, higher means more danger. Either subtract those values from the benefit score or treat them as veto conditions. A high-risk payment test may still deserve automation, but not as a brittle UI script if the same risk can be checked through a stable service boundary.

Apply weights that match the business

A high-velocity SaaS team should weight business criticality, execution frequency, and diagnostic value heavily. A regulated enterprise should give additional weight to compliance exposure, evidence quality, and audit traceability. Don't use one universal formula across departments with different consequences.

A simple operating rule works well:

  • Automate first: High criticality, frequent execution, stable behavior, strong diagnostics, and manageable environment cost.
  • Redesign before automating: High business risk paired with high flakiness or poor diagnostics.
  • Explore manually: Low repeatability, changing requirements, subjective judgment, or one-off investigation.
  • Retire: Low criticality, low diagnostic value, infrequent execution, and maintenance cost that exceeds risk reduction.

Review the matrix with product and engineering, not QA alone. Product owners know which journeys affect customers. Engineers know which boundaries are stable. Security and compliance stakeholders know which evidence matters. The risk-based software testing framework gives teams a useful basis for making that discussion explicit.

Decision test: If nobody can explain what action follows a failure, don't automate the test yet.

Real Candidate Tests Worth Automating First

Consider a payment flow from price calculation through settlement. I wouldn't start by scripting every browser click. I'd split the journey into contracts and controls, then add one narrow end-to-end smoke test to prove the wiring.

The pricing API contract comes first. Verify accepted inputs, currency handling, required fields, error responses, and the structure consumed by the checkout service. This test is repeatable, runs quickly, and points directly to a broken interface when it fails.

Next, automate idempotency under retries. A payment request that reaches the processor and is submitted again shouldn't create an unintended second charge. This is a business-critical behavior with a clear expected result, and it belongs below the UI where retry conditions can be controlled precisely.

Webhook signature verification deserves an executable security boundary. Send valid, invalid, altered, and replayed webhook payloads, then assert which events the application accepts. Don't rely on a successful checkout screen to prove that an untrusted callback can't change payment state.

The refund authorization rule should test identity, role, tenant, and object ownership. A low-privilege user must not refund another account's transaction. This is exactly the kind of threat-model risk that should become a repeatable assertion rather than a manual reminder.

Finally, automate settlement reconciliation against controlled fixtures. The test should detect mismatches between payment records, ledger entries, and settlement status. Its diagnostic output should identify the transaction and reconciliation rule that failed.

The browser smoke test still has a place. Run the smallest realistic path: sign in, select a product, submit payment in a safe environment, observe confirmation, and verify the resulting state. It proves the main customer journey works across the assembled system.

Keep these activities human-led:

  • Refund dispute reasoning: A person should assess unusual evidence and policy interpretation.
  • New checkout UX: Exploratory testing should find confusion, awkward flows, and accessibility problems while the design is changing.
  • Unfamiliar processor behavior: Investigate first, then automate the stable rule discovered.
  • Visual judgment: Use people for subjective layout and interaction quality.

The judgment isn't “automate payments.” It's “automate the stable, consequential claims inside the payment journey.”

Automation Anti-Patterns That Quietly Kill Pipelines

A bloated suite usually reveals its history. Someone automated a demo, copied a scenario, added a retry, and never removed the test when the feature changed. The result feels complete while becoming less useful every month.

A comparison chart showing healthy automation signals versus anti-patterns that negatively impact software testing pipelines.

Counting scripts instead of protecting decisions

Script count is easy to report and difficult to connect to customer risk. Teams add low-value cases because the metric rewards volume. The correction is to report prevented escapes, useful feedback, and actionable failures instead.

Automating exploratory work

Exploration changes direction as the tester learns. A fixed script removes that discovery behavior and preserves yesterday's questions. Keep exploration manual, then convert a discovered, repeatable failure into a regression test once the expected behavior is clear.

Accepting flaky tests as normal

Google's analysis of 4.2 million tests found that almost 16% showed some flakiness, while about 1.5% of test runs were flaky. It also found that 84% of pass-to-fail transitions involved flaky tests, and reruns consumed 2% to 16% of compute resources, as reported in this study of flaky tests. Those findings explain why “rerun until green” is not a quality strategy.

A retry can help classify a transient failure, but it must preserve the original result. Repeated nondeterminism needs an owner, quarantine status, and repair deadline.

Driving every assertion through the UI

A UI test often combines application logic, network timing, browser state, test data, and selectors. If the rule can be asserted through an API or service boundary, put it there. Keep the UI layer for a small number of customer-critical compositions.

Leaving the suite without an owner

Unowned tests rot. Nobody updates the fixture, investigates a changed contract, or decides whether a failure blocks release. Assign ownership by domain and require each failure to produce a decision: fix, quarantine, redesign, or retire.

What to Automate When the Product Is AI-Powered

An AI feature changes the risk surface. A conventional regression test can confirm that a prompt field submits and a response renders, while missing whether the system exposed private data, selected an unauthorized tool, ignored a refusal rule, or returned malformed structured output.

Stack Overflow's 2025 developer survey found that 84% of surveyed developers use or plan to use AI tools in development, while 46% said they don't trust AI output accuracy. The same survey reported that 31% currently use AI agents. Adoption and confidence are therefore moving at different speeds, which makes evaluation automation a release control rather than an optional research exercise.

Automate invariants, not one perfect answer

Don't assert that an open-ended assistant must produce one exact paragraph. Assert properties that must hold across acceptable responses:

  • Schema validity: Structured output parses and contains required fields.
  • Permission boundaries: The model or agent cannot retrieve another tenant's data or invoke a tool beyond its authority.
  • Tool constraints: Calls use approved tools, parameters, and sequencing.
  • Safety behavior: Disallowed requests receive the required refusal or safe handling.
  • Retrieval grounding: Responses meet defined evidence or citation requirements where those requirements apply.
  • Operational limits: Latency, token usage, and cost remain within the product's accepted thresholds.

Run deterministic checks in CI when possible. Use curated real-world scenarios and offline evaluation for quality patterns that need broader comparison. Route ambiguous judgments, nuanced tone, and novel failure modes to human review.

The right question isn't whether every response matches a reference answer. It's whether the system preserves its contracts while giving reviewers enough evidence to assess quality.

A practical overview of this discipline is AI testing for modern products. Treat prompt routing, retrieval, tool authorization, and output validation as separate test surfaces, because a passing interface check doesn't establish that the underlying agent behaved safely.

Your 30-60-90 Day Plan to Automate Less but Better

A cleanup program needs a deadline, owners, and removal decisions. Don't launch another framework project before you know which tests deserve to survive.

A 30-60-90 day strategic plan infographic for automating software tests efficiently, focusing on scoring, pruning, and measuring.

Days 0 to 30 should expose the backlog's real value

Export the existing suite into a reviewable list. Score each candidate for criticality, failure probability, frequency, diagnostic value, flakiness risk, and environment cost. Mark tests that have no owner, no clear assertion, or no release decision attached.

During this phase, remove tests for retired behavior and quarantine tests that repeatedly produce unexplained failures. Replace duplicated UI checks with API or contract checks where the business rule is already represented at a more stable layer.

Days 31 to 60 should place and prune

Map the surviving tests to unit, integration, API, security, performance, AI evaluation, smoke, or broader end-to-end layers. Put fast deterministic checks on pull requests, contract and interaction coverage on merge or scheduled pipelines, and a narrow customer journey set on release gates.

Create failure governance. A failed test should retain its first result, artifacts, duration, environment, retry history, and classification. Assign an owner and a repair deadline for quarantined tests. Google's research shows why this matters: uncontrolled flakiness consumes compute and weakens confidence in the entire pipeline.

Days 61 to 90 should measure trust

Add evaluation harnesses for AI features if the product uses models, agents, retrieval, or tool calls. Then audit the remaining low-signal tests. Keep a test only when its release-risk reduction justifies its execution and maintenance cost.

Track four operating metrics:

  1. Prevented escaped defects, tied to failures caught before production.
  2. Trustworthy feedback time, measured by how quickly a credible result reaches the engineer.
  3. Failure diagnosability, based on whether the first investigation identifies a likely cause.
  4. Actionable pipeline results, the share of results that lead to a clear decision.

DORA's 2025 research, based on nearly 5,000 technology professionals worldwide, found a positive relationship between AI adoption and throughput but a negative relationship with delivery stability. Its 2025 State of AI-assisted Software Development report emphasizes automated testing, mature version control, and fast feedback loops as control systems for that tension.

Don't reward the team for keeping every script. Reward it for making releases safer with fewer distractions. If a test cannot protect a meaningful decision, diagnose a credible failure, or run reliably enough to earn trust, retire it.

AskYourQA designs and implements test automation across frontend, backend APIs, mobile products, security, performance, and AI features, with CI/CD integration for smoke and regression feedback. If your backlog is noisy or your critical journeys remain under-tested, visit AskYourQA to turn this prioritization method into an owned automation system.

Want this level of confidence in your releases?

We build test automation frameworks 5× faster than in-house teams. Free 20-min call — we map your critical flows.

Book a call