A $99 penetration test sounds wonderfully simple. Someone has asked for a pentest, the deadline is close and, for less than the cost of a team lunch, the problem appears to go away.
I understand the appeal. Security requests often arrive with very little explanation and a great deal of implied urgency. Nobody wants to spend weeks comparing providers when a customer, insurer or tender is waiting.
Automation can genuinely help. It can explore faster, repeat checks consistently and give skilled testers more time to investigate the things that matter. We use it ourselves. The awkward bit is working out how a $99 service pays for enough of that work to produce a result you can trust.
Start with the uncomfortable maths
The $99 has to pay for the whole service. That includes infrastructure, discovery, model inference, repeated exploration, orchestration, evidence capture, validation, deduplication, reporting, support and the provider's operating costs.
Whatever remains for AI inference will be less than $99, potentially much less. That is before a person has looked at the result or answered a question.
A thorough assessment cannot simply read a homepage once and produce a credible report. Modern applications contain routes, APIs, authentication flows, roles, parameters, error states and business processes. Useful testing needs to explore those surfaces, form hypotheses, test them safely, revisit promising paths and compare observations across the application.
That consumes time and compute. It also requires repeated reasoning. A model may need to inspect the same behaviour from several angles before it can distinguish a real weakness from an odd but harmless response.
At $99 all-in, something has to give. Perhaps the scope is tiny. Perhaps the service is subsidised. Perhaps the provider has found a genuinely clever operating model. Any of those could be reasonable, but you deserve a straight answer.
Automation is useful when it has room to work
The cheapest possible model call is a strange thing to optimise in a penetration test. You want enough exploration to find less obvious paths, enough context to understand them and enough validation to avoid sending your engineers on a very expensive wild-goose chase.
Ask the provider:
- How much of the application is actually explored?
- Does testing follow links, APIs, authentication states and multi-step workflows?
- Can the system revisit and deepen an investigation, or does each surface receive one pass?
- Are potential findings validated before they reach the report?
- Are related observations deduplicated into one clear underlying issue?
- Does the evidence show what happened and why it matters?
- Is the remediation specific enough for an engineer to act on?
- What happens when you question a finding or need help interpreting it?
Good automation expands coverage and consistency. A severely constrained budget can still force shallow exploration, small context windows, limited validation and templated reporting, however capable the underlying technology may be.
A long report can still leave you alone
There is a particular kind of disappointment in opening a security report and finding 40 pages of alerts with no clear sense of what to do next. It is the security equivalent of receiving flat-pack furniture without the instructions.
False positives create unnecessary anxiety and send engineers towards problems that are not real. Poor deduplication can turn one underlying issue into a dozen findings. Generic remediation leaves your team to work out whether the advice applies to your architecture at all.
Even a technically correct observation may provide little assurance. “A security header is missing” does not tell you whether an attacker can use that gap, what the business impact could be or which fix should come first.
The report should reduce uncertainty. It should help your team move from “Are we exposed?” to “Here is what matters, here is the evidence, and here is what we will fix.”
Ask where your vulnerability data goes
Testing can expose sensitive information about your URLs, technology, configuration, authentication behaviour and weaknesses that have not yet been fixed.
Before submitting a target, ask:
- Where is the data processed and stored?
- Which service providers and AI models can receive it?
- In which countries does processing occur?
- How long are prompts, responses, evidence and reports retained?
- Can any of that information be used for model training or product improvement?
- Who can access it, and how is it deleted?
Price and polished marketing will not answer those questions. A credible provider should explain its data handling plainly.
Buy the level of assurance you actually need
A $99 automated scan may be perfectly useful for a quick hygiene check when its limits are clear. That can be a perfectly sensible product.
The standard changes when a customer, insurer, auditor, tender panel or leadership team will rely on the result. In that situation, you need enough coverage, validation and evidence to support a real decision. You also need findings your team can act on and a provider prepared to stand behind the work.
Before you buy, ask the awkward question: after everyone and everything involved has been paid, how much testing can $99 realistically contain?
A good answer leaves you clearer and more confident. Box-ticking relief tends to wear off quickly.