Offensive SecurityJune 7, 2026 · 13 min read

How to choose a penetration testing company

The brief is the same everywhere; the work is not. See how to tell a real offensive team from a scan with an invoice, and the questions to put in your RFP.
A penetration tester reviewing findings and attack paths on a laptop.

Choosing a penetration testing company comes down to one question. Can this firm find the ways an attacker would actually break into your systems, explain them in terms your engineers can fix, and give you evidence an auditor will accept. Most buyers get distracted by logos, certifications on a slide, and a low day rate. Those signals tell you almost nothing about the quality of the testing. The work that matters happens in the hands of the individual testers assigned to your scope, and in the report you receive at the end. This guide explains what to look for, what to ignore, and how to run a selection process that protects you from buying a scan dressed up as a pentest.

We do this work. We are senior engineers who scope, test, and write reports for regulated companies in finance, healthcare, technology, infrastructure, and government, and increasingly for teams shipping AI systems. We have seen what separates a useful engagement from an expensive PDF. The patterns below come from that experience, and they apply whether you are buying your first test or replacing a vendor that disappointed you. If you want the short version, read the section on red flags first, then come back to the start.

What a penetration test actually is, and what it is not

A penetration test is a manual, goal-driven security assessment in which a tester acts like an attacker against an agreed scope. The tester chains weaknesses together to demonstrate real impact, such as reading data they should not see, escalating privilege, or moving from one system into another. The output is a narrative of how they got in, ranked findings with proof, and concrete remediation guidance. That is different from a vulnerability scan, which is an automated tool that matches your systems against a database of known issues and produces a list. Both have a place, but they are not the same product, and you should never pay pentest prices for scan output.

The distinction matters because it drives everything else in your selection. A scan finds missing patches and weak configurations. A pentest finds the logic flaw in your checkout flow, the authorization gap that lets one tenant read another tenant's records, and the chain that turns a low-severity bug into a full compromise. If you are unsure which you need, we wrote a deeper comparison in pentest vs scan and a broader walkthrough of testing types in what is VAPT. For most regulated buyers the answer is both, run continuously, with manual testing on the parts of your estate that carry real risk.

The types of testing, and matching them to your risk

Penetration testing is not one service. The label covers several disciplines, and a firm that is excellent at one may be mediocre at another. Before you compare vendors you need to know which type your scope demands, because the skills, tooling, and report structure differ in each case. Buying a generalist when you need a specialist is one of the most common and expensive mistakes we see.

  • Web application testing focuses on authentication, authorization, business logic, and the OWASP Top 10 classes of flaw, and it rewards testers who understand how your application is supposed to behave, covered in our web application penetration testing guide.
  • API testing targets the endpoints behind your apps and integrations, where broken object level authorization and excessive data exposure dominate, mapped to the OWASP API Security Top 10 and explained in API security testing.
  • Network penetration testing examines external and internal infrastructure, segmentation, and the paths an attacker takes once inside, which we break down in network penetration testing.
  • Cloud assessments review identity, configuration, and trust boundaries across providers like AWS, Azure, and GCP, where misconfigured roles cause more breaches than exploits, detailed in cloud security assessment.
  • AI and LLM testing probes prompt injection, data leakage, and abuse of model-backed features, a young discipline where most firms have little real experience, illustrated in prompt injection in the real world.

Match the test to where your risk concentrates. A SaaS company with a single large web application and a public API should buy deep application and API testing, not a broad network sweep. A bank with a sprawling internal estate needs internal network and segmentation testing alongside application work. If a vendor proposes the same scope for both, they are selling a template, not an assessment.

How a good engagement runs, step by step

The process tells you as much about a firm as the credentials. A serious engagement follows a clear sequence, and you should expect each stage to be visible to you rather than hidden behind a portal. When a vendor cannot describe how they will spend the days you are paying for, that is a signal the work will be thin.

  • Scoping comes first, where the testers, not a salesperson, agree the targets, the rules of engagement, the test accounts, and the goals that define success for the assessment.
  • Reconnaissance and mapping follow, where testers enumerate the attack surface and understand how the system is built before they attack it, which is why external attack surface work in external attack surface management often feeds the scope.
  • Active testing is the core, where testers manually probe, exploit, and chain weaknesses, validating each finding so you never receive a false positive copied from a tool.
  • Reporting turns findings into a document with an executive summary, technical detail, reproduction steps, evidence, and remediation guidance ranked by real business impact.
  • Remediation support and retesting close the loop, where the firm verifies your fixes actually work and issues a clean letter or updated report you can show auditors and customers.

Ask any vendor to walk you through these stages for your specific scope. The quality of their answer, especially on scoping and retesting, predicts the quality of the engagement better than any certificate on the wall.

What you should actually receive at the end

The deliverable is the product you are buying. A weak report buries the findings that matter, copies tool output verbatim, and gives generic advice that your engineers cannot act on. A strong report reads like a senior engineer sat down with your team and explained exactly what they did, what they found, and what to fix first. You should be able to hand it to a developer and to a board member and have both groups understand the parts that concern them.

  • An executive summary that states the overall risk posture in plain language, without jargon, so a non-technical leader understands what is at stake.
  • A findings section where each issue carries a severity, a clear description, the steps to reproduce it, evidence such as screenshots or requests, and specific remediation guidance.
  • A narrative of attack paths that shows how individual weaknesses combined into real impact, because a chain of three medium issues is often more dangerous than one high.
  • A remediation roadmap that ranks fixes by effort and impact, so your team knows what to do this week versus this quarter.
  • A retest provision and a formal attestation or letter you can share with auditors, customers, and partners who ask for proof of testing.

Ask for a sanitized sample report before you sign. Read it the way your engineers will. If you cannot tell what to fix from it, your team will not be able to either, and the engagement will have produced a document instead of an outcome.

The value of a penetration test is the report, the retest, and the judgment of the engineer who produced both.

How to evaluate the people doing the work

Penetration testing is a craft, and the firm's brand is a poor proxy for the skill of the individual assigned to your scope. Large vendors routinely win contracts on reputation and then staff junior testers running automated tools. The work suffers and you never see the gap until you compare against a better firm. Push past the brand and ask about the humans who will touch your systems.

  • Ask who specifically will test your scope, what their background is, and whether the same person who scopes the work will perform it.
  • Ask whether testing is manual and goal-driven or whether the engagement leans on a scanner with light human review, which our piece on PTaaS vs pentest vs automated scanning untangles.
  • Ask for a redacted report written by the actual tester, not a polished marketing sample produced by someone else.
  • Ask how they handle a finding that needs creativity, such as a business logic flaw, and listen for a real story rather than a process diagram.
  • Ask whether they will get on a call to walk your engineers through findings, because the firms that avoid this usually cannot defend their work.

Certifications held by individual testers, such as OSCP or CREST-aligned qualifications, are a reasonable baseline. They prove someone passed a hard practical exam. They do not prove the firm will assign that person to you, which is why the staffing question matters more than the certificate count on a capability statement.

Independence, and why it changes the result

An independent testing firm has no incentive to soften findings or to sell you the product that caused the weakness. Many large vendors test systems built or managed by another division of the same company, which creates a quiet pressure to underreport. Independence is one reason we built Raptoric the way we did, and it is a fair question to put to any vendor. Ask whether they sell the tools or managed services they will then assess, and whether their compensation depends on a clean result. The honest answer shapes how much you can trust the report. Our offensive testing approach is described on the offensive security service page, where independence is the default rather than a feature.

Red flags that should end the conversation

Some signals reliably predict a disappointing engagement. When you see several of these together, walk away regardless of price or brand. A low day rate is never worth a report your auditor rejects or a test that misses the flaw that later gets exploited.

  • A fixed price quoted before any scoping conversation, which means the firm is selling a template and will spend the same hours on every client.
  • A sample report that is mostly scanner output, with severity ratings copied from a tool and no narrative of how findings chain together.
  • Reluctance to name the testers or to put them on a call with your engineers, which usually hides junior staffing.
  • No retest included, so you pay again to confirm your own fixes, or worse, never confirm them at all.
  • Vague scoping that lists technologies rather than goals, because a test without defined objectives produces a list without priorities.
  • Pressure to skip a kickoff and start immediately, which trades the planning that makes testing useful for speed that looks good on a timeline.

The inverse of each red flag is a green flag. A firm that scopes carefully, names its testers, includes retesting, and writes reports your team can act on is worth more than a cheaper alternative that does none of those things.

Cost, effort, and what drives the number

Penetration testing is priced by the number of expert days the scope requires, so the honest answer to what it costs is that it depends on what you ask the firm to test and how deeply. The size and complexity of the application, the number of user roles, the size of the network, and the depth of testing all move the figure. Be suspicious of a quote that arrives before anyone has understood your scope. We break the drivers down in detail in penetration testing cost, and we explain the broader service landscape in penetration testing services explained. Your internal effort matters too. Provisioning test accounts, preparing a staging environment, and assigning an engineer to answer questions will make the days you pay for far more productive.

When to test, and how often

Annual testing is the floor for most regulated companies, and it is rarely enough on its own. Systems change continuously, and a test is a snapshot of one moment. The right cadence depends on how fast you ship and how much risk each release carries. Test before a major launch, after a significant architectural change, and on the schedule your compliance obligations demand. Between full tests, continuous attack surface monitoring and detection coverage close the gap, which is why teams pair offensive testing with detection through services like those in managed detection and response and the firm's own threat detection and response capability. A point-in-time test plus continuous monitoring beats one large annual exercise that leaves you blind for the other eleven months.

How testing connects to your compliance obligations

For regulated buyers, testing is not optional, and the framework you answer to shapes what you need. The NIS2 Directive (EU) 2022/2555 requires essential and important entities to manage cyber risk and test their measures. DORA, Regulation (EU) 2022/2554, has applied to financial entities since January 2025 and expects regular threat-led testing of critical systems. ISO/IEC 27001:2022, with its 93 Annex A controls across four themes, treats technical assessment as evidence that your controls work in practice. SOC 2 examinations against the Trust Services Criteria expect security testing as part of a credible program. A good report gives you the evidence each of these regimes asks for, written so an assessor accepts it without a second engagement.

  • For NIS2, testing supports the risk management and incident readiness obligations, which we summarize in NIS2 explained and on the NIS2 compliance page.
  • For DORA, threat-led testing of critical functions is central, with a practical walkthrough in the DORA compliance checklist and the difference from NIS2 set out in NIS2 vs DORA.
  • For ISO 27001, testing produces evidence your Annex A controls operate as intended, covered in the ISO 27001 certification guide.
  • For SOC 2, testing strengthens the security criteria, though a clean report is not the same as real security, a point we make in why SOC 2 is not security and on the SOC 2 compliance page.
  • For AI systems, testing intersects with emerging obligations under the EU AI Act, where model-backed features create new classes of risk.

Treat the framework as the minimum, not the goal. Passing an audit and being secure are related but separate outcomes, and the firm you choose should care about both. A vendor that only optimizes for the audit will leave you compliant and exposed at the same time.

Where to start

Choosing well is mostly about asking the right questions and refusing to be rushed. Define where your risk concentrates, demand a sample report, insist on knowing who will test your systems, and check that retesting and a usable attestation are included. The firm that scopes carefully and writes for your engineers will give you more than a cheaper one that ships scanner output. If you want to see how we run offensive engagements, read the offensive security service page, and when you are ready to define a scope for your environment, book a scoping call and we will walk through it with you.

Frequently asked questions

How long does a penetration test take?
Most application or network tests run between one and three weeks of active testing, plus time for scoping beforehand and reporting afterward. The exact duration depends on the size and complexity of the scope. Be cautious of any firm that promises a thorough manual test of a large application in a day or two, because that timeline only fits an automated scan.
What is the difference between a pentest and a vulnerability scan?
A scan is an automated tool that lists known issues, while a penetration test is a manual assessment where a person exploits and chains weaknesses to show real impact. Scans are fast and cheap and good for continuous coverage. Pentests are slower and more expensive and find the logic and authorization flaws that scanners miss. Most regulated companies need both.
How often should we run a penetration test?
At least annually, and additionally before major launches and after significant changes to your architecture. If you ship frequently or carry high risk, pair annual testing with continuous attack surface monitoring and detection so you are not blind between tests. Your compliance framework may set a minimum cadence, but treat that as the floor.
Do certifications guarantee a good test?
No. Certifications held by individual testers, such as OSCP, prove a baseline of practical skill, and firm-level accreditations show process maturity. Neither guarantees the firm will assign a skilled tester to your scope or write a report your team can act on. Ask who will do your work and read a sample report before you decide.
What should we prepare before a test starts?
Provision test accounts for each user role, prepare a stable staging environment that mirrors production, document the scope and goals, and assign an engineer to answer the testers' questions. Good preparation lets the testers spend their days finding issues rather than waiting on access, which directly improves what you get for the budget.

Sources

  1. 1NIST. SP 800-115: Technical Guide to Information Security Testing and Assessment. National Institute of Standards and Technology, 2008. Link
  2. 2OWASP. Web Security Testing Guide. OWASP Foundation, 2024. Link
Related service
Offensive Security
Want this tested on your own systems?
Our team will scope it with you on a 30-minute call.
Book a scoping call