Pentest, red team, TLPT and purple team
These four terms are often used interchangeably, although they describe services with different purposes, costs and prerequisites in organisational maturity. The distinction is not academic: it decides whether the service you commission answers the question you actually have.
A pentest is a planned engagement, coordinated with the client, within a defined scope - a web application, an internal network, an Active Directory estate. The defending team usually knows it is happening, and the goal is to find and describe technical weaknesses together with the path to exploiting them. A red team engagement goes further: the offensive team plays the part of a specific adversary and uses that adversary's full repertoire, including social engineering and physical entry. The crucial difference is that the object of study stops being the infrastructure and becomes the organisation's ability to detect the attack and respond to it - which is why the defenders are usually not told.
TLPT, threat-led penetration testing, is a red team engagement wrapped in regulatory rigour. The process framework is TIBER-EU [4], published by the European Central Bank in 2018, and DORA [2] made that approach mandatory for designated financial entities. The test rests on verified intelligence about threats aimed at the particular institution, runs with the formal involvement of the supervisory authority, and ends with a defined set of reports. A purple team exercise inverts the logic of the other three: the offensive and defensive teams work together in real time, and the aim is not to surprise the defence but to improve detection. It is the cheapest way of turning the results of an earlier test into an actual change in monitoring configuration.
For most organisations a sensible order is straightforward. Local authorities and smaller public bodies can start with an external and internal pentest at a frequency derived from risk, major changes and applicable requirements. Entities covered by the Polish national cybersecurity system act may add a scope-limited red team engagement once they have monitoring worth testing. Under article 26 of DORA, TLPT at least once every three years is mandatory only for the financial entities identified by the competent authority under the applicable criteria, not for every entity within a listed financial subsector.
Methods: which framework to choose
There is no single universal pentest methodology, and no serious provider claims to use one exclusively. In practice several are combined, because each describes a different part of the work well.
The most operational is PTES [6], which divides a test into seven phases - from pre-engagement, through reconnaissance, threat modelling, vulnerability analysis and exploitation, to post-exploitation and reporting. Its strength is the pre-engagement phase, which covers what is most often neglected: agreeing the scope, the rules of engagement and the emergency procedure. The older OSSTMM 3 [7] from 2010 contributes something the others lack - a formal attack surface metric that lets successive tests be compared numerically rather than only descriptively. NIST SP 800-115 [8] is less detailed but carries the weight of a recognised reference and structures network testing well; it has an article of its own.
For the application layer the reference remains OWASP WSTG [9] version 4.2 - a checklist of web application tests organised into twelve categories, from reconnaissance and configuration, through authentication and authorisation, to business logic, the client side and APIs. Its mobile counterpart is the OWASP MASTG together with the MASVS requirements model. For regulated testing the process framework is TIBER-EU [4], and the CREST standard [10] helps verify a provider's competence.
The practical choice therefore follows the object under test rather than the provider's preference: the external attack surface is tested to PTES supplemented by WSTG, the internal network and Active Directory to PTES and NIST SP 800-115, mobile applications to MASTG, and financial sector engagements to TIBER-EU. Declaring a methodology means little if no trace of it is visible in the report.
How much should the tester know
The second axis of choice is how much information the tester is given before starting. In a black box engagement they get nothing and begin exactly where a real attacker from the internet would. That model reproduces the initial conditions of an attack most faithfully, but a substantial share of the paid time goes on reconnaissance the organisation could have supplied in an hour. In a white box engagement the tester has architecture documentation, source code and privileged accounts, which gives the highest coverage and the most findings, at the cost of realism.
In practice grey box dominates, with the tester given user accounts and basic knowledge of the architecture. That matches the most common real scenario - an adversary who has already obtained someone's credentials - and lets the budget go on checking the safeguards rather than guessing at them. The rule is simple: the more the test is meant to measure detection, the less information; the more it is meant to measure completeness of the safeguards, the more.
For a typical local authority with a few dozen workstations and a handful of servers, the sensible arrangement is grey box from the inside and black box from the internet. White box is worth reserving for bespoke applications, where the business logic and the permission model are unique and, without knowing them, the tester can only guess which transition between process steps is an abuse.
What the report has to contain
A pentest report is not a list of vulnerabilities sorted by CVSS score. It is a document written for three different readers asking three different questions.
The board needs two pages without jargon: how many attack paths lead to serious consequences, what the worst realistic scenario is, and how the organisation's overall resilience compares. The security team needs something else - a description organised by attack chains rather than by individual findings. A chain is the sequence of steps from entry point to consequence, and it is the chain that shows how three medium-severity weaknesses together produce a critical outcome that none of them produces alone. The remediation team, in turn, needs a list of items, each with a description, evidence, reproduction steps, a recommendation and a mapping to MITRE ATT&CK [5] and CWE.
The best test of report quality is reproducibility. "Administrative privileges were obtained" is not a finding; a finding is a timestamped sequence of commands that somebody with the right access can repeat and see the same result. A red team report adds a detection timeline - a minute-by-minute record of what the defenders noticed and what they missed. That is usually the most valuable part of the whole document, because it points at specific gaps in monitoring (see the SOC article). A TLPT report follows the structure imposed by TIBER-EU [4]: a threat intelligence report, a test report, a summary for the supervisory authority and a closing document with a remediation plan.
The most common mistakes
The commonest mistake is buying a vulnerability scan under the name of a pentest. Running an automated scanner and transcribing its output into a document is work worth a fraction of the price of a penetration test, and it creates a false sense of having checked. A professional test starts with threat modelling and manual verification; the scanner is a supporting tool in it, not the source of the content. Automated scanning has its own perfectly sensible role - described in the article on vulnerability scanning - but it is a complementary role, performed far more often and far more cheaply.
The second common mistake is too narrow a scope. An engagement limited to a single web application leaves out DNS, open source reconnaissance, the internal network and the employee vector. According to DBIR 2026 [11] the most frequent initial breach vector was exploitation of a vulnerability at 31 per cent, ahead of credential abuse at 13 per cent - precisely the areas such a scope excludes.
The third mistake concerns the provider. The job title is not a protected term in this profession, so it is worth checking certifications, accreditation [10] and experience in an environment resembling your own. A special case is commissioning the test from your own IT provider: it is often cheaper and more convenient, but it produces a report assessing the work of its own author, which removes the very function the test was bought for.
The remaining three mistakes concern what happens around the test. Testing production without an agreed emergency plan can take a service down - old network hardware is sometimes sensitive to traffic nobody would call an attack. A report with no retest of selected findings leaves open the question of whether the fixes worked at all. Finally, after every test it is worth checking what your own monitoring recorded of its activity and which alerts were ignored; that is the only opportunity to measure detection against an attack whose details you already know.
Legislation that requires testing
None of the instruments in force in Poland uses the word "pentest" as an obligation for organisations generally, but several require an outcome that is hard to reach without one. NIS2 [1], in article 21(2), lists both vulnerability handling (point e) and policies to assess the effectiveness of the measures adopted (point f) among the risk management measures; in Poland the applicable instrument is the amended national cybersecurity system act, in force in its new wording since 3 April 2026 (see the NIS2 article). The GDPR [12], in article 32(1)(d), requires regular testing, assessing and evaluating the effectiveness of technical and organisational measures, which covers applications processing personal data.
The most concrete requirements come from the financial sector. DORA [2] divides testing into two levels: a baseline digital operational resilience testing programme binding on all covered entities, and TLPT, required at least once every three years of entities designated by the competent authority - in Poland the Polish Financial Supervision Authority (see the DORA article). Beyond EU law the card industry has its own requirements: PCI DSS version 4.0 mandates penetration testing at least annually and after significant changes.
ISO/IEC 27001:2022 [13] is not legislation, but its annex A names controls whose demonstrated effectiveness leads straight to testing - A.5.7 on threat intelligence and A.8.29 on security testing in development and acceptance. The AI Act [3] forms a category of its own: article 15 requires high-risk systems to achieve an appropriate level of accuracy, robustness and cybersecurity, demonstrated among other things through adversarial testing (see the AI Act article).
Where the market is heading
The fastest growing area is testing systems built on large language models. As they move into everyday work, a class of problems has appeared that a classical pentest does not cover: prompt injection through input data, circumventing model restrictions, leakage of data from context, and abuse of the permissions granted to integrations. The AI Act [3] gives this work a regulatory basis for high-risk systems, but the methods here are visibly less mature than in web application testing.
The second direction is the shift from a single test to a continuous service billed by subscription, with results delivered as they are found and fed into the ticketing system. The model works well where the application changes frequently and an annual test cycle always examines an out-of-date version. It does not replace an in-depth test, though: it changes the frequency, not the depth.
The third direction is a move away from generic testing towards scenarios modelled on the adversaries genuinely interested in the organisation. In the financial sector TIBER-EU [4] fills that role; in other sectors the starting point is often the threat landscape analysis published by ENISA [14] together with the reports of CERT Polska [15], which describe campaigns actually observed against Polish targets.
What to settle before commissioning a test
The value of a test is decided by a dozen or so choices made before the contract is signed. It is worth working through them in this order.
Start with scope and purpose. The scope is a specific list of addresses, domains, applications, network segments and - if the test includes social engineering - groups of staff, together with an equally clear statement of what stays outside it. The purpose is often skipped, yet it shapes the whole engagement: a test meant to confirm compliance looks different from one meant to assess overall resilience, which in turn differs from an in-depth examination of a single application. The purpose determines the access model and the chosen methodology, and both decisions should leave a trace in the report.
The second group of questions concerns the provider and how the test is run. Establish the qualifications of the specific people rather than of the company, and write down the rules of engagement: testing hours, a contact reachable while they run, the escalation path and the range of permitted actions. Decide separately whether the tester may exploit the vulnerabilities found, move laterally once access is obtained, and work on real data - these are decisions of a different weight from scanning alone.
The third group concerns what is left afterwards. Beyond the report format, agree in advance whether a retest of the highest-severity findings is included in the price and within what period, and settle the confidentiality and liability terms: a non-disclosure agreement, a named responsible person on the provider's side, insurance, and a procedure for deleting the data collected during the work.
Frequently asked questions
- Pentest or red team - which should we choose?
If this is the first offensive test, a pentest. Its purpose is to find and fix technical weaknesses. A red team engagement only makes sense after the first two or three pentests, once the organisation has a mature SOC and procedures and wants to know how it would behave in a real crisis.
- How long does a web application pentest take?
A typical CMS application: 5 to 7 person-days. A typical bespoke application such as a citizen portal or a patient portal: 8 to 15 person-days. A large sector-specific application: 15 to 30 person-days. As a fixed-scope product, from PLN 10,000 net; bespoke work is quoted individually.
- Can a pentest break production?
Carried out professionally, very rarely, and only where agreed in advance, as with an exploit or a denial-of-service test. A standard pentest is non-disruptive. In our own project history the rate is below one per cent; the typical event is old layer 2 switching hardware failing under intensive ARP scanning.
- Can we ask our own IT provider to run the pentest?
Independence is preferable because it reduces conflicts of interest, but it is not a universal legal rule for every pentest. An IT provider may perform the work if responsibilities are separated, competence is demonstrated and the conflict is managed transparently. A specific regulation, contract or assurance framework may nevertheless require an external or organisationally independent tester.
- What counts as a successful pentest?
No critical findings can be a valid result and is not, by itself, proof of either security or a poor test. Success is assessed from the adequacy of the scope, methodology, evidence and coverage, followed by risk-based remediation and retesting. There is no defensible universal number of high or critical findings that a professional pentest should produce.
- What exactly is TLPT under DORA?
Threat-led penetration testing in line with TIBER-EU [4] - a simulation of a real adversary based on threat intelligence gathered about the specific client. Required of financial entities designated by the competent authority, in Poland the Polish Financial Supervision Authority. Cycle: at least every three years. Duration: 12 to 30 weeks. Cost: PLN 200,000 to 800,000, with the regulator involved.
- Does a red team engagement include social engineering and physical entry?
It may, but only where social engineering and physical entry are expressly included in the written rules of engagement. The authorisation should define permitted actions, locations, dates, escalation contacts and stop conditions. It demonstrates the client's consent within the agreed scope; it cannot guarantee that the tester will not be questioned or detained, and it does not protect conduct outside that scope.
- What should we do if a pentest found no critical issues?
It is a result that still needs interpretation. Review the following before drawing a conclusion:
- The tester was not competent enough - a junior with a scanner instead of an experienced pentester. Check the methodology and the traces of it in the report.
- The scope was too narrow - a web-application-only pentest leaves out the network vector, Active Directory and the supply chain.
- The environment does not reflect production - testing on a sterile development or staging system with no real data and no real configuration.
- Too little time - two or three person-days for an enterprise web app is not enough.
PTES [6] and NIST SP 800-115 [8] provide guidance on documenting the scope, constraints and test plan. Zero critical findings may reflect a resilient environment, but it may also follow from narrow coverage or an unsuitable test design; the report should provide enough evidence to distinguish those possibilities.
- Is a black box test of a web application enough?
For a standard CMS application such as WordPress or Drupal, yes, black box is usually sufficient. For a bespoke application - a citizen portal, a patient portal, a purpose-built system - no: grey box or white box is preferable, because:
- Business logic and authorisation are where black box performs worst. The business logic testing and authorisation testing categories in the OWASP WSTG [9] assume knowledge of roles, permissions and the intended flow of the process. Without it the tester can only guess which transition between steps is an abuse and which is a normal path.
- PTES [6] recommends access to architecture documentation for bespoke applications in its threat modelling section.
- Effort shifts towards reconnaissance and mapping functionality in a black box engagement, leaving less time for deeper testing within the same budget.
Best practice: grey box with two or three accounts of different roles, plus black box as a separate component to validate the external attacker perspective.
- Is it safe to pentest OT and SCADA equipment?
Only with a dedicated OT-aware team working to IEC 62443 [16]. Active testing of control devices such as PLCs and RTUs can damage them physically - buffer overflows in old firmware, denial-of-service against a watchdog, race conditions. The guidance from ICS-CERT and NIST SP 800-82 Rev. 3 (also cited in our SOC article) is:
- Passive mode - traffic capture and static configuration analysis - for production.
- Active testing preferably on a representative testbed or during a planned shutdown. Testing in production requires a documented risk assessment, explicit authorisation, agreed safeguards and stop conditions appropriate to the industrial process.
- Work with the process operator - testers who do not know the industrial process can trigger an emergency stop.
- The tester wants to place an implant on our network - is that lawful and safe?
Yes on both counts, subject to three conditions:
- Written authorisation in the rules of engagement, naming the type of implant (a C2 beacon, a persistence agent), its location on the network, its lifetime and the removal mechanism. Without it, there is potential exposure under articles 267 to 269b of the Polish criminal code.
- An isolated C2 environment - implant traffic must never leave for the public internet unprotected.
- A cleanup procedure - after the test the tester removes the implant and supplies proof of removal, from C2 logs and verification on the endpoint. PTES [6] requires this explicitly in its post-exploitation cleanup section.
An implant is standard red team and TLPT technique [4]. Without one there is no way to verify detection of persistence (MITRE ATT&CK TA0003 [5]).
- How does a mobile application pentest differ from a web one?
The main methodological differences, following the OWASP MASTG (Mobile Application Security Testing Guide):
- Static analysis of the code (APK or IPA) as the first step - obfuscation, hardcoded secrets, weak cryptography.
- Dynamic analysis - runtime hooking with Frida, anti-debug bypass, certificate pinning bypass.
- The backend API - often the larger part of the attack surface, tested to the OWASP WSTG [9].
- Storage - analysis of data in SharedPreferences and UserDefaults, temporary files and the keychain.
Effort: typically 1.5 to 2 times a web pentest of comparable complexity.
Need consulting in this area?
A free 30-60 minute consultation. No obligations. We discuss needs, scale and a high-level timeline.
Related content
Other competence areas
- IT security audit
- Vulnerability scanning
- Device and system hardening
- Email security audit
- KRI compliance audit
- KSC and NIS2 audit
- GDPR compliance audit
- Information security policy
- ISMS - information security management system
- Security awareness - onsite and online
- SOC 24/7 - monitoring and response
Compliance and regulation
Bibliography and sources
All cited sources are publicly available. ISO/IEC standards, IETF RFCs, EU directives and national legal acts link to the original documents.
- [1]regulationParlament Europejski, Rada UE (2022). Dyrektywa (UE) 2022/2555 (NIS2) w sprawie środków na rzecz wysokiego wspólnego poziomu cyberbezpieczeństwa. Dziennik Urzędowy UE, L 333, 27.12.2022 · https://eur-lex.europa.eu/legal-content/PL/TXT/?uri=CELEX:32022L2555
- [2]regulationParlament Europejski, Rada UE (2022). Rozporządzenie (UE) 2022/2554 (DORA) w sprawie operacyjnej odporności cyfrowej sektora finansowego. Dziennik Urzędowy UE, L 333, 27.12.2022 · https://eur-lex.europa.eu/legal-content/PL/TXT/?uri=CELEX:32022R2554
- [3]regulationParlament Europejski, Rada UE (2024). Rozporządzenie (UE) 2024/1689 (AI Act) ustanawiające zharmonizowane przepisy dotyczące sztucznej inteligencji. Dziennik Urzędowy UE, L seria, 12.7.2024 · https://eur-lex.europa.eu/legal-content/PL/TXT/?uri=CELEX:32024R1689
- [4]guidelineEuropean Central Bank (2018). TIBER-EU Framework: Threat Intelligence-Based Ethical Red Teaming. ECB · https://www.ecb.europa.eu/paym/cyber-resilience/tiber-eu/html/index.en.html
- [5]standardMITRE Corporation (2024). MITRE ATT&CK Framework v15 (Enterprise, Mobile, ICS). MITRE · https://attack.mitre.org/
- [6]guidelinePTES Team (2014). Penetration Testing Execution Standard (PTES). pentest-standard.org · http://www.pentest-standard.org/
- [7]guidelineHerzog, P. (ISECOM) (2010). Open Source Security Testing Methodology Manual (OSSTMM) 3. Institute for Security and Open Methodologies · https://www.isecom.org/OSSTMM.3.pdf
- [8]standardScarfone, K., Souppaya, M., Cody, A., Orebaugh, A. (2008). NIST SP 800-115: Technical Guide to Information Security Testing and Assessment. National Institute of Standards and Technology · DOI: 10.6028/NIST.SP.800-115
- [9]guidelineOWASP Foundation (2024). OWASP Web Security Testing Guide (WSTG) v4.2. OWASP · https://owasp.org/www-project-web-security-testing-guide/
- [10]guidelineCREST (2023). CREST Defensible Penetration Test Standard. CREST International · https://www.crest-approved.org/
- [11]reportVerizon Business (2026). 2026 Data Breach Investigations Report (DBIR). Verizon, wydanie 19. Zbiór obejmuje incydenty zarejestrowane od 1 listopada 2024 r. do 31 października 2025 r. · https://www.verizon.com/business/resources/reports/dbir/
- [12]regulationParlament Europejski, Rada UE (2016). Rozporządzenie (UE) 2016/679 (RODO) w sprawie ochrony osób fizycznych w związku z przetwarzaniem danych osobowych. Dziennik Urzędowy UE, L 119, 4.5.2016 · https://eur-lex.europa.eu/legal-content/PL/TXT/?uri=CELEX:32016R0679
- [13]standardInternational Organization for Standardization (2022). ISO/IEC 27001:2022 - Information security, cybersecurity and privacy protection - Information security management systems - Requirements. ISO/IEC · https://www.iso.org/standard/27001
- [14]reportEuropean Union Agency for Cybersecurity (ENISA) (2025, wersja 1.2 z 9 stycznia 2026 r.). ENISA Threat Landscape 2025. ENISA · https://www.enisa.europa.eu/publications/enisa-threat-landscape-2025
- [15]reportCERT Polska / NASK (2024). Raport roczny CERT Polska 2023. NASK PIB, Warszawa · https://cert.pl/uploads/docs/Raport_CP_2023.pdf
- [16]standardInternational Electrotechnical Commission (2018). IEC 62443 - Industrial communication networks - Network and system security. IEC · https://www.iec.ch/cyber-security