Why traditional toolsets are failing your most experienced Red Team operators.
It’s a conversation we’re having with security teams more frequently. Their red team operators are skilled; the engagement is scoped properly and yet the simulation is over before it’s begun. We see a clear pattern in why this keeps happening.
Picture the scenario: your red team has planned a thorough engagement. The objectives are well-defined, the operators are experienced, and the rules of engagement are signed off. Execution begins and within the first half-hour, the endpoint defence has flagged the activity, the process is killed, and the engagement has effectively stalled before it produced a single meaningful finding.
When security teams realise their standard C2 frameworks and public toolsets are flagged before establishing initial access, the immediate reaction is usually to pivot internal resources into custom tool development.
However, custom payload R&D is a double-edged sword.
Read more: Simulating the Adversary: Why Elite Red Teams are Moving Beyond Open-Source Tools with OST.
The symptoms we see most often
Before exploring what’s driving this, it’s worth flagging the specific signals that tell us a red team programme is hitting this wall. They’re usually some combination of the following:
Flagged at first execution
The payload or stager is detected and killed before any meaningful activity occurs. The engagement never gets past initial access.
Repeat false positives on known tools
Standard offensive security tools that worked reliably two years ago are now instant detections, regardless of configuration.
Shallow engagement findings
Reports consistently identify the same early-stage detections but never surface the deeper risks — lateral movement, privilege escalation, data access.
Eroding executive confidence
Stakeholders start to question the value of the red team programme when every engagement tells the same limited story.
The Core Problem
The red team’s job is to answer the question: “Could a sophisticated attacker reach our critical assets undetected?” But that question never gets answered if the engagement is over at the first step. The security posture appears strong, but the test that would reveal its real gaps hasn’t actually been run.
Why the landscape has shifted
EDR technology has matured significantly in a short window theres been a fundamental shift in defensive capabilities. The platforms deployed by enterprise organisations today don’t operate primarily on file signatures. They observe behaviour at runtime, monitor API calls at a kernel level, and real-world behavioural heuristics that feed into cloud-backed detection models that learn from activity across enormous fleets of endpoints.
The practical consequence of this is that many of the offensive tools that formed the foundation of red team work and that are still widely used have had their behavioural DNA catalogued in exhaustive detail. When those tools execute, their activity patterns are recognised not because a signature matched, but because the sequence of actions they generate is known.
- The more widely a tool is used in the red team community, the more thoroughly defenders have studied it. Broad adoption and broad detection go hand in hand as popularity becomes a liability.
- Obfuscation is not evasion, so changing how a payload looks is not the same as changing what it does. Behavioural detection sees through surface-level variation because it’s watching execution, not appearance.
- EDR vendors invest continuously in detection research, whilst community offensive tools evolve more slowly and publicly.
What this means for your Red Team Programme
This is where the problem becomes strategically significant. If your red team can only operate with commodity toolsets, you’re not testing your defences against a capable adversary, you’re testing them against a known and heavily-studied set of tools. When those tools get caught, it’s a confirmation of signature-based detection, not a measure of real security posture.
The threat actors you actually need to worry about are not using Mimikatz in default configuration. They’re not leaving the fingerprints that standard toolsets generate. Your red team needs to be able to replicate that level of operational sophistication or the engagement isn’t answering the question it’s supposed to answer.
The organisations we work with that are getting the most value from their red team programmes share a common characteristic: they’ve aligned their offensive toolset to the maturity level of their target environment.
In a heavily-defended enterprise, that means moving away from commodity tools and toward purpose-built offensive capability designed with operational security as a primary requirement.
The Enterprise dilemma: Custom R&D vs. Threat Emulation
When faced with persistent EDR flags, security leaders generally fall into one of two traps:
The Red Team Trade-Off
The Public Tooling Trap
- Reliance on open-source frameworks / scripts.
- Signatures instantly flagged by EDRs.
- High burn rate on initial execution.
Custom R&D drain
- Red team spends 80% of time engineering C# / C++.
- High maintenance cost for internal codebase.
- Focus moves away from emulating real TTPs.
The primary duty of an enterprise red team is not to build custom C2 software from scratch. Their primary objective is to evaluate organisational resilience, stress-test incident response protocols, and expose gaps in detections.
To break this cycle, enterprise security teams must adopt purpose-built, evasive tooling designed specifically to bypass modern endpoint mechanisms during initial execution.
A different kind of tooling
Outflank Security Tooling (OST)
Outflank, developed specifically for red teams operating against environments with mature endpoint defences.
Where community tools are designed to be capable, OST is designed to be undetected. It’s built from the ground up with operational security as the design requirement not an afterthought. The result is tooling that allows red team operators to conduct realistic, sophisticated simulations in environments where commodity tools simply don’t survive first contact.
For security teams that are consistently hitting the detection wall, OST represents a meaningful step change in what’s possible. It’s not a tweak to existing tradecraft it’s a different category of capability.
Read more: Outflank Security Tooling (OST) overview
Not sure where your Red Team stands?
Visit our self-assessment checklist, a quick diagnostic to help you identify whether your red team programme is equipped to operate in a hardened EDR environment.
Next Steps
Ready to elevate your red teaming capabilities and move past the limitations of open-source tools? Book a Consultation with S4 Applications today to learn more about OST.

